A method, apparatus and equipment for training large models
By replacing less important network layers with lightweight layers and adjusting the parameters of more important network layers in a large model, the problems of data privacy leakage and resource consumption during the training of large models are solved, and efficient and secure cross-domain fine-tuning is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
- Filing Date
- 2024-06-07
- Publication Date
- 2026-05-26
AI Technical Summary
During the training of large models, how can we ensure model performance while avoiding data privacy leaks and reducing resource consumption, especially during cross-domain fine-tuning?
By acquiring the importance information of each network layer in the pre-trained large model, the less important network layers are replaced with lightweight network layers, and the parameters of the more important network layers are allowed to be adjusted to generate a simulation model. The model is then trained using business data from the data owner and finally, the target large model is generated.
It improves the accuracy of processing results for large target models, reduces training resource consumption, and ensures data privacy and security, avoiding the leakage of business data and the privacy of large models.
Smart Images

Figure CN118586284B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large model technology, and in particular to a large model training method, apparatus and equipment. Background Technology
[0002] With the continuous development of technology, large-scale modeling technology is being applied more and more widely across various industries. Currently, to further improve the performance of large-scale models in business applications, the model provider typically sends a pre-trained large-scale model to the data owner. The data owner can then fine-tune the pre-trained large-scale model using their own business data; alternatively, the data owner can send their own business data to the model provider, allowing the provider to fine-tune the pre-trained large-scale model using that data. However, this approach fails to meet the requirement that data from all parties remain within their respective domains. Furthermore, since large-scale models typically have a large number of parameters, directly fine-tuning a pre-trained large-scale model using business data often consumes significant resources.
[0003] Therefore, how to ensure that the trained large model has good performance while avoiding data privacy leaks during the training process and reducing the resources required for model training has become a technical problem that needs to be solved. Summary of the Invention
[0004] The large model training method, apparatus, and device provided in the embodiments of this specification can avoid data privacy leakage during the large model training process while ensuring that the trained large model has good performance, and can also reduce the resources consumed in the model training process.
[0005] To solve the above-mentioned technical problems, the embodiments in this specification are implemented as follows:
[0006] This specification provides an embodiment of a large model training method, including:
[0007] Obtain information on the importance of each pre-defined network layer in a pre-trained large model;
[0008] Based on the importance information, settings are applied to the first and second network layers in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the settings are used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer;
[0009] Send the simulation model to the device of the data owner;
[0010] The system receives adjusted parameter data of the second network layer from the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using service data from the data owner.
[0011] Based on the adjusted parameter data of the second network layer and the pre-trained large model, a target large model is generated.
[0012] This specification provides an embodiment of a large model training method, including:
[0013] The simulation model of the pre-trained large model is obtained from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layer according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0014] The simulation model is trained using business data from the data owner to obtain the adjusted parameter data for the second network layer;
[0015] The adjusted parameter data of the second network layer is sent to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0016] This specification provides an embodiment of a large model training device, comprising:
[0017] The first acquisition module is used to acquire information on the importance of each preset network layer in the pre-trained large model.
[0018] The setting module is used to perform setting processing on the first network layer and the second network layer in the preset network layers according to the importance information, so as to obtain the simulation model of the pre-trained large model; wherein, the setting processing is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer;
[0019] The first sending module is used to send the simulation model to the device of the data owner;
[0020] The receiving module is used to receive the adjusted parameter data of the second network layer fed back by the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using the service data at the data owner;
[0021] The generation module is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0022] This specification provides an embodiment of a large model training device, comprising:
[0023] The acquisition module is used to acquire a simulation model of a pre-trained large model from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layers according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer;
[0024] The training module is used to train the simulation model using business data from the data owner to obtain the adjusted parameter data of the second network layer.
[0025] The first sending module is used to send the adjusted parameter data of the second network layer to the model provider's device; the model provider's device is used to generate the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0026] This specification provides an embodiment of a large model training device, comprising:
[0027] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0028] Obtain information on the importance of each pre-defined network layer in a pre-trained large model;
[0029] Based on the importance information, settings are applied to the first and second network layers in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the settings are used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer;
[0030] Send the simulation model to the device of the data owner;
[0031] The system receives adjusted parameter data of the second network layer from the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using service data from the data owner.
[0032] Based on the adjusted parameter data of the second network layer and the pre-trained large model, a target large model is generated.
[0033] This specification provides an embodiment of a large model training device, comprising:
[0034] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0035] The simulation model of the pre-trained large model is obtained from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layer according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0036] The simulation model is trained using business data from the data owner to obtain the adjusted parameter data for the second network layer;
[0037] The adjusted parameter data of the second network layer is sent to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0038] At least one embodiment provided in this specification can achieve the following beneficial effects:
[0039] The model provider can modify the less important first network layer to a lightweight one based on the importance information of each preset network layer in the pre-trained large model, and adjust the parameters of the more important second network layer to obtain a simulation model of the pre-trained large model. Subsequently, after the data owner trains this simulation model using business data and provides feedback on the adjusted parameters for the second network layer, the model provider can combine the adjusted parameters of the second network layer with the pre-trained large model to generate the target large model. This not only improves the accuracy of the target large model's processing results for the data owner's business data, but also avoids the leakage of the model provider's complete pre-trained large model and the data owner's business data, ensuring data privacy and security when fine-tuning the large model across domains. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a schematic diagram illustrating an application scenario of a large model training method provided in the embodiments of this specification;
[0042] Figure 2 A flowchart illustrating a large model training method provided in the embodiments of this specification;
[0043] Figure 3 A flowchart illustrating another large model training method provided in the embodiments of this specification;
[0044] Figure 4 The embodiments provided in this specification correspond to Figure 2 and Figure 3 A schematic diagram of the swimlane process for training large models in China;
[0045] Figure 5 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a large model training device;
[0046] Figure 6 The embodiments provided in this specification correspond to Figure 3 A schematic diagram of the structure of a large model training device;
[0047] Figure 7 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a large model training device;
[0048] Figure 8 The embodiments provided in this specification correspond to Figure 3 A schematic diagram of the structure of a large model training device. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of one or more embodiments of this specification clearer, the technical solutions of one or more embodiments of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of one or more embodiments of this specification.
[0050] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0051] In existing technologies, traditional methods for fine-tuning pre-trained large models mainly fall into two categories: closed-source fine-tuning and open-source fine-tuning. Closed-source fine-tuning typically requires the data owner to send their business data to the model provider, who then uses this data to fine-tune the pre-trained large model. This method often fails to guarantee the security and privacy of the data owner's business data. Open-source fine-tuning, on the other hand, usually involves the model provider sending the pre-trained large model to the data owner, allowing the data owner to fine-tune it using their own business data. However, this approach often allows malicious actors to obtain the complete structure of the pre-trained large model, making it more vulnerable to white-box attacks and threatening its security. Furthermore, since the pre-trained large models provided by model providers often have hundreds of billions to trillions of parameters, fine-tuning the entire pre-trained large model using business data often consumes a significant amount of resources.
[0052] Currently, some model providers first use lossy compression techniques to remove half of the network layers from the pre-trained large model, thus transforming it into an emulator. This emulator is then passed to the data owner. The data owner can then train a specific adapter based on this emulator and their own business data, and feed the trained adapter back to the model provider. The model provider can then plug the adapter into the pre-trained large model to complete the fine-tuning process. However, because the emulator obtained by the data owner is lossy compressed, the performance of the final fine-tuned pre-trained large model generated based on this emulator is relatively poor.
[0053] Therefore, the current technical problem to be solved is how to ensure that the trained large model has good performance while avoiding the leakage of privacy of both business data and the large model, and reducing the resources required for model training.
[0054] To address the shortcomings of existing technologies, this solution provides the following embodiments:
[0055] Figure 1 This is a schematic diagram illustrating an application scenario of a large model training method provided in the embodiments of this specification.
[0056] like Figure 1 As shown, when a model provider needs to provide services related to a large model to a data owner, it can use its device 101 to obtain the importance information of each preset network layer in the pre-trained large model. Then, based on the importance information, it performs setting processing on the first and second network layers in the preset network layers to obtain a simulation model of the pre-trained large model. The setting processing can be used to change the first network layer to a preset lightweight network layer and to indicate that the parameters of the second network layer can be adjusted. The importance of the second network layer is higher than that of the first network layer. Subsequently, the model provider's device 101 can send the simulation model to the data owner's device 102.
[0057] After obtaining the simulation model of the pre-trained large model from the model provider's device 101, the data owner's device 102 can use the business data from the data owner to train the simulation model in order to obtain the adjusted parameter data of the second network layer. By sending the adjusted parameter data of the second network layer to the model provider's device 101, the model provider's device 101 can generate the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0058] Next, a large model training method provided in the embodiments of the specification will be described in detail with reference to the accompanying drawings:
[0059] Figure 2 This is a flowchart illustrating a large model training method provided in an embodiment of this specification. From a programming perspective, the entity executing this process can be the model provider's device, or the application program on the model provider's device. Figure 2 As shown, the process may include the following steps:
[0060] Step 202: Obtain the importance information of each preset network layer in the pre-trained large model.
[0061] In the embodiments of this specification, a large model can refer to a machine learning model with a large number of parameters and a complex structure. Because large models have a large number of parameters and can handle massive amounts of data, they can be applied to handle large-scale data and complex problems. A pre-trained large model can refer to a large model that has been pre-trained using a large amount of training data based on unsupervised or self-supervised learning methods, thereby storing a large amount of knowledge. Pre-trained large models typically have strong generalization capabilities, and therefore can be applied to various downstream business tasks through simple fine-tuning.
[0062] In practical applications, depending on the structure and function of the large model, it can be including but not limited to large language model (LLM), computer vision (CV) large model, automatic speech recognition (ASR) large model, recommendation system large model, reinforcement learning (RL) large model, generative adversarial network (GAN) large model, dialogue system large model, etc.
[0063] In the embodiments of this specification, the pre-trained large model often has multiple preset network layers, and the degree of influence of different preset network layers on the performance of the pre-trained large model often varies. Generally, the greater the influence of a preset network layer on the performance of the pre-trained large model, the higher its importance in the pre-trained large model. Based on this, importance information reflecting the importance of each preset network layer in the pre-trained large model can be determined for use when adjusting the model structure of the pre-trained large model to generate its simulation model.
[0064] Step 204: Based on the importance information, perform setting processing on the first network layer and the second network layer in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the setting processing is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0065] In this embodiment, to ensure the security of the pre-trained large model, it is necessary to avoid providing the fully structured pre-trained large model to the data owner. Therefore, the structure of the pre-trained large model needs to be modified to generate a simulation model of the pre-trained large model. Specifically, to ensure that the performance of the simulation model is close to that of the pre-trained large model and to reduce the resource consumption of the data owner when training the simulation model using business data, the less important preset network layers (e.g., the first network layer) can be replaced with preset lightweight network layers. The data owner is allowed to adjust the parameters of the more important preset network layers (e.g., the second network layer), and is prohibited from adjusting the parameters of other preset network layers, thus obtaining the simulation model of the pre-trained large model. The number of parameters in the preset lightweight network layers is usually less than that in the first network layer, reducing the number of parameters in the simulation model, thereby reducing the resource consumption required by the data owner to train the simulation model and improving model training efficiency.
[0066] In practical applications, the second network layer and the first network layer are usually different preset network layers. However, if the number of preset network layers in the pre-trained large model is limited, it is possible to use the same preset network layer as both the first and second network layers. In this case, the preset network layer can usually be replaced with a preset lightweight network layer first, and then the data owner can be allowed to adjust the parameters of the preset lightweight network layer. There are no specific limitations on this.
[0067] Step 206: Send the simulation model to the device of the data owner.
[0068] Step 208: Receive the adjusted parameter data of the second network layer fed back by the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using the service data at the data owner.
[0069] In the embodiments of this specification, the business data used by the data owner when training the simulation model can be set according to actual needs. This business data is usually related to the business needs processed by the data owner when subsequently calling the target large model; therefore, the specific types and content of the business data are not limited. For ease of understanding, examples are provided. For instance, suppose the data owner needs to use the target large model to determine the sentiment type of user comment data, then the business data may include historical user comment data and data reflecting the sentiment type. Alternatively, if the data owner needs to use the target large model to determine whether a user-submitted ID image is forged, then the business data may include historical user ID image data and data reflecting whether it is forged. Or, if the data owner needs to use the target large model to automatically generate comment information based on article information, then the business data may also include historical article information and corresponding historical article comment information, etc. Since subsequent... Figure 3 The solution in the document will explain in detail the implementation principle of how the data owner obtains the adjusted parameter data of the second network layer, so it will not be elaborated here.
[0070] Step 210: Generate the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0071] In the embodiments of this specification, the pre-trained large model typically contains historical parameter data of the second network layer. Since the historical parameter data of the second network layer usually does not carry knowledge data of the business data of the data owner, the adjusted parameter data of the second network layer that has learned the knowledge of the business data and the pre-trained large model can be combined to generate the target large model, so as to improve the processing performance of the target large model for the business data of the data owner.
[0072] Figure 2 The proposed method replaces less important network layers in a pre-trained large model with pre-defined lightweight network layers as the simulation model. Compared to existing methods that directly delete half of the network layers in a pre-trained large model, this approach improves the consistency of structure and performance between the simulation model and the pre-trained large model, while reducing the number of parameters in the simulation model. This allows for improved effectiveness of the adjusted parameter data of the more important second network layer generated from the simulation model, while reducing the resources required for model training. This ensures that the target large model generated based on the adjusted parameter data of the second network layer and the pre-trained large model achieves good performance. Furthermore, since the model provider cannot access the business data from the data owner, and the data owner cannot access the structurally complete pre-trained large model, privacy leaks between the business data and the large model are avoided, ensuring data privacy and security during cross-domain fine-tuning of the large model.
[0073] based on Figure 2 In addition to the method described in the embodiments of this specification, some specific implementation schemes of the method are also provided, which will be described below.
[0074] To facilitate understanding, a method for obtaining the importance information of preset network layers in a pre-trained large model is provided here. Specifically, step 202: obtaining the importance information of each preset network layer in the pre-trained large model may include:
[0075] Using a reinforcement learning algorithm, a reinforcement learning reward value is determined for implementing a first action policy on the pre-trained large model; wherein, the first action policy is used to instruct the replacement of the third network layer in the preset network layers within the pre-trained large model with a preset lightweight network layer, and the reinforcement learning reward value is positively correlated with the loss value generated by the pre-trained large model processing the first sample data after implementing the first action policy.
[0076] Based on the reinforcement learning reward value, the historical importance score of the third network layer is updated to obtain the target importance score of each preset network layer in the pre-trained large model; wherein, the reinforcement learning reward value is positively correlated with the target importance score of the third network layer.
[0077] In the embodiments described in this specification, reinforcement learning (RL) is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an agent learning strategies to maximize rewards or achieve specific goals during interactions with its environment. Reinforcement learning views learning as a trial-and-error evaluation process. The agent selects an action for the environment; after the environment accepts the action, its state changes, generating a reinforcement signal (reward or penalty) fed back to the agent. Based on the reinforcement signal and the current state of the environment, the agent selects its next action, aiming to increase the probability of receiving a positive reinforcement (reward). The selected action not only affects the immediate reinforcement value but also the environment's state at the next moment and the final reinforcement value.
[0078] Based on this, when it is necessary to determine the importance information of each preset network layer in the pre-trained large model, the replacement of a portion of the preset network layers (i.e., the third network layer) in the pre-trained large model with a preset lightweight network layer can be used as the first action policy to be implemented in the reinforcement learning algorithm. The reinforcement learning reward value corresponding to the first action policy is calculated by combining the loss value generated by the pre-trained large model after implementing the first action policy. This allows the importance score of the replaced preset network layer (i.e., the third network layer) involved in the first action policy to be updated based on the reinforcement learning reward value, so as to use the target importance score of the preset network layer as the importance information of the preset network layer.
[0079] In practical applications, the loss value generated by the pre-trained large model after implementing the first action strategy and processing the first sample data can be used as an indicator to measure the difference between the model's prediction and the true label of the first sample data. The level of the loss value reflects the degree of fit of the model to the training data under the current parameters. Generally speaking, the lower the loss value, the smaller the difference between the model's prediction and the true label, indicating better model performance and a more accurate prediction of the target variable. In practical applications, the types of loss values can be various, such as those calculated based on L1 norm loss, mean squared error loss (MSELoss), cross-entropy loss (CrossEntropyLoss), and KL divergence loss (KLDivLoss), etc., without specific limitations.
[0080] In the embodiments of this specification, if the loss value generated by the pre-trained large model after implementing the first action strategy when processing the first sample data is small, it indicates that the replaced third network layer corresponding to the first action strategy has a small impact on the performance of the pre-trained large model, that is, the importance of the third network layer is low, which can lead to a lower target importance score for the third network layer. Clearly, there can be a positive correlation between the loss value and the target importance score of the third network layer. Furthermore, since it is necessary to avoid replacing the preset network layer with a preset lightweight network layer (i.e., to avoid executing the first action strategy with a smaller loss value), there can also be a positive correlation between the loss value and the reinforcement learning reward value corresponding to the first action strategy. Based on this, there can also be a positive correlation between the reinforcement learning reward value and the target importance score of the third network layer.
[0081] To facilitate understanding, a specific implementation of calculating the reinforcement learning reward value corresponding to the first action policy is provided here.
[0082] In the embodiments of this specification, the number of the first action policies can typically be greater than 1. Correspondingly, determining the reinforcement learning reward value for implementing the first action policy on a pre-trained large model using a reinforcement learning algorithm can include:
[0083] Based on the historical importance scores of each preset network layer in the pre-trained large model, a first number of third network layers that need to be replaced by preset lightweight network layers are selected from the preset network layers to obtain a first action policy; wherein, the probability of the preset network layer being selected as the third network layer is negatively correlated with the historical importance score of the preset network layer.
[0084] The first sample data is processed using the pre-trained large model after implementing the first action strategy to obtain the loss value corresponding to the first action strategy.
[0085] Based on the loss value corresponding to each of the first action strategies, the reinforcement learning reward value for implementing each of the first action strategies on the pre-trained large model is calculated according to the first formula.
[0086] The first formula can be:
[0087]
[0088] r i The loss is the reinforcement learning reward value of implementing the i-th first action policy on a pre-trained large model. i Let loss be the loss value corresponding to the i-th first action policy. j Let be the loss value corresponding to the j-th first action policy, T be the number of first action policies, and j be a positive integer in the range of 1 to T. loss for e i Power of 1.
[0089] In the embodiments described in this specification, to better compare and measure the importance of different preset network layers, multiple first action policies are typically generated during a single reinforcement learning algorithm execution. In practical applications, the number of third network layers to be replaced indicated by different first action policies is usually consistent. If each first action policy indicates multiple third network layers to be replaced, the third network layers corresponding to different first action policies are usually not completely consistent. However, different first action policies can have one or more identical third network layers, providing good flexibility.
[0090] In the embodiments described in this specification, it is generally necessary to avoid replacing highly important preset network layers with preset lightweight network models to ensure the performance of large models. Therefore, the probability of a preset network layer being selected as the third network layer to be replaced can be negatively correlated with its historical importance score. It is worth noting that the higher the historical importance score of a preset network layer, the lower its probability of being selected as the third network layer, but it is not equal to 0. Therefore, a highly important preset network layer may also be selected as the third network layer, thereby continuously updating its historical importance score.
[0091] In practical applications, there are multiple ways to select a third network layer from the various preset network layers based on their historical importance scores. For example, the sampling probability pi for any preset network layer mi can be the absolute value of the difference between a randomly selected value from the range of 0 to sigmoid(si) and 1, where si can be the historical importance score of the preset network layer mi, and sigmoid(si) is an S-shaped growth curve with a value range of 0 to 1. Since a higher historical importance score si for the preset network layer mi results in a larger sigmoid(si) value, the higher the probability that the randomly selected value from the range of 0 to sigmoid(si) will be larger. Consequently, the absolute value of the difference between 1 and sigmoid(si) is more likely to be smaller, thus reducing the probability that the preset network layer mi will be selected as the third network layer. Obviously, in this case, the probability of the preset network layer being selected as the third network layer can have a negative correlation with its historical importance score. Alternatively, the historical importance scores of each preset network layer can be mapped to values between 0 and 1, and the difference between 1 and the normalized historical importance score of each preset network layer can be used as the probability of it being selected as the third network layer. In this case, the probability of a preset network layer being selected as the third network layer can also have a negative correlation with its historical importance score. Of course, other implementation methods are also possible, and no specific limitations are imposed on them.
[0092] In practical applications, the probability of each preset network layer being selected as the third network layer can be made the same. In this case, the selection of each preset network layer as the third network layer can also be unrelated to the historical importance score of the preset network layer, and no specific restrictions are imposed on this.
[0093] In the embodiments of this specification, for each first action strategy, the first sample data needs to be processed after the first action strategy is implemented once on the pre-trained large model to obtain the loss value of the pre-trained large model implementing the first action strategy. That is, each first action strategy has a corresponding loss value. The first sample data can be general data that the model provider can obtain; therefore, the specific content and meaning of the first sample data are not limited. It is understood that after processing the first sample data using the pre-trained large model after implementing the first action strategy to obtain the loss value, the parameters of the currently existing preset lightweight network layer in the pre-trained large model may or may not be updated according to the loss value; neither is specifically limited in this regard.
[0094] In the embodiments of this specification, since the reinforcement learning reward values for implementing each first action strategy for the pre-trained large model can be calculated according to the first formula described above, they will not be elaborated here.
[0095] To facilitate understanding, this paper provides a specific implementation method for determining the target importance score of each preset network layer in the pre-trained large model based on the reinforcement learning reward value corresponding to each first action policy.
[0096] Specifically, updating the historical importance score of the third network layer based on the reinforcement learning reward value to obtain the target importance score of each preset network layer in the pre-trained large model may include:
[0097] For any of the third network layers, the target importance score of the third network layer is calculated according to the second formula based on the reinforcement learning reward value and the historical importance score of the third network layer.
[0098] The second formula is:
[0099] s cd =s hd +∑r q ,
[0100] s cd s is the target importance score of the third network layer. hd r is the historical importance score of the third network layer. q This is used to indicate the reinforcement learning reward value corresponding to the first action policy that replaces the third network layer.
[0101] In the embodiments of this specification, the reinforcement learning reward value corresponding to each first action policy is typically only related to the third network layer that the first action policy indicates to be replaced, and is unrelated to other preset network layers that have not been replaced. Furthermore, as the preceding analysis shows, the reinforcement learning reward value corresponding to each first action policy is positively correlated with the target importance score of the third network layer that the first action policy indicates to be replaced. Therefore, for each third network layer, the historical importance score of the third network layer can be added to the reinforcement learning reward values corresponding to each first action policy indicating the replacement of the third network layer to obtain the target importance score of the third network layer.
[0102] In practical applications, if multiple first action policies all indicate the replacement of the same third network layer, the maximum, minimum, or average value of the reinforcement learning reward values corresponding to the multiple first action policies can be added to the historical importance score of the third network layer to obtain the target importance score of the third network layer. No specific limitation is made in this regard.
[0103] It is understandable that, since the other preset network layers besides the third network layer are not replaced during the reinforcement learning algorithm processing, the reinforcement learning reward value corresponding to each first action policy does not affect the importance scores of the other preset network layers. This means that there is no need to update the historical importance scores of the other preset network layers besides the third network layer, but their historical importance scores can be directly used as the target importance scores, thereby obtaining the target importance scores of all preset network layers. This will not be elaborated further.
[0104] In the embodiments of this specification, in order to further improve the performance of the target large model obtained through training, the pre-trained large model and the second sample data can be combined to pre-train the preset lightweight network layers required for the simulation model that generates the pre-trained large model, so as to ensure that each preset lightweight network layer can learn the knowledge of a large amount of sample data.
[0105] Based on this Figure 2 The method described herein may further include:
[0106] Based on the target importance scores of each preset network layer in the pre-trained large model, a second number of fourth network layers that need to be replaced by lightweight network layers are selected from the preset network layers to obtain a second action strategy; wherein, the probability of the preset network layer being selected as the fourth network layer is negatively correlated with the target importance score of the preset network layer.
[0107] Based on the second action strategy, the fourth network layer in the pre-trained large model is replaced by a lightweight network layer in the lightweight network layer set that has a unique correspondence with the fourth network layer, thereby obtaining the pre-trained large model after implementing the second action strategy; wherein, each lightweight network layer in the lightweight network layer set corresponds one-to-one with each preset network layer.
[0108] The pre-trained large model after implementing the second action strategy is used to process the second sample data to obtain the loss value corresponding to the second action strategy.
[0109] Based on the loss value corresponding to the second action strategy, the parameters of the lightweight network layer in the pre-trained large model after implementing the second action strategy are adjusted to obtain a preset lightweight network layer that has a unique correspondence with the fourth network layer.
[0110] The lightweight network layer in the set of lightweight network layers that corresponds to other preset network layers is determined as a preset lightweight network layer that has a unique correspondence with the other preset network layers.
[0111] Specifically, the setting process is used to change the first network layer to a preset lightweight network layer that has a unique correspondence with the first network layer; the first action strategy is used to instruct the third network layer to be changed to a preset lightweight network layer that has a unique correspondence with the third network layer.
[0112] In the embodiments described in this specification, typically, a lightweight network layer with rich knowledge and unique correspondences can be trained for each preset network layer in the pre-trained large model. For ease of understanding, the pre-trained large model can be denoted as G = {g1, g2, ..., g...}. n}, where g i The i-th preset network layer can be represented. The set of lightweight network layers can be denoted as H = {h1, h2, ..., h...}. n}, where h i This can represent the i-th lightweight network layer or a preset lightweight network layer, where the i-th lightweight network layer and the i-th preset lightweight network layer are essentially the same network layer before and after the parameter update, respectively. i With g i There is a unique correspondence between them. In practical applications, whether it's the reinforcement learning algorithm processing, the deep learning process, or the simulation model generation process, only h is typically used. i Replace g i This ensures the effectiveness of the model training process. In addition, the target importance scores of each pre-defined network layer in the pre-trained large model can be denoted as S = {s1, s2, ..., s...}.n}, where s i It can represent the target importance score of the i-th preset network layer.
[0113] Since generating simulation models for pre-trained large models requires prioritizing the replacement of less important pre-defined network layers with pre-defined lightweight network layers, in deep learning, parameter optimization can also be prioritized for the pre-defined lightweight network layers corresponding to less important pre-defined network layers. Based on this, the probability of a pre-defined network layer being the fourth network layer to be replaced can be negatively correlated with its target importance score. In practical applications, the fourth network layer with the lowest target importance score can simply be selected as the fourth network layer, or other methods can be used to select the pre-defined network layers to be the fourth network layer; no specific limitations are imposed. At this point, a second action strategy can be generated to instruct the replacement of each fourth network layer with the corresponding lightweight network layer.
[0114] To make this easier to understand, let's illustrate with an example. Suppose that the second action policy instruction will be g. x With g y Replace with h x with h y Then, in the pre-trained large model, the preset network layer g x Replace with a lightweight network layer h x and the preset network layer g y Replace with a lightweight network layer h y Afterwards, a pre-trained large model with the second action strategy implemented can be obtained. Subsequently, after processing the second sample data using the pre-trained large model with the second action strategy implemented to obtain the loss value, minimizing the loss value can be used as the deep learning objective for the lightweight network layer h. x with h y Perform parameter adjustments and then apply the adjusted lightweight network layer h. x with h y Each is used as a pre-defined network layer g. x and g y A pre-defined lightweight network layer with a unique correspondence.
[0115] In practical applications, typically only one second action policy is generated in each deep learning process. To ensure model training efficiency, multiple deep learning processes can be executed consecutively. Furthermore, the fourth network layer corresponding to the second action policy generated in each deep learning process can be either the same or different; no specific restrictions are imposed on this.
[0116] In practical applications, the aforementioned deep learning and reinforcement learning processes can often be executed concurrently. This helps ensure the good performance of the final generated lightweight network layer and the accuracy of the target importance scores obtained from the network layer. Based on this, Figure 2 The methods mentioned may also include:
[0117] Obtain the pre-trained large model, the importance score set of each preset network layer, and the lightweight network layer set.
[0118] The types of processing required to be performed on the pre-trained large model are determined, including reinforcement learning processing or deep learning processing.
[0119] If deep learning processing needs to be performed on a pre-trained large model, it can be performed according to the implementation method provided in the above embodiment, and the lightweight network layer set can be updated using the obtained preset lightweight network layers. Subsequent reinforcement learning processing can be implemented based on the latest version of the lightweight network layer set.
[0120] If reinforcement learning processing needs to be performed on a pre-trained large model, it can be performed according to the implementation method provided in the above embodiment, and the importance score set can be updated using the obtained target importance score. Subsequent deep learning processing can be implemented based on the latest version of the importance score set.
[0121] In practical applications, the target number of executions of reinforcement learning and / or deep learning processes can be preset. After reaching the target number of executions, the importance information of each preset network layer in step 202 is generated based on the latest version of the importance score set. Furthermore, a simulation model of the pre-trained large model is generated using the latest version of the lightweight network layer set. Alternatively, reinforcement learning and deep learning processes can be stopped when the accuracy of the latest version of the importance score set is good and / or the performance of the latest version of the lightweight network layer set is good. In this case, a simulation model of the pre-trained large model can be generated based on the latest version of the importance score set and the lightweight network layer set. Further details are omitted here.
[0122] In this embodiment of the specification, step 204: Based on the importance information, setting processing is performed on the first network layer and the second network layer in the preset network layers to obtain the simulation model of the pre-trained large model, which may specifically include:
[0123] Based on the importance information, the third number of preset network layers with the highest importance in the pre-trained large model are selected to obtain the second network layer.
[0124] An adapter for a pre-trained large model is generated based on the second network layer; wherein the parameters of the adapter can be adjusted.
[0125] Based on the importance information, the fourth number of preset network layers with the lowest importance in the pre-trained large model are selected to obtain the first network layer.
[0126] The first network layer is replaced by a preset lightweight network layer to obtain a simulation model carrying an adapter.
[0127] Correspondingly, step 208: receiving the adjusted parameter data of the second network layer fed back by the device of the data owner may specifically include: receiving the adjusted parameter data for the adapter fed back by the device of the data owner.
[0128] In the embodiments of this specification, an adapter is a trainable module that can be plugged into a pre-trained large model or a simulation model of the pre-trained large model. Based on this, when the data owner needs to fine-tune the pre-trained large model, an adapter can be generated based on the third most important preset network layer (i.e., the second network layer) in the pre-trained large model, allowing the data owner to adjust the parameters of the adapter. In addition, the fourth least important preset network layer (i.e., the first network layer) needs to be replaced with a preset lightweight network layer. At this point, the compression processing of the pre-trained large model has been completed, resulting in a simulation model carrying the adapter. Subsequently, after the data owner trains the simulation model carrying the adapter, they can feed back the adjusted parameter data for the adapter to the model provider, which is convenient and quick. In practical applications, the third and fourth quantities can be set according to actual needs. For example, if higher model performance is required, the third quantity can be set larger, and / or the fourth quantity can be set smaller. If higher training efficiency is required for the model, the fourth quantity can be set to a larger value, etc., which will not be elaborated further.
[0129] Specifically, the step of generating the adapter for the pre-trained large model based on the second network layer may include:
[0130] The adapter is generated based on the second network layer that has been granted parameter adjustment permissions; or, the adapter is obtained by setting up a plugin for fine-tuning the pre-trained large model for the second network layer.
[0131] Correspondingly, the adjusted parameter data for the adapter fed back by the receiving device can specifically include:
[0132] Receive adjusted parameter data for the second network layer that has been granted parameter adjustment authority from the device of the data owner; or, receive adjusted parameter data for the plug-in from the device of the data owner.
[0133] Correspondingly, step 210: Based on the adjusted parameter data of the second network layer and the pre-trained large model, generate the target large model, which may specifically include:
[0134] The target large model is obtained by replacing the historical parameter data of the second network layer in the pre-trained large model with the adjusted parameter data of the second network layer for which parameter adjustment authority has been granted; or, the target large model is generated by combining the adjusted parameter data of the plugin with the historical parameter data of the pre-trained large model.
[0135] This specification provides several implementation methods for the adapter in the embodiments. For example, the second network layer can be directly used as the network layer within the adapter, allowing the data owner to directly adjust the parameters of the second network layer. Subsequently, the model provider will use the parameter-updated second network layer to replace the second network layer in the pre-trained large model to obtain the target large model, which is convenient and quick. Alternatively, a plugin carrying adjustable parameters can be additionally set for the second network layer, and this plugin can be used as the adapter. In this case, the data owner will not directly adjust the parameters of the second network layer, but will adjust the parameters within the additional plugin. Subsequently, the model provider can directly connect the parameter-adjusted plugin to the pre-trained large model to obtain the target large model, offering good flexibility.
[0136] In practical applications, an emulator can be a small model obtained by the model provider through lossy compression pre-training a large model. The emulator mimics the input-output behavior of the pre-trained large model, but with performance deviations. The model provider then provides both the emulator and adapter to the data owner for their operation. Therefore, the emulator and adapter can constitute a simulation model with an adapter. Specifically, if the second network layer is directly used as the adapter, the emulator can include all network layers in the simulation model except for the second network layer. However, if a plugin with adjustable parameters specifically designed for the second network layer is used as the adapter, the emulator will include the second network layer, a preset lightweight network layer, and other preset network layers within the simulation model, but will not include the plugin specifically designed for the second network layer. This will not be elaborated further.
[0137] In this embodiment of the specification, after step 210: generating the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model, it may further include:
[0138] Obtain the call request from the data owner for the target large model; the call request carries the business data to be processed at the data owner's location.
[0139] The target large model is used to process the business data to be processed, and the business data processing result is obtained.
[0140] The business data processing result is sent to the device of the data owner.
[0141] In the embodiments of this specification, the business data to be processed carried in the call request of the data owner can be consistent with the business data used when generating the adjusted parameter data of the second network layer, thereby enabling the target large model to generate business data processing results with better accuracy based on the business data to be processed. In practical applications, the types and meanings of the business data to be processed and the business data processing results can be set according to actual business needs. For example, the business data to be processed and the business data processing results can be respectively: user comment data to be sentiment-recognized, and the sentiment type to which the user comment data belongs; or respectively: user ID image to be verified for authenticity, and the authenticity verification result of the user ID image; or can be: article data to generate review information, and the review information generated for the article data, etc.
[0142] Based on and Figure 2 Following the same approach as the scheme shown, this specification also provides another method for training large models. Figure 3 This is a flowchart illustrating another large model training method provided in the embodiments of this specification. The execution entity of this process can be the data owner's device, or an application running on the data owner's device. Figure 3 As shown, the process may include:
[0143] Step 302: Obtain the simulation model of the pre-trained large model from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layers according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0144] In the embodiments of this specification, the simulation model of the pre-trained large model obtained in step 302 can be based on... Figure 2 The examples generated by these and their implementations will not be described in detail here.
[0145] Step 304: Train the simulation model using the business data from the data owner to obtain the adjusted parameter data of the second network layer.
[0146] In the embodiments of this specification, minimizing the loss value generated by the simulation model can be used as the training objective. Business data from the data owner can be used to train and optimize the parameters of the second network layer in the simulation model that can be adjusted, so as to obtain the adjusted parameter data of the second network layer.
[0147] Step 306: Send the adjusted parameter data of the second network layer to the device of the model provider; the device of the model provider is used to generate the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0148] In the embodiments of this specification, the adjusted parameter data of the second network layer sent to the device of the model provider in step 302 is... Figure 2 The adjusted parameter data of the second network layer received in step 208 will not be described in detail.
[0149] Figure 3 The proposed method replaces less important network layers in a pre-trained large model with pre-defined lightweight network layers to serve as the simulation model. This improves the consistency of structure and performance between the simulation model and the pre-trained large model while reducing the resources required for model training. This helps ensure that the target large model generated from the adjusted parameters of the second network layer and the pre-trained large model achieves good performance. Furthermore, the core data from both the data owner and the model provider remains within their respective domains, thus guaranteeing data privacy and security during cross-domain fine-tuning of the large model.
[0150] In the embodiments described in this specification, the simulation model may include an adapter generated based on the second network layer; the parameters of the adapter can be adjusted.
[0151] Correspondingly, step 304: training the simulation model using the business data from the data owner to obtain the adjusted parameter data of the second network layer may specifically include:
[0152] The simulation model is used to process the business data at the data owner's location to obtain the loss value corresponding to the simulation model; based on the loss value corresponding to the simulation model, the parameter data of the adapter is adjusted to obtain the adjusted parameter data for the adapter.
[0153] Specifically, adjusting the adapter's parameter data based on the loss value corresponding to the simulation model to obtain adjusted parameter data for the adapter may include:
[0154] If the adapter is generated using the second network layer that has been granted parameter adjustment authority, then the parameter data of the second network layer that has been granted parameter adjustment authority is adjusted according to the loss value corresponding to the simulation model to obtain the adjusted parameter data for the adapter; or,
[0155] If the adapter is generated based on a plugin for fine-tuning the pre-trained large model set for the second network layer, then the parameter data of the plugin is adjusted according to the loss value corresponding to the simulation model to obtain the adjusted parameter data for the adapter.
[0156] In the embodiments of this specification, due to Figure 2 The structure of the adapter and the training optimization principle have already been explained in the embodiments, so they will not be repeated here.
[0157] In this embodiment of the specification, after step 306: sending the adjusted parameter data of the second network layer to the device of the model provider, it may further include:
[0158] Send a call request for the target large model to the device of the model provider; the call request carries the business data to be processed at the data owner.
[0159] The system receives the business data processing results obtained by processing the business data to be processed using the target large model, based on feedback from the model provider's device.
[0160] In the embodiments of this specification, the business data to be processed and the business data for training the simulation model in step 304 can usually be consistent, so that the target large model can generate business data processing results with better accuracy for the business data to be processed. This will not be elaborated further.
[0161] Figure 4 The embodiments provided in this specification correspond to Figure 2 and Figure 3 A schematic diagram of the swimlane process for training large models in China. (See attached diagram.) Figure 4 As shown, the large model training process can involve execution entities such as the equipment of the model provider and the equipment of the data owner.
[0162] During the model training phase, the model provider's device can, on the one hand, select a third network layer from the preset network layers that needs to be replaced with a preset lightweight network layer based on the historical importance scores of the preset network layers in the pre-trained large model, thus obtaining the first action policy. The pre-trained large model, after implementing the first action policy, processes the first sample data to obtain the loss value corresponding to the first action policy. Based on the loss values corresponding to each first action policy, the reinforcement learning reward value for implementing each first action policy is determined. This allows the historical importance score of the third network layer to be updated based on the reinforcement learning reward value, and combined with the historical importance scores of other preset network layers, the target importance score of each preset network layer is obtained.
[0163] On the other hand, the model provider's device can also select a fourth network layer from the preset network layers that needs to be replaced with a preset lightweight network layer based on the target importance score of the preset network layers, thus obtaining a second action policy. The second sample data is then processed using the pre-trained large model after implementing the second action policy to obtain the loss value corresponding to the second action policy. Based on the loss value, the parameters of the lightweight network layers in the pre-trained large model after implementing the second action policy are adjusted to obtain a preset lightweight network layer that has a unique correspondence with the fourth network layer. Furthermore, the lightweight network layers whose parameters have not been adjusted and which correspond to other preset network layers are determined as preset lightweight network layers that have a unique correspondence with other preset network layers. This results in preset lightweight network layers that have a unique correspondence with each of the preset network layers.
[0164] Subsequently, the model provider's device can configure the first and second network layers in the preset network layers of the pre-trained large model according to the target importance score, obtain the simulation model of the pre-trained large model, and send the simulation model to the data owner's device; wherein, the configuration process is used to change the first network layer to a preset lightweight network layer with a unique correspondence, and to indicate that the parameters of the second network layer can be adjusted, and the importance of the second network layer is usually higher than that of the first network layer.
[0165] The data owner's equipment can use the data owner's business data to train the simulation model, obtain and feed back the adjusted parameter data of the second network layer to the model provider's equipment. This allows the model provider's equipment to generate the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0166] During the model invocation phase, the data owner's device can generate and send an invocation request for the target large model carrying the business data to be processed at the data owner's location to the model provider's device. After the model provider's device processes the business data to be processed using the target large model to obtain the business data processing result, it can send the business data processing result to the data owner's device, thereby satisfying the data owner's business processing needs.
[0167] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods. Figure 5 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a large model training device. (See attached diagram.) Figure 5 As shown, the device may include:
[0168] The first acquisition module 502 is used to acquire the importance information of each preset network layer in the pre-trained large model.
[0169] Setting module 504 is used to perform setting processing on the first network layer and the second network layer in the preset network layer according to the importance information to obtain the simulation model of the pre-trained large model; wherein, the setting processing is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0170] The first sending module 506 is used to send the simulation model to the device of the data owner.
[0171] The receiving module 508 is used to receive the adjusted parameter data of the second network layer fed back by the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using the service data at the data owner.
[0172] The generation module 510 is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0173] based on Figure 5 The embodiments of this specification also provide some specific implementations of the device, which will be described below.
[0174] Optionally, the first acquisition module 502 may include:
[0175] A determining unit is used to determine, using a reinforcement learning algorithm, the reinforcement learning reward value of implementing a first action policy on the pre-trained large model; wherein the first action policy is used to instruct the replacement of the third network layer in the preset network layers in the pre-trained large model with a preset lightweight network layer, and the reinforcement learning reward value is positively correlated with the loss value generated by the pre-trained large model processing the first sample data after implementing the first action policy.
[0176] The update unit is used to update the historical importance score of the third network layer according to the reinforcement learning reward value, so as to obtain the target importance score of each preset network layer in the pre-trained large model; wherein the reinforcement learning reward value is positively correlated with the target importance score of the third network layer.
[0177] Optionally, the number of the first action strategies can be greater than 1. Correspondingly, the determining unit can specifically be used to: based on the historical importance scores of each preset network layer in the pre-trained large model, select a first number of third network layers that need to be replaced by preset lightweight network layers from the preset network layers to obtain a first action strategy; wherein, the probability of the preset network layer being selected as the third network layer is negatively correlated with the historical importance score of the preset network layer.
[0178] The first sample data is processed using the pre-trained large model after implementing the first action strategy to obtain the loss value corresponding to the first action strategy.
[0179] Based on the loss value corresponding to each of the first action strategies, the reinforcement learning reward value for implementing each of the first action strategies on the pre-trained large model is calculated according to the first formula.
[0180] The first formula can be:
[0181]
[0182] r i The loss is the reinforcement learning reward value of implementing the first action policy for the i-th action in the pre-trained large model. i Let loss be the loss value corresponding to the i-th first action policy. j Let be the loss value corresponding to the j-th first action strategy, and T be the number of first action strategies.
[0183] The update unit can be specifically used for:
[0184] For any of the third network layers, the target importance score of the third network layer is calculated according to the second formula based on the reinforcement learning reward value and the historical importance score of the third network layer.
[0185] The second formula can be:
[0186] s cd =s hd +∑r q ,
[0187] s cd s is the target importance score of the third network layer. hd r is the historical importance score of the third network layer. q This is used to indicate the reinforcement learning reward value corresponding to the first action policy that replaces the third network layer.
[0188] Optionally, the setting module may include:
[0189] The first filtering unit is used to filter out the third number of preset network layers with the highest importance in the pre-trained large model based on the importance information, and obtain the second network layer.
[0190] The first generation unit is configured to generate an adapter for the pre-trained large model based on the second network layer; wherein the parameters of the adapter can be adjusted.
[0191] The second filtering unit is used to filter out the fourth number of preset network layers with the lowest importance in the pre-trained large model based on the importance information, and obtain the first network layer.
[0192] The replacement unit is used to replace the first network layer using a preset lightweight network layer to obtain a simulation model carrying the adapter.
[0193] Optionally, the receiving module may include:
[0194] A receiving unit is used to receive adjusted parameter data for the adapter fed back by the device of the data owner.
[0195] Optionally, the first generation unit may be used to: generate the adapter according to the second network layer which has been granted parameter adjustment permissions; or, set up a plugin for fine-tuning the pre-trained large model for the second network layer to obtain the adapter.
[0196] The receiving unit can be specifically used to: receive adjusted parameter data for the second network layer that has been granted parameter adjustment authority, fed back by the device of the data owner; or, receive adjusted parameter data for the plug-in fed back by the device of the data owner.
[0197] Optionally, the generation module may include:
[0198] The second generation unit is used to replace the historical parameter data of the second network layer in the pre-trained large model with the adjusted parameter data of the second network layer for which parameter adjustment authority has been granted, to obtain the target large model; or...
[0199] The third generation unit is used to combine the adjusted parameter data for the plugin with the historical parameter data of the pre-trained large model to generate the target large model.
[0200] Optional, Figure 5 The device may also include:
[0201] The filtering module is used to filter out a second number of fourth network layers that need to be replaced with lightweight network layers from the preset network layers in the pre-trained large model based on the target importance scores of each preset network layer, thereby obtaining a second action strategy; wherein the probability of the preset network layer being filtered into the fourth network layer is negatively correlated with the target importance scores of the preset network layer.
[0202] The replacement module is used to replace the fourth network layer in the pre-trained large model with a lightweight network layer in the lightweight network layer set that has a unique correspondence with the fourth network layer, based on the second action strategy, to obtain the pre-trained large model after implementing the second action strategy; wherein, each lightweight network layer in the lightweight network layer set corresponds one-to-one with each preset network layer.
[0203] The sample data processing module is used to process the second sample data using the pre-trained large model after implementing the second action strategy, and obtain the loss value corresponding to the second action strategy.
[0204] The parameter adjustment module is used to adjust the parameters of the lightweight network layer in the pre-trained large model after implementing the second action strategy based on the loss value corresponding to the second action strategy, so as to obtain a preset lightweight network layer that has a unique correspondence with the fourth network layer.
[0205] The determining module is used to determine the lightweight network layer in the set of lightweight network layers that corresponds to other preset network layers as a preset lightweight network layer that has a unique correspondence with the other preset network layers.
[0206] Specifically, the setting process is used to change the first network layer to a preset lightweight network layer that has a unique correspondence with the first network layer.
[0207] The first action strategy is specifically used to instruct the third network layer to be changed to a preset lightweight network layer that has a unique correspondence with the third network layer.
[0208] Optional, Figure 5 The device may further include:
[0209] The second acquisition module is used to acquire the call request from the data owner for the target large model; the call request carries the business data to be processed at the data owner.
[0210] The business data processing module is used to process the business data to be processed using the target large model to obtain the business data processing result.
[0211] The second sending module is used to send the business data processing result to the device of the data owner.
[0212] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods. Figure 6 The embodiments provided in this specification correspond to Figure 3 A schematic diagram of the structure of a large model training device. (See attached diagram.) Figure 6 As shown, the device may include:
[0213] The acquisition module 602 is used to acquire a simulation model of a pre-trained large model from the device of the model provider; wherein the simulation model is obtained by setting the first network layer and the second network layer in the preset network layers according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0214] The training module 604 is used to train the simulation model using business data from the data owner to obtain the adjusted parameter data of the second network layer.
[0215] The first sending module 606 is used to send the adjusted parameter data of the second network layer to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0216] based on Figure 6The embodiments of this specification also provide some specific implementations of the device, which will be described below.
[0217] Optionally, the simulation model may include an adapter generated based on the second network layer; the parameters of the adapter can be adjusted. Correspondingly, the training module may include:
[0218] The business data processing unit is used to process the business data from all data owners using the simulation model to obtain the loss value corresponding to the simulation model.
[0219] The parameter adjustment unit is used to adjust the parameter data of the adapter according to the loss value corresponding to the simulation model, so as to obtain the adjusted parameter data for the adapter.
[0220] Specifically, the parameter adjustment unit may include:
[0221] The first adjustment subunit is configured to, if the adapter is generated using the second network layer with granted parameter adjustment authority, adjust the parameter data of the second network layer with granted parameter adjustment authority according to the loss value corresponding to the simulation model, to obtain adjusted parameter data for the adapter; or,
[0222] The second adjustment subunit is used to adjust the parameter data of the plugin according to the loss value corresponding to the simulation model if the adapter is generated based on the plugin for fine-tuning the pre-trained large model set for the second network layer, so as to obtain the adjusted parameter data for the adapter.
[0223] Optional, Figure 6 The device may further include:
[0224] The second sending module is used to send a call request for the target large model to the device of the model provider; the call request carries the business data to be processed at the data owner.
[0225] The receiving module is used to receive the business data processing result obtained by processing the business data to be processed using the target large model, which is fed back by the device of the model provider.
[0226] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0227] Figure 7 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a large model training device. (See diagram below.) Figure 7As shown, device 700 may include: at least one processor 710; and a memory 730 communicatively connected to the at least one processor; wherein the memory 730 stores instructions 720 executable by the at least one processor 710, the instructions being executed by the at least one processor 710 to enable the at least one processor 710 to:
[0228] Obtain information on the importance of each pre-set network layer in a pre-trained large model.
[0229] Based on the importance information, settings are applied to the first and second network layers in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the settings are used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0230] The simulation model is sent to the device of the data owner.
[0231] The device receives the adjusted parameter data of the second network layer from the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using service data from the data owner.
[0232] Based on the adjusted parameter data of the second network layer and the pre-trained large model, a target large model is generated.
[0233] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.
[0234] Figure 8 The embodiments provided in this specification correspond to Figure 3 A schematic diagram of the structure of a large model training device. (See diagram below.) Figure 8 As shown, device 800 may include: at least one processor 810; and a memory 830 communicatively connected to the at least one processor; wherein the memory 830 stores instructions 820 executable by the at least one processor 810, the instructions being executed by the at least one processor 810 to enable the at least one processor 810 to:
[0235] The simulation model of the pre-trained large model is obtained from the device of the model provider; wherein the simulation model is obtained by setting the first network layer and the second network layer in the preset network layer according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer.
[0236] The simulation model is trained using business data from the data owner to obtain the adjusted parameter data for the second network layer.
[0237] The adjusted parameter data of the second network layer is sent to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
[0238] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on its differences from other embodiments. In particular, for... Figure 7 and Figure 8 As the device shown is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.
[0239] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0240] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0241] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0242] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0243] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0244] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0245] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0246] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0247] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0248] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0249] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0250] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0251] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0252] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0253] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A large model training method, applied to model providers, including: Obtain information on the importance of each pre-defined network layer in a pre-trained large model; Based on the importance information, settings are applied to the first and second network layers in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the settings are used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer; Send the simulation model to the device of the data owner; The system receives adjusted parameter data for the second network layer from the device of the data owner; the adjusted parameter data for the second network layer is obtained by training the simulation model using service data from the data owner. Based on the adjusted parameter data of the second network layer and the pre-trained large model, a target large model is generated.
2. The method as described in claim 1, wherein obtaining the importance information of each preset network layer in the pre-trained large model specifically includes: Using a reinforcement learning algorithm, a reinforcement learning reward value is determined for implementing a first action policy on the pre-trained large model; wherein, the first action policy is used to instruct the replacement of the third network layer in the preset network layers in the pre-trained large model with a preset lightweight network layer, and the reinforcement learning reward value is positively correlated with the loss value generated by the pre-trained large model processing the first sample data after implementing the first action policy. Based on the reinforcement learning reward value, the historical importance score of the third network layer is updated to obtain the target importance score of each preset network layer in the pre-trained large model; wherein, the reinforcement learning reward value is positively correlated with the target importance score of the third network layer.
3. The method as described in claim 2, wherein the number of the first action strategies is greater than 1; The step of using a reinforcement learning algorithm to determine the reinforcement learning reward value for implementing the first action policy on the pre-trained large model specifically includes: Based on the historical importance scores of each preset network layer in the pre-trained large model, a first number of third network layers that need to be replaced by preset lightweight network layers are selected from the preset network layers to obtain a first action policy; wherein, the probability of the preset network layer being selected as the third network layer is negatively correlated with the historical importance score of the preset network layer; The first sample data is processed using the pre-trained large model after implementing the first action strategy to obtain the loss value corresponding to the first action strategy. Based on the loss value corresponding to each of the first action strategies, the reinforcement learning reward value for implementing each of the first action strategies on the pre-trained large model is calculated according to the first formula. The first formula is: , a reinforcement learning reward value for implementing the ith first action policy for the pre-trained large model, a loss value corresponding to the ith first action policy, a loss value corresponding to the jth first action policy, and T is the number of the first action policies.
4. The method as described in claim 3, wherein updating the historical importance score of the third network layer based on the reinforcement learning reward value to obtain the target importance score of each preset network layer in the pre-trained large model specifically includes: For any of the third network layers, the target importance score of the third network layer is calculated according to the second formula based on the reinforcement learning reward value and the historical importance score of the third network layer. The second formula is: , The target importance score of the third network layer. The historical importance score of the third network layer. This is used to indicate the reinforcement learning reward value corresponding to the first action policy that replaces the third network layer.
5. The method according to any one of claims 2-4, further comprising: Based on the target importance scores of each preset network layer in the pre-trained large model, a second number of fourth network layers that need to be replaced by lightweight network layers are selected from the preset network layers to obtain a second action policy; wherein, the probability of the preset network layer being selected as the fourth network layer is negatively correlated with the target importance score of the preset network layer. Based on the second action strategy, the fourth network layer in the pre-trained large model is replaced by a lightweight network layer in the lightweight network layer set that has a unique correspondence with the fourth network layer, thereby obtaining the pre-trained large model after implementing the second action strategy; wherein, each lightweight network layer in the lightweight network layer set corresponds one-to-one with each preset network layer. The pre-trained large model after implementing the second action strategy is used to process the second sample data to obtain the loss value corresponding to the second action strategy; Based on the loss value corresponding to the second action strategy, the parameters of the lightweight network layer in the pre-trained large model after implementing the second action strategy are adjusted to obtain a preset lightweight network layer that has a unique correspondence with the fourth network layer. The lightweight network layer in the set of lightweight network layers that corresponds to other preset network layers is determined as a preset lightweight network layer that has a unique correspondence with the other preset network layers. Specifically, the setting process is used to change the first network layer to a preset lightweight network layer that has a unique correspondence with the first network layer. The first action strategy is specifically used to instruct the third network layer to be changed to a preset lightweight network layer that has a unique correspondence with the third network layer.
6. The method as described in claim 1, wherein the step of setting the first and second network layers in the preset network layers according to the importance information to obtain the simulation model of the pre-trained large model specifically includes: Based on the importance information, the third number of the preset network layers with the highest importance in the pre-trained large model are selected to obtain the second network layer; Based on the second network layer, an adapter for the pre-trained large model is generated; wherein the parameters of the adapter can be adjusted. Based on the importance information, the fourth number of preset network layers with the lowest importance in the pre-trained large model are selected to obtain the first network layer; The first network layer is replaced by a preset lightweight network layer to obtain a simulation model carrying the adapter; The adjusted parameter data of the second network layer fed back by the device receiving the data specifically includes: The device receives adjusted parameter data for the adapter from the data owner's device.
7. The method of claim 6, wherein generating the adapter for the pre-trained large model based on the second network layer specifically includes: The adapter is generated based on the second network layer, which has been granted parameter adjustment permissions. or, A plugin for fine-tuning the pre-trained large model is set for the second network layer to obtain the adapter; The adjusted parameter data for the adapter fed back by the device receiving the data specifically includes: The device receiving the adjusted parameter data of the second network layer, for which the parameter adjustment authority has been granted, is fed back by the device of the data owner. or, The device receives the adjusted parameter data for the plugin from the data owner's device.
8. The method of claim 7, wherein generating the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model specifically includes: By using the adjusted parameter data of the second network layer (which has been granted parameter adjustment authority), the historical parameter data of the second network layer in the pre-trained large model is replaced to obtain the target large model; or... By combining the adjusted parameter data for the plugin with the historical parameter data of the pre-trained large model, a target large model is generated.
9. The method of claim 1, further comprising, after generating the target large model based on the adjusted parameter data of the second network layer and the pre-trained large model: Obtain the data owner's invocation request for the target large model; The call request carries the business data to be processed from the data owner. The target large model is used to process the business data to be processed, and the business data processing result is obtained; The business data processing result is sent to the device of the data owner.
10. A large model training method, applied to data owners, including: The simulation model of the pre-trained large model is obtained from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layer according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer. The simulation model is trained using business data from the data owner to obtain the adjusted parameter data for the second network layer; The adjusted parameter data of the second network layer is sent to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
11. The method of claim 10, wherein the simulation model carries an adapter generated based on the second network layer; the parameters of the adapter can be adjusted; The step of training the simulation model using business data from the data owner to obtain adjusted parameter data for the second network layer specifically includes: The simulation model is used to process the business data at the data owner's location to obtain the loss value corresponding to the simulation model; Based on the loss value corresponding to the simulation model, the parameter data of the adapter is adjusted to obtain the adjusted parameter data for the adapter.
12. The method as described in claim 11, wherein adjusting the parameter data of the adapter based on the loss value corresponding to the simulation model to obtain adjusted parameter data for the adapter specifically includes: If the adapter is generated using the second network layer that has been granted parameter adjustment authority, then the parameter data of the second network layer that has been granted parameter adjustment authority is adjusted according to the loss value corresponding to the simulation model to obtain the adjusted parameter data for the adapter. or, If the adapter is generated based on a plugin for fine-tuning the pre-trained large model set for the second network layer, then the parameter data of the plugin is adjusted according to the loss value corresponding to the simulation model to obtain the adjusted parameter data for the adapter.
13. The method of claim 10, further comprising, after sending the adjusted parameter data of the second network layer to the device of the model provider: Send a request for the target large model to the device of the model provider; The call request carries the business data to be processed from the data owner. The system receives the business data processing results obtained by processing the business data to be processed using the target large model, based on feedback from the device provided by the model provider.
14. A large model training device, applied to a model provider, comprising: The first acquisition module is used to acquire information on the importance of each preset network layer in the pre-trained large model. The setting module is used to perform setting processing on the first network layer and the second network layer in the preset network layers according to the importance information, so as to obtain the simulation model of the pre-trained large model; wherein, the setting processing is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer; The first sending module is used to send the simulation model to the device of the data owner; The receiving module is used to receive the adjusted parameter data of the second network layer fed back by the device of the data owner; the adjusted parameter data of the second network layer is obtained by training the simulation model using the service data at the data owner; The generation module is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
15. The apparatus of claim 14, wherein the first acquisition module comprises: A determining unit is used to determine, using a reinforcement learning algorithm, the reinforcement learning reward value of implementing a first action policy on the pre-trained large model; wherein, the first action policy is used to instruct the replacement of the third network layer in the preset network layers in the pre-trained large model with a preset lightweight network layer, and the reinforcement learning reward value is positively correlated with the loss value generated by the pre-trained large model processing the first sample data after implementing the first action policy. The update unit is used to update the historical importance score of the third network layer according to the reinforcement learning reward value, so as to obtain the target importance score of each preset network layer in the pre-trained large model; wherein the reinforcement learning reward value is positively correlated with the target importance score of the third network layer.
16. The apparatus of claim 15, wherein the number of the first action strategies is greater than 1; The determining unit is specifically used for: Based on the historical importance scores of each preset network layer within the pre-trained large model, a first number of third network layers that need to be replaced with preset lightweight network layers are selected from the preset network layers to obtain a first action policy; wherein... The probability that the preset network layer is selected as the third network layer is negatively correlated with the historical importance score of the preset network layer; The first sample data is processed using the pre-trained large model after implementing the first action strategy to obtain the loss value corresponding to the first action strategy. Based on the loss value corresponding to each of the first action strategies, the reinforcement learning reward value for implementing each of the first action strategies on the pre-trained large model is calculated according to the first formula. The first formula is: , The reinforcement learning reward value for implementing the first action policy for the i-th time in the pre-trained large model. Let i be the loss value corresponding to the first action strategy. Let be the loss value corresponding to the j-th first action strategy, and T be the number of first action strategies; The update unit is specifically used for: For any of the third network layers, the target importance score of the third network layer is calculated according to the second formula based on the reinforcement learning reward value and the historical importance score of the third network layer. The second formula is: , The target importance score of the third network layer. The historical importance score of the third network layer. This is used to indicate the reinforcement learning reward value corresponding to the first action policy that replaces the third network layer.
17. The apparatus of any one of claims 15-16, further comprising: The filtering module is used to filter out a second number of fourth network layers that need to be replaced with lightweight network layers from the preset network layers in the pre-trained large model based on the target importance score of each preset network layer, and obtain a second action strategy; wherein the probability of the preset network layer being filtered into the fourth network layer is negatively correlated with the target importance score of the preset network layer. The replacement module is used to replace the fourth network layer in the pre-trained large model with a lightweight network layer in the lightweight network layer set that has a unique correspondence with the fourth network layer, based on the second action strategy, to obtain the pre-trained large model after implementing the second action strategy; wherein, each lightweight network layer in the lightweight network layer set corresponds one-to-one with each preset network layer. The sample data processing module is used to process the second sample data using the pre-trained large model after implementing the second action strategy, and obtain the loss value corresponding to the second action strategy. The parameter adjustment module is used to adjust the parameters of the lightweight network layer in the pre-trained large model after implementing the second action strategy based on the loss value corresponding to the second action strategy, so as to obtain a preset lightweight network layer that has a unique correspondence with the fourth network layer. The determining module is used to determine the lightweight network layer in the set of lightweight network layers that corresponds to other preset network layers as a preset lightweight network layer that has a unique correspondence with the other preset network layers. Specifically, the setting process is used to change the first network layer to a preset lightweight network layer that has a unique correspondence with the first network layer. The first action strategy is specifically used to instruct the third network layer to be changed to a preset lightweight network layer that has a unique correspondence with the third network layer.
18. The apparatus of claim 14, wherein the setting module comprises: The first filtering unit is used to filter out the third number of preset network layers with the highest importance in the pre-trained large model based on the importance information, and obtain the second network layer. The first generation unit is configured to generate an adapter for the pre-trained large model based on the second network layer; wherein the parameters of the adapter can be adjusted. The second filtering unit is used to filter out the fourth number of preset network layers with the lowest importance in the pre-trained large model according to the importance information, so as to obtain the first network layer. The replacement unit is used to replace the first network layer with a preset lightweight network layer to obtain a simulation model carrying the adapter. The receiving module includes: A receiving unit is used to receive adjusted parameter data for the adapter fed back by the device of the data owner.
19. The apparatus of claim 18, wherein the first generating unit is specifically configured to: The adapter is generated based on the second network layer, which has been granted parameter adjustment permissions; or... A plugin for fine-tuning the pre-trained large model is set for the second network layer to obtain the adapter; The receiving unit is specifically used for: Receive the adjusted parameter data of the second network layer, which has been granted parameter adjustment authority, from the device of the data owner; or... Receive the adjusted parameter data for the plugin from the device of the data owner; The generation module includes: The second generation unit is used to replace the historical parameter data of the second network layer in the pre-trained large model with the adjusted parameter data of the second network layer that has been granted parameter adjustment authority, so as to obtain the target large model. or, The third generation unit is used to combine the adjusted parameter data for the plugin with the historical parameter data of the pre-trained large model to generate the target large model.
20. The apparatus of claim 14, further comprising: The second acquisition module is used to acquire the data owner's call request for the target large model; The call request carries the business data to be processed from the data owner. The business data processing module is used to process the business data to be processed using the target large model to obtain the business data processing result; The second sending module is used to send the business data processing result to the device of the data owner.
21. A large model training apparatus, applied to data owners, comprising: The acquisition module is used to acquire a simulation model of a pre-trained large model from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layers according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer; The training module is used to train the simulation model using business data from the data owner to obtain the adjusted parameter data of the second network layer. The first sending module is used to send the adjusted parameter data of the second network layer to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.
22. The apparatus of claim 21, wherein the simulation model carries an adapter generated based on the second network layer; The parameters of the adapter can be adjusted; The training module includes: A business data processing unit is used to process the business data at the data owner using the simulation model to obtain the loss value corresponding to the simulation model. The parameter adjustment unit is used to adjust the parameter data of the adapter according to the loss value corresponding to the simulation model, so as to obtain the adjusted parameter data for the adapter. The parameter adjustment unit specifically includes: The first adjustment subunit is configured to, if the adapter is generated using the second network layer with granted parameter adjustment authority, adjust the parameter data of the second network layer with granted parameter adjustment authority according to the loss value corresponding to the simulation model, to obtain adjusted parameter data for the adapter; or, The second adjustment subunit is used to adjust the parameter data of the plugin based on the loss value corresponding to the simulation model if the adapter is generated based on the plugin for fine-tuning the pre-trained large model set for the second network layer, so as to obtain the adjusted parameter data for the adapter.
23. The apparatus of claim 21, further comprising: The second sending module is used to send a call request for the target large model to the device of the model provider; The call request carries the business data to be processed from the data owner. The receiving module is used to receive the business data processing result obtained by processing the business data to be processed using the target large model, which is fed back by the device of the model provider.
24. A large model training device, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: Obtain information on the importance of each pre-defined network layer in a pre-trained large model; Based on the importance information, settings are applied to the first and second network layers in the preset network layers to obtain the simulation model of the pre-trained large model; wherein, the settings are used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer; Send the simulation model to the device of the data owner; The system receives adjusted parameter data for the second network layer from the device of the data owner; the adjusted parameter data for the second network layer is obtained by training the simulation model using service data from the data owner. Based on the adjusted parameter data of the second network layer and the pre-trained large model, a target large model is generated.
25. A large model training device, comprising: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The simulation model of the pre-trained large model is obtained from the device of the model provider; wherein, the simulation model is obtained by setting the first network layer and the second network layer in the preset network layer according to the importance information of each preset network layer in the pre-trained large model; the setting process is used to change the first network layer to a preset lightweight network layer, and to indicate that the parameters of the second network layer can be adjusted; the importance of the second network layer is higher than that of the first network layer. The simulation model is trained using business data from the data owner to obtain the adjusted parameter data for the second network layer; The adjusted parameter data of the second network layer is sent to the device of the model provider; the device of the model provider is used to generate a target large model based on the adjusted parameter data of the second network layer and the pre-trained large model.