Model obtaining method and device
Through the denoising processing and parameter update method based on the diffusion model, the problem of large workload and poor performance in the existing technology is solved, and efficient and reliable model parameter adjustment and performance improvement are achieved.
Patent Information
- Application Number
- CN202510376751.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-13
AI Technical Summary
When adjusting model parameters in the prior art, the workload is large and the adjusted model performance is poor, making it difficult to respond to dynamic task requirements instantly.
By obtaining the first conditional information required by the initial model to perform the target task, denoising the initial parameters based on the diffusion model, obtaining the target parameters, and updating the initial model with the target parameters to form the target model to perform the target task.
It realizes efficient adjustment of model parameters, improves model performance, can respond to new task requirements in real time, and enhances the reliability and adaptability of the model.
Smart Images

Figure CN120146112A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular, to a method and apparatus for obtaining a model. Background Art
[0002] In practical applications, for different tasks, it is necessary to adjust the parameters of the model to varying degrees. However, currently, the workload of adjusting the parameters of the model is large and the performance of the adjusted model is not good. Summary of the Invention
[0003] In view of this, the present disclosure provides a method and apparatus for obtaining a model.
[0004] According to a first aspect of the present disclosure, there is provided a method for obtaining a model, including: obtaining first condition information required for an initial model to execute a target task, where the first condition information represents a result requirement expected for the initial model to execute the target task; using the first condition information as a constraint, performing denoising processing on first initial parameters based on a diffusion model to obtain target parameters, where the first initial parameters represent parameter noise of the initial model regarding the target task; and updating the initial model with the target parameters to obtain a target model for executing the target task through the target model.
[0005] A second aspect of the present disclosure provides a model obtaining apparatus, including: an obtaining module configured to obtain first condition information required for an initial model to execute a target task, where the first condition information represents a result requirement expected for the initial model to execute the target task; a processing module configured to perform denoising processing on first initial parameters based on a diffusion model with the first condition information as a constraint to obtain target parameters, where the first initial parameters represent parameter noise of the initial model regarding the target task; and a loading module configured to update the initial model with the target parameters to obtain a target model for executing the target task by using the target model.
[0006] A third aspect of the present disclosure provides an electronic device, including: one or more processors; a memory configured to store one or more programs, where when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned model obtaining method.
[0007] A fourth aspect of the present disclosure further provides a computer-readable storage medium having executable instructions stored thereon, where when the instructions are executed by a processor, the processor is caused to execute the above-mentioned model obtaining method.
[0008] A fifth aspect of the present disclosure further provides a computer program product including a computer program, where when the computer program is executed by a processor, the above-mentioned model obtaining method is implemented.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily apparent from the following description. Description of the Drawings
[0010] Through the following description of the embodiments of the present disclosure with reference to the drawings, the above and other objects, features, and advantages of the present disclosure will become more apparent. In the drawings:
[0011] Figure 1 A flowchart of a method for obtaining a model according to an embodiment of the present disclosure is schematically shown;
[0012] Figure 2 A flowchart of a method for training a diffusion model according to an embodiment of the present disclosure is schematically shown;
[0013] Figure 3 A flowchart of a method for obtaining first sample data according to an embodiment of the present disclosure is schematically shown;
[0014] Figure 4 A schematic diagram of a method for obtaining a model according to an embodiment of the present disclosure is schematically shown;
[0015] Figure 5 A schematic view of a method for obtaining a model according to an embodiment of the present disclosure is schematically shown;
[0016] Figure 6 One of the effect diagrams of a method for obtaining a model according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 7 Another effect diagram of a method for obtaining a model according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 8 A structural block diagram of a device for obtaining a model according to an embodiment of the present disclosure is schematically shown; and
[0019] Figure 9 A block diagram of an electronic device suitable for implementing the method for obtaining a model according to an embodiment of the present disclosure is schematically shown. Detailed Description of the Embodiments
[0020] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments may be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0021] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising" and the like used herein indicate the presence of features, steps, operations and / or components, but do not preclude the presence or addition of one or more other features, steps, operations or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.
[0023] In the case of using expressions such as "at least one of A, B, and C, etc.", generally it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0024] Embodiments of the present disclosure provide a method and apparatus for obtaining a model. Before introducing the technical solutions provided by the embodiments of the present disclosure, the related technologies involved in the present disclosure will be described first.
[0025] In practical applications, for different tasks, it is necessary to adjust the parameters of the model to varying degrees. However, at present, the workload of adjusting the parameters of the model is large and the performance of the adjusted model is not good.
[0026] In one example, in a commercial recommendation system, the task requirements of the platform or users often change dynamically (for example, the preference for the accuracy or diversity of the recommendation results changes). In order to be able to respond to dynamic task requirements immediately, a combined overall optimization goal can be obtained by combining multiple objectives (such as accuracy, diversity, etc.), and then the model is retrained to adapt to this new optimization goal. However, retraining the model is costly, time-consuming, and difficult to respond to new task requirements in a timely manner. At the same time, model fusion can also be used to combine the output results of multiple independent models to adapt to this new optimization goal. However, simple model parameter fusion is difficult to have fine-grained control over the performance of the fused model, and the performance of the fused model is often unstable.
[0027] Before further elaborating on the embodiments of the present disclosure, the nouns and terms involved in the embodiments of the present disclosure will be described. The nouns and terms involved in the embodiments of the present disclosure are applicable to the following explanations.
[0028] Diffusion Models are a type of generative model that do not require an additional classifier. Instead, by jointly training conditional and unconditional diffusion models, conditional information is directly introduced into the model.
[0029] Classifier-Free Diffusion Models. In conditional diffusion models, additional conditional information (such as class labels, text descriptions, performance parameters, etc.) is introduced during the generation process to generate data (or parameters) that meet specific conditions.
[0030] Embodiments of the present disclosure provide a method for obtaining a model, including: obtaining first conditional information required for an initial model to perform a target task, where the first conditional information represents the result requirements expected for the initial model to perform the target task; using the first conditional information as a constraint, based on a diffusion model, denoising a first initial parameter to obtain a target parameter, where the first initial parameter represents the parameter noise of the initial model regarding the target task; and updating the initial model using the target parameter to obtain a target model for performing the target task through the target model.
[0031] The following will be through Figures 1 to 7 A detailed description of the method for obtaining a model according to embodiments of the present disclosure will be given.
[0032] Figure 1 Schematically shows a flowchart of the method for obtaining a model according to an embodiment of the present disclosure.
[0033] As Figure 1 shown, the method for obtaining a model in this embodiment includes operations S210 to S230.
[0034] In operation S210, obtain first conditional information required for an initial model to perform a target task, where the first conditional information represents the result requirements expected for the initial model to perform the target task.
[0035] Exemplarily, the initial model may be a model to perform the target task, but the result of the initial model's current model parameters performing the target task may not meet the first conditional information.
[0036] The target task refers to the specific task or goal that the model needs to complete in a specific application scenario. For example, in the text-to-image scenario, the target task may be to generate an image that meets the requirements according to the text. In the field of natural language processing, the target task may be to generate a piece of text that meets the expectations according to the input prompt. For example: automatic writing. In e-commerce or social platforms, the target task may be to recommend relevant products or content based on the user's historical behavior.
[0037] The first conditional information is usually in text form and is used to constrain the denoising process of the diffusion model. In this embodiment, the first conditional information can be used to describe the effects that the model is expected to achieve when performing the target task. The first conditional information can include one or more execution effect requirements. The first conditional information can be an increase in the type of execution effect requirement, or it can be a change in the weight of the execution effect requirement. The first conditional information may vary depending on the target task. The first conditional information may also vary depending on the weights of the execution effect requirements. For example, in a recommendation task, the first conditional information can be that the weights of the accuracy and diversity of the recommendation results satisfy a first preset ratio. For example, accuracy:diversity is 8:2. In another recommendation task, the first conditional information can be that the weights of the accuracy and diversity of the recommendation results satisfy a second preset ratio. For example, accuracy:diversity is 5:5. In the task of generating images from text, the first conditional information can be the accuracy, diversity, and image quality of the generated images.
[0038] In operation S220, using the first conditional information as a constraint, the first initial parameter is denoised based on the diffusion model to obtain a target parameter, where the first initial parameter represents the parameter noise of the initial model regarding the target task.
[0039] Exemplarily, the first initial parameter can be the noise distribution of the original parameters (such as weights, biases, etc.) of the initial model when performing the target task. The first initial parameter can be obtained through the following operations: obtaining the original parameters of the initial model and adding noise to the original parameters according to multiple time steps to obtain the first initial parameter.
[0040] If the initial model uses the original parameters to perform the target task, the execution result may not fully meet the first conditional information. The original parameters can be randomly initialized, or they can conform to a Gaussian distribution, or they can conform to a normal distribution. Adding noise to the original parameters step by step according to multiple time steps to obtain the first initial parameter, and the first initial parameter is a pure noise distribution.
[0041] Here, the original parameters can be some of the parameters required for the initial model to perform the target task. That is, based on the existing model, only a small number of parameters are adjusted to adapt to the new task, rather than retraining all the parameters of the model. The essence of "updating some parameters" is to achieve task adaptation by updating some parameters while retaining the general capabilities of the model. The original parameters can also be all the parameters of the initial model when performing the target task.
[0042] The diffusion model is used to take the first conditional information as a condition to denoise the first initial parameter. During the denoising process, the first initial parameter is optimized and adjusted to generate a target parameter that meets the requirements of the initial model when performing the target task. Applying the target parameter by the initial model to perform the target task can meet the result requirements corresponding to the first conditional information.
[0043] Different execution effect requirement weights correspond to different conditional information of the target task, and different conditional information corresponds to different target parameters. For the same target task, when the execution effect requirement weight changes, the target parameters also change accordingly.
[0044] In operation S230, the initial model is updated using the target parameters to obtain a target model, so as to execute the target task through the target model.
[0045] Exemplarily, the target model is obtained by applying the target parameters to the initial model. The parameters of the initial model are adjusted using the target parameters to form a target model adapted to the first conditional information for executing the target task, so as to respond to the first conditional information of the target task.
[0046] It can be understood that by denoising the parameters of the initial model through a diffusion model, the target parameters required to meet the result requirements for executing the target task can be obtained; and it can instantaneously respond to the new task requirements of the model.
[0047] At the same time, the diffusion model is guided by the first conditional information to optimize the parameters of the initial model, making the parameters of the optimized initial model more accurate and stable, and the optimized initial model more in line with the task requirements, thereby improving the reliability of model application.
[0048] As described above, in operation S220, using the first conditional information as a constraint, the first initial parameters are denoised based on the diffusion model to obtain the target parameters. In one implementable manner, this operation may further include operation S221.
[0049] In operation S221, using the first conditional information as a constraint, the noise of the first initial parameters is gradually removed step by step according to multiple time steps to obtain the target parameters; the target parameters represent the model parameters required by the initial model when the execution result meets the first requirement information.
[0050] Exemplarily, during the denoising process, the diffusion model takes the first conditional information as a conditional input to guide the diffusion model to recover the target parameters that meet specific requirements (the execution result meets the first requirement information) from the noise. That is, the execution result of the initial model applying the target parameters to execute the target task can meet the first conditional information.
[0051] In one example, for a recommendation model, in a new recommendation task, the first conditional information is that the weights of the accuracy and diversity of the recommendation results satisfy a first preset ratio of 8:2. In the process of obtaining the model parameters of the recommendation model for the new recommendation task, the original parameters of the recommendation model are added with noise for T time steps to obtain the first initial parameters. The first initial parameters are input into the diffusion model, and given the first conditional information, the diffusion model starts from the first initial parameters and denoises for T time steps to generate target parameters that satisfy the first conditional information.
[0052] As described above, during the denoising process, the noise is predicted based on the current state at the current time step, and the current state is updated using the predicted noise to obtain the state of the next time step. The state is the intermediate representation of the diffusion model at the time step, and the state includes parameter information and noise information.
[0053] The state can be the intermediate representation at a certain time step during the diffusion process, that is, a part of the parameters are optimized and the other part of the parameters are noise. At time step t, the state corresponding to time step t is neither pure data nor pure noise, but a mixture of both. As time step t decreases, the noise in the state gradually decreases and the data component gradually increases.
[0054] In one example, for each time step t = T, T - 1, …, 0, the diffusion model predicts the noise based on the current state at the current time step t , and then uses the predicted noise to calculate the state of the next time step t - 1 through the denoising update formula . Through iterative denoising from t = T to t = 0, the target parameters are generated.
[0055] For example, denoising step by step according to the sampling method of DDPM (Denoising Diffusion Probabilistic Models), the denoising update formula is as follows:
[0056]
[0057] where is a set of hyperparameters, which are restricted to start from = 10−4t and increase linearly with time t to = 0.02, = 1 - , is the noise variance, is the random noise.
[0058] As described above, the predicted noise includes conditional noise and unconditional noise. The conditional noise is the noise generated under the constraint of the first conditional information, and the unconditional noise represents the noise generated without constraints.
[0059] In one example, for each time step, the diffusion model predicts the noise at that step. In the Classifier-Free paradigm, the conditional and unconditional noises are recombined to generate the final noise prediction:
[0060]
[0061] where is a hyperparameter. represents the output of the diffusion model obtained with the conditional information of task n as the input and the predicted noise of the output of the diffusion model obtained with as the input .
[0062] It can be understood that predicting the conditional noise and unconditional noise of the diffusion model in the Classifier-Free manner not only simplifies the process of generating conditional noise but also maintains the prediction ability of unconditional noise.
[0063] As described above, the method for obtaining the model may further include obtaining the first attribute information of the user and the second attribute information of multiple objects; inputting the first attribute information and the second attribute information of the multiple objects into the target model to obtain multiple target objects, and the multiple target objects satisfy the first demand condition information.
[0064] The first attribute information may be information associated with the user. The first attribute information may include: the basic information and behavior information of the user. The basic information may be the age, gender, region, interest, preference, etc. of the user. The behavior information may be the user's web browsing data, commodity purchase data, etc.
[0065] The second attribute information may be information associated with the object to be recommended. The second attribute information may be the price, brand, function, etc. of the commodity.
[0066] In one example, the initial model is a movie recommendation model. For the requirements of a new recommendation task: the weights of the accuracy and diversity of the recommendation results satisfy the preset ratio of 8:2. The diffusion model is processed to obtain the target parameters. The target parameters are loaded into the initial model to obtain the target model for responding to the new recommendation task. When applying the target model, the first attribute information of the user and the second attribute information of the movie (such as year, type, lead actor, etc.) are input into the target model, and the target model outputs multiple recommended movies (target objects) for the user, and the accuracy and diversity of the multiple recommended movies satisfy the preset ratio of 8:2.
[0067] Figure 2 Schematically shows a flowchart of a diffusion model training method according to an embodiment of the present disclosure.
[0068] As described above, as Figure 2 shown, the diffusion model is obtained through the following operations S310 to S330.
[0069] In operation S310, first sample data is obtained. The first sample data includes second condition information and historical parameters. The second condition information represents the result requirements achieved by the initial model in performing historical tasks, and the historical parameters represent the model parameters required by the initial model when the execution result satisfies the second condition information.
[0070] In operation S320, with the second condition information as a constraint, the second initial parameters are denoised based on the diffusion model to obtain predicted parameters corresponding to each second condition information. The second initial parameters represent the parameter noise of the diffusion model regarding historical tasks.
[0071] In operation S330, according to the difference between the predicted parameters and the historical parameters, the parameters of the diffusion model are adjusted until the difference between the predicted parameters and the historical parameters is less than a threshold or the model converges, and a trained diffusion model is obtained.
[0072] Exemplarily, the first sample data is the original data set required for training the diffusion model. The first sample data includes second condition information and historical parameters. The second condition information can be the result requirements achieved by the initial model in performing historical tasks. The historical parameters can be the model parameters of the initial model when the execution result of performing historical tasks satisfies the second condition information.
[0073] For example, when the initial model is a recommendation model, the first sample data may include: the parameters of the recommendation model when the weights of the accuracy and diversity of the recommendation results satisfy 5:5, the parameters of the recommendation model when the weights of the accuracy and diversity of the recommendation results satisfy 1:9, the parameters of the recommendation model when the weights of the accuracy and diversity of the recommendation results satisfy 3:7, etc., historical parameters corresponding to the condition information of multiple historical tasks.
[0074] The second initial parameters can be the noise distribution of the original parameters (such as weights, biases, etc.) of the initial model in performing historical tasks. The second initial parameters can be obtained through the following operations: obtaining the original parameters of the initial model and adding noise to the original parameters according to multiple time steps to obtain the second initial parameters.
[0075] In one example, the first sample data is input into the diffusion model to be trained. The diffusion model uses the second conditional information of the historical task as a condition to denoise the second initial parameters. During the denoising process, the second initial parameters are optimized and adjusted, and the predicted parameters required for the initial model to execute the historical task are obtained. By comparing the predicted parameters with the historical parameters, the difference between the two is calculated, and the parameters of the diffusion model are continuously adjusted until the difference between the prediction result of the diffusion model and the historical parameters (the ideal result) is less than a threshold (for example, 0.5 or 1, etc.), and the trained diffusion model is obtained. The trained diffusion model can generate the target parameters of the initial model according to the conditional information of the new target task, and the execution result of the initial model applying the target parameters to execute the new target task can meet the first conditional information.
[0076] In another example, by comparing the predicted parameters with the historical parameters and continuously adjusting the parameters of the diffusion model until the predicted parameters output by the diffusion model no longer change significantly and the model converges, the trained diffusion model is obtained.
[0077] As described above, in operation S320, with the second conditional information as a constraint, the second initial parameters are denoised based on the diffusion model to obtain the predicted parameters corresponding to each second conditional information. In one implementable way, this operation may further include operation S321.
[0078] In operation S321, with the second conditional information as a constraint, the noise of the second initial parameters is gradually removed step by step according to multiple time steps to obtain the predicted parameters.
[0079] Among them, during the denoising process, the noise is predicted according to the current state at the current time step, and the current state is updated using the predicted noise to obtain the state at the next time step. The state is the intermediate representation of the diffusion model at the time step, and the state includes parameter information and noise information.
[0080] The predicted noise includes conditional noise and unconditional noise. The conditional noise is the noise generated under the constraint of the second conditional information, and the unconditional noise represents the noise generated without constraints. During the process of predicting the noise, the unconditional noise is predicted and generated with a target probability.
[0081] For the denoising during the training process of the diffusion model, reference can be made to the description of operation S221 above, which will not be elaborated here.
[0082] In one example, during the training process, the objective function of the diffusion model can be defined as follows:
[0083]
[0084] Among them, represents the historical task, Represents the time step of the diffusion process. Represents a diffusion model, whose input is the task At the step in the diffusion process , the task 's conditional information (for example, including the accuracy weight for condition one and the weight for diversity of condition two ), represents Gaussian noise.
[0085] Classifier-Free conditional training. To enable the diffusion model to recognize specific conditions and retain general generation capabilities, the conditions are randomly set to be empty.
[0086]
[0087] Among them, the weight of the task is randomly set to be empty according to the target probability.
[0088] Conditional noise: The noise predicted by the diffusion model according to conditional information (such as condition one); Unconditional noise: The noise directly predicted by the diffusion model without relying on any conditional information. At each step of denoising, the diffusion model recombines the conditional and unconditional noise predictions to generate the final predicted noise. During training, the conditional information is randomly discarded, and with a certain probability (for example, 10%, 20%, etc.), the conditional information is set to be empty, allowing the diffusion model to perform unconditional generation. Thus, the diffusion model learns to handle both conditional and unconditional situations simultaneously.
[0089] It can be understood that by randomly discarding the conditional information, the model is prevented from overly relying on the conditional information, thereby improving the flexibility and robustness of the model generation.
[0090] Figure 3 Schematically shows a flowchart of a method for obtaining first sample data according to an embodiment of the present disclosure.
[0091] As described above, in operation S310, first sample data is obtained. In one implementable manner, as Figure 3 shown, this operation may further include operations S410 to S440.
[0092] In operation S410, second sample data is obtained. The second sample data includes training data and second conditional information. The training data includes first attribute information of multiple users and second attribute information of multiple objects.
[0093] In operation S420, based on the second condition information of each historical task, input the training data into the target model to determine multiple target objects corresponding to the user.
[0094] In operation S430, based on the multiple target objects, determine the prediction condition information corresponding to the historical task.
[0095] In operation S440, adjust the parameters of the initial model until the difference between the prediction condition information and the second condition information is less than the threshold or the model converges, to obtain the historical parameters of the initial model for each second condition information.
[0096] Exemplarily, the first attribute information may be information associated with the user. The first attribute information may include: the basic information and behavioral information of the user. The basic information may be the age, gender, region, interests, preferences, etc. of the user. The behavioral information may be the user's web browsing data, commodity purchase data, etc.
[0097] The second attribute information may be the associated information of the object to be recommended. The second attribute information may be the price, brand, function, etc. of the commodity.
[0098] The second condition information characterizes the degree of emphasis shown by the execution results of the expected multiple target objects under different conditions.
[0099] The prediction condition information characterizes the degree of emphasis shown by the execution results of the multiple target objects under different conditions. For example, the weight ratio of the accuracy and diversity of the recommendation results for multiple objects for the user is 2:8.
[0100] In one example, the target model is a recommendation model. Input the first attribute information of the user and the second attribute information of the movie (such as year, type, lead actor, etc.) into the target model. The target model outputs multiple recommended movies (target objects) for the user. The weight ratio of the accuracy and diversity of the multiple recommended movies is 3:7, which is different from the second condition information (the weight ratio of the accuracy and diversity of the recommended movies is 2:8). Continuously adjust the parameters of the recommendation model until the difference between the prediction condition information of the recommendation result of the recommendation model and the second condition information (the ideal result) is less than the threshold (such as 0.5 or 1, etc.), to obtain the model parameters (historical parameters) of the recommendation model under the second condition information: the weight ratio of the accuracy and diversity of the recommended movies is 2:8. The obtained model parameters and the second condition information can be used as the first sample data for the diffusion model training.
[0101] Among them, use the relevant algorithm to calculate the gradient of the objective function for each parameter in the initial model. The objective function is the result of combining the loss functions corresponding to multiple requirements in the second condition information.
[0102] Given a specific task requirement , respectively represent the weights for different conditions under a specific task. Using the LinearScalarization method to optimize the objective, the model parameters suitable for the specific task are obtained:
[0103]
[0104] Among them, represents the weights for condition one (e.g., accuracy) and condition two (e.g., diversity). respectively represent the loss functions for the two conditions. represents the target parameters of the model.
[0105] According to the gradient, update the parameters of the initial model through the gradient descent method until the objective function reaches the minimum value, and the historical parameters of the initial model regarding each second condition information are obtained.
[0106] As described above, in operation S230, update the initial model using the target parameters to obtain the target model. In an implementable manner, this operation may further include operation S231.
[0107] In operation S231, load the target parameters into the adapter module of the initial model to obtain the target model.
[0108] Exemplarily, the adapter module may be a network layer inserted into the initial model for adapting to new tasks. When receiving a new task, fix the backbone parameters of the initial model and only update the parameters of the adapter module to obtain a target model adapted to the requirements of the new task.
[0109] The adapter parameters of the new task generated by the diffusion model will be directly loaded into the adapter module connected to the backbone network of the connection recommendation model, thereby forming a new customized recommendation model to respond to the recommendation result requirements of the new task.
[0110] Figure 4 Schematically shows the schematic diagram of the model acquisition method according to an embodiment of the present disclosure; Figure 5 Schematically shows the schematic diagram of the model acquisition method according to an embodiment of the present disclosure.
[0111] The model acquisition method of this embodiment includes three parts. The first part: Use the model to generate the training data of the diffusion model. The second part: Use the training data to train the diffusion model. The third part: Use the trained diffusion model to generate the model parameters corresponding to the condition information of the new task.
[0112] 1. The first part:
[0113] The first sample data for diffusion model training can be obtained by training a recommendation model. The acquisition of the first sample data includes operations S510 to S540.
[0114] In operation S510, second sample data is acquired. The second sample data includes training data and second condition information. The training data includes first attribute information of multiple users and second attribute information of multiple objects.
[0115] In operation S520, based on the second condition information of each historical task, the training data is input into the target model to determine multiple target objects corresponding to the user.
[0116] In operation S530, according to the multiple target objects, prediction condition information corresponding to the historical task is determined.
[0117] In operation S540, the parameters of the initial model are adjusted until the difference between the prediction condition information and the second condition information is less than the threshold or the model converges, and the historical parameters of the initial model for each second condition information are obtained.
[0118] In an example, referring to Figure 5 , the target model is a recommendation model, and a structure of an adapter for adapting to different tasks is added to the recommendation model. The basic parameters of the recommendation model are fixed, and the adapter parameters of the adapter module corresponding to the task in the backbone network need to be adjusted according to different task requirements (i.e., condition information). Correspondingly, adapter tuning (i.e., obtaining the first sample data): training the adapter parameters according to the requirements of different tasks to obtain the adapter parameters corresponding to the result requirements of different tasks, that is, task 1 corresponds to adapter 1, task 2 corresponds to adapter 2,..., task N corresponds to adapter N, and each adapter contains adapter parameters corresponding to the task. Taking a historical task as an example, the first attribute information of the user and the second attribute information of the movie (for example, year, type, starring, etc.) are input into the target model, and the target model outputs multiple recommended movies (target objects) for the user. The weight ratio of the accuracy and diversity of the multiple recommended movies is 3:7, which is different from the second condition information (the weight ratio of the accuracy and diversity of the recommended movies is 2:8). Update the parameters of the recommendation model according to the difference until the difference between the prediction condition information of the recommendation result of the recommendation model and the second condition information (ideal result) is less than the threshold (for example, 0.5 or 1, etc.), and the model parameters (historical parameters) of the recommendation model under the second condition information: the weight ratio of the accuracy and diversity of the recommended movies is 2:8 are obtained. The obtained model parameters and the second condition information can be used as the first sample data for diffusion model training.
[0119] 2. The second part:
[0120] The training process of the diffusion model includes operations S610 to S630.
[0121] In operation S610, obtain first sample data, where the first sample data includes second condition information and historical parameters. The second condition information represents the result requirements achieved by the initial model in performing historical tasks, and the historical parameters represent the model parameters required by the initial model when the execution result meets the second condition information.
[0122] In operation S620, using the second condition information as a constraint, perform denoising processing on the second initial parameters based on the diffusion model to obtain predicted parameters corresponding to each second condition information. The second initial parameters represent the parameter noise of the diffusion model regarding historical tasks.
[0123] In operation S630, according to the difference between the predicted parameters and the historical parameters, adjust the parameters of the diffusion model until the difference between the predicted parameters and the historical parameters is less than a threshold or the model converges, obtaining a trained diffusion model.
[0124] In one example, continue to refer to Figure 5 , input the first sample data obtained by adapter tuning into the diffusion model to be trained for multi-task parameter diffusion to update the parameters of the diffusion model. The diffusion model realizes data generation through a forward process (adding noise) and a reverse process (denoising). Forward process: Obtain the original parameters generated by the recommendation model according to the second condition information of the historical task, and add noise to the original parameters according to T time steps to obtain second initial parameters (Gaussian noise). Reverse process: The diffusion model uses the second condition information of the historical task as a condition, and performs denoising on the second initial parameters according to T time steps. During the denoising process, optimize and adjust the second initial parameters to predict the predicted parameters required for the initial model to perform historical tasks. By comparing the predicted parameters and the historical parameters, calculate the difference between the two, and continuously adjust the parameters of the diffusion model until the difference between the predicted result of the diffusion model and the historical parameters (ideal result) is less than a threshold (for example, 0.5 or 1, etc.), obtaining a trained diffusion model. The trained diffusion model can generate the target parameters of the initial model according to the condition information of the new target task (for example, the weights of the accuracy and diversity of the execution result in the target task condition information), and the execution result of the initial model applying the target parameters to perform the new target task can meet the first condition information.
[0125] 3. The third part:
[0126] The model obtaining method of this embodiment may include operation S640 to operation S660.
[0127] In operation S640, obtain the first condition information required for the initial model to perform the target task. The first condition information represents the result requirements expected for the initial model to perform the target task.
[0128] In operation S650, using the first conditional information as a constraint, denoise the first initial parameter based on the diffusion model to obtain the target parameter, where the first initial parameter characterizes the parameter noise of the initial model regarding the target task.
[0129] In operation S660, update the initial model using the target parameter to obtain the target model for performing the target task through the target model.
[0130] In one example, continue to refer to Figure 5 , for the recommendation model, in a new recommendation task, the first conditional information is that the weights of the accuracy and diversity of the recommendation result satisfy the first preset ratio of 8:2. Generate the adapter parameter using the diffusion model according to the conditional weights of the new task. During the process of obtaining the model parameters (adapter parameters) of the recommendation model regarding the new recommendation task, forward process: Add noise to the original parameters of the recommendation model over T time steps to obtain the first initial parameter. Reverse process: Input the first initial parameter into the diffusion model, given the first conditional information, the diffusion model starts from the first initial parameter and denoises over T time steps to generate the target parameter that satisfies the first conditional information. Directly load the target parameter into the adapter connecting the backbone network of the recommendation model, thereby forming a new customized recommendation model to respond to the recommendation result requirements of the new task.
[0131] According to different new task lists (including the conditional information of the new task), generate the adapter parameters of different new tasks using the trained diffusion model, and then the recommendation models adapted to different new tasks can be obtained.
[0132] Figure 6 Schematically shows one of the effect diagrams of the model acquisition method according to an embodiment of the present disclosure; Figure 7 Schematically shows the second effect diagram of the model acquisition method according to an embodiment of the present disclosure.
[0133] To verify the effect of applying the target parameter of the conditional information regarding the target task obtained by the model acquisition method of the embodiments of the present disclosure to the recommendation model and performing the target task. Among them, the backbone network of the recommendation model respectively adopts a sequential recommendation model based on the self-attention mechanism (SASRec), a sequential recommendation model based on the Gated Recurrent Unit (GRU4Rec), and a sequential recommendation model based on the time-based self-attention mechanism (TiSASRec).
[0134] The algorithms include:
[0135] (1) Retrain: Retrain the obtained recommendation model according to the conditional information of the target task.
[0136] (2)CMR: Controllable multi-objective re-ranking with policy hypernetworks is a controllable re-ranking method for multi-objective ranking that combines a policy hypernetwork to optimize the ranking tasks of multiple objectives. CMR can be used for effective re-ranking under multiple ranking criteria and can adjust these objectives according to requirements.
[0137] (3)Soup: Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time is a method that improves the accuracy of a model by averaging the weights of multiple fine-tuned models.
[0138] (4)MMR: The use of mmr, diversity-based reranking for reordering documents and producing summaries is a document ranking and summary generation technique aimed at selecting the most representative documents or sentences in a retrieval system.
[0139] (5)PaDiRec is a recommendation model that applies the target parameters of the conditional information about the target task obtained by the model acquisition method of the present disclosure embodiment.
[0140] (6)LLM (Large Language Model) is a large language model capable of understanding and generating natural language text.
[0141] For Task 1: Movie Recommendation, Task 2: Food Recommendation on Platform A; and Task 3: Industrial Datasets on Recommendation respectively.
[0142] Evaluation Metric 1: The correlation between Condition 1 and the algorithm (Avg.HV) is obtained by averaging the Hypervolume of the multi-objective task under multiple task weights, which measures the overall performance of the algorithm on a single task.
[0143] Evaluation Metric 2: The correlation between Condition 2 and the algorithm (Pearsonr-aPearsonr-d), where Pearsonr-aPearsonr-d measures the correlation between two objectives (accuracy and diversity) and the algorithm of the retrained recommendation model respectively.
[0144] The recommendation model obtained by the present disclosure embodiment and the recommendation models using different algorithms respectively execute the same target task, and the results are asFigure 6 As shown. Analysis Figure 6 It can be obtained that: The recommendation model applied to the target parameters of the conditional information about the target task obtained by using the model acquisition method of the present disclosure embodiment has good overall performance in the recommendation task. In the backbone network using the self-attention mechanism for the movie recommendation task, the recommendation model of this method can maintain a small gap with the performance of the recommendation model retrained according to the conditional information of the target task. For other tasks and algorithms, the overall performance of the recommendation model using this method is better than the overall performance of the recommendation model retrained according to the conditional information of the target task.
[0145] When the recommendation model obtained by using the present disclosure embodiment and the recommendation models using different algorithms respectively execute multiple target tasks, the performance regarding the result requirements is as follows Figure 7 As shown. Analysis Figure 7 It can be obtained that: Under multiple sets of conditional information: the weights of accuracy and diversity, the change trends of the change curves of the recommendation model applied to the target parameters of the conditional information about the target task obtained by using the model acquisition method of the present disclosure embodiment and the recommendation model retrained according to the conditional information of the target task are the closest.
[0146] Based on the above model acquisition method, the present disclosure also provides a model acquisition device. The following will be combined with Figure 8 to describe this device in detail.
[0147] Figure 8 The structural block diagram of the model acquisition device according to the embodiment of the present disclosure is schematically shown.
[0148] As Figure 8 shown, the model acquisition device 700 of this embodiment includes an acquisition module 710, a processing module 720, and an update module 730.
[0149] The acquisition module 710 is used to acquire the first conditional information required for the initial model to execute the target task, and the first conditional information characterizes the result requirements expected for the initial model to execute the target task. In one embodiment, the acquisition module 710 can be used to execute the operation S210 described above, which will not be elaborated here.
[0150] The processing module 720 is used to perform denoising processing on the first initial parameter based on the diffusion model with the first conditional information as a constraint to obtain the target parameter, and the first initial parameter characterizes the parameter noise of the initial model regarding the target task. In one embodiment, the processing module 720 can be used to execute the operation S220 described above, which will not be elaborated here.
[0151] The update module 730 is configured to update the initial model using the target parameters to obtain the target model, so as to perform the target task using the target model. In one embodiment, the update module 730 may be configured to perform the operation S230 described above, which will not be elaborated herein.
[0152] According to an embodiment of the present disclosure, any plurality of the acquisition module 710, the processing module 720, and the update module 730 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present disclosure, at least one of the acquisition module 710, the processing module 720, and the update module 730 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application specific integrated circuit (ASIC), or any other reasonable manner of integrating or packaging circuits, etc., implemented by hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in a suitable combination of any several of them. Alternatively, at least one of the acquisition module 710, the processing module 720, and the update module 730 may be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.
[0153] Figure 9 A block diagram of an electronic device suitable for implementing the model acquisition method according to an embodiment of the present disclosure is schematically shown.
[0154] As Figure 9 shown, the electronic device 800 according to an embodiment of the present disclosure includes a processor 801, which can perform various appropriate actions and processes according to a program stored in a read only memory (ROM) 802 or a program loaded from a storage section 808 into a random access memory (RAM) 803. The processor 801 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 801 may also include on board memory for caching purposes. The processor 801 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0155] In the RAM 803, various programs and data required for the operation of the electronic device 800 are stored. The processor 801, the ROM 802, and the RAM 803 are connected to each other via the bus 804. The processor 801 performs various operations of the method flow according to the embodiments of the present disclosure by executing the programs in the ROM 802 and / or the RAM 803. It should be noted that the programs can also be stored in one or more memories other than the ROM 802 and the RAM 803. The processor 801 can also perform various operations of the method flow according to the embodiments of the present disclosure by executing the programs stored in one or more memories.
[0156] According to an embodiment of the present disclosure, the electronic device 800 may further include an input / output (I / O) interface 805, and the input / output (I / O) interface 805 is also connected to the bus 804. The electronic device 800 may further include one or more of the following components connected to the I / O interface 805: an input portion 806 including a keyboard, a mouse, etc.; an output portion 807 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 808 including a hard disk, etc.; and a communication portion 809 including a network interface card such as a LAN card, a modem, etc. The communication portion 809 performs communication processing via a network such as the Internet. The drive 810 is also connected to the I / O interface 805 as needed. A removable medium 811, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 810 as needed so that a computer program read from it can be installed into the storage portion 808 as needed.
[0157] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiments of the present disclosure is implemented.
[0158] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include one or more memories other than the above-described ROM 802 and / or RAM 803 and / or ROM 802 and RAM 803.
[0159] An embodiment of the present disclosure further includes a computer program product, which includes a computer program, and the computer program contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to enable the computer system to implement the model acquisition method provided by the embodiment of the present disclosure.
[0160] When the computer program is executed by the processor 801, it executes the above functions defined in the system / apparatus of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0161] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program may also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 809, and / or installed from the removable medium 811. The program code included in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0162] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 809, and / or installed from the removable medium 811. When the computer program is executed by the processor 801, it executes the above functions defined in the system of the embodiment of the present disclosure. According to an embodiment of the present disclosure, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.
[0163] According to embodiments of the present disclosure, program code for executing the computer programs provided by the embodiments of the present disclosure may be written in any combination of one or more programming languages. Specifically, these computing programs may be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0164] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0165] Those skilled in the art can understand that the features recited in the various embodiments and / or claims of the present disclosure can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in the various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0166] The embodiments of the present disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A model acquisition method, comprising: Acquire first condition information required for the initial model to perform the target task, where the first condition information represents the result requirement expected to be achieved by the initial model in performing the target task; Using the first condition information as a constraint, denoising the first initial parameter based on a diffusion model to obtain a target parameter, where the first initial parameter represents parameter noise of the initial model with respect to the target task; The initial model is updated using the target parameters to obtain a target model, so as to perform the target task through the target model.
2. The method according to claim 1, taking the first condition information as a constraint, performing denoising processing on the first initial parameter based on a diffusion model to obtain a target parameter, comprising: Using the first condition information as a constraint, gradually removing noise from the first initial parameter according to multiple time steps to obtain a target parameter; The target parameters represent model parameters required by the initial model when the execution result meets the first requirement information.
3. According to the method of claim 2, in the denoising process, the noise is predicted based on the current state of the current time step, and the current state is updated with the predicted noise to obtain the state of the next time step, wherein the state is the intermediate representation of the diffusion model at the time step, and the state includes parameter information and noise information.
4. According to the method of claim 3, the predicted noise includes conditional noise and unconditional noise, the conditional noise is noise generated under the constraint of the first condition information, and the unconditional noise represents noise generated without constraint.
5. The method according to claim 1, further comprising: Acquire first attribute information of a user and second attribute information of multiple objects; The first attribute information and the second attribute information of the multiple objects are input into the target model to obtain multiple target objects, and the multiple target objects satisfy the first condition information.
6. According to the method of claim 1, the diffusion model is obtained by the following operations: Acquire first sample data, where the first sample data includes second condition information and historical parameters, where the second condition information represents the result requirement achieved by the initial model when executing the historical task, and the historical parameters represent the model parameters required by the initial model when the execution result satisfies the second condition information; Using the second condition information as a constraint, denoising the second initial parameters based on the diffusion model to obtain prediction parameters corresponding to each second condition information, wherein the second initial parameters represent parameter noise of the diffusion model with respect to the historical tasks; According to the difference between the predicted parameter and the historical parameter, the parameters of the diffusion model are adjusted until the difference between the predicted parameter and the historical parameter is less than a threshold or the model converges, thereby obtaining a trained diffusion model.
7. The method according to claim 6, taking the second condition information as a constraint, performing denoising processing on the second initial parameters based on a diffusion model, and obtaining prediction parameters corresponding to each of the second condition information comprises: Taking the second condition information as a constraint, gradually removing noise from the second initial parameter according to multiple time steps to obtain a predicted parameter; Wherein, in the denoising process, the noise is predicted according to the current state of the current time step, and the predicted noise is used to update the current state to obtain the state of the next time step, the state is the intermediate representation of the diffusion model in the time step, and the state includes parameter information and noise information; The predicted noise includes conditional noise and unconditional noise. The conditional noise is the noise generated under the constraint of the second condition information, and the unconditional noise represents the noise generated without constraints. In the process of predicting noise, the unconditional noise is generated by predicting with a target probability.
8. The method according to claim 6, obtaining the first sample data comprises: Acquire second sample data, where the second sample data includes training data and the second condition information, where the training data includes first attribute information of multiple users and second attribute information of multiple objects; Based on the second condition information of each of the historical tasks, inputting the training data into the target model to determine a plurality of target objects corresponding to the user; Determining prediction condition information corresponding to the historical task according to the multiple target objects; The parameters of the initial model are adjusted until the difference between the predicted condition information and the second condition information is less than a threshold or the model converges, thereby obtaining the historical parameters of the initial model with respect to each of the second condition information.
9. The method according to claim 1, using the target parameters to update the initial model to obtain the target model, comprising: The target parameters are loaded into the adapter module of the initial model to obtain a target model.
10. A model obtaining device, comprising: An acquisition module, used to acquire first condition information required for the initial model to execute the target task, wherein the first condition information represents the result requirement expected to be achieved by the initial model in executing the target task; A processing module, configured to perform denoising processing on a first initial parameter based on a diffusion model using the first condition information as a constraint to obtain a target parameter, wherein the first initial parameter represents parameter noise of the initial model with respect to the target task; The loading module is used to update the initial model using the target parameters to obtain a target model so as to perform the target task using the target model.