Model updating method and device, equipment, storage medium and program product
By dynamically adjusting the guidance ratio of the teacher model, the problems of low model reasoning efficiency and degraded teacher model output quality caused by CFG technology are solved, and efficient model updates and output quality improvements are achieved.
Patent Information
- Application Number
- CN202510737095.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-09-19
AI Technical Summary
When existing machine learning models use Classifier-Free Guidance (CFG) technology, the reasoning steps are numerous and the reasoning time is long, which affects the efficiency of the model in executing tasks. In addition, during the model distillation process, the output quality of the teacher model may decline, affecting the update effect of the student model.
By dynamically adjusting the guidance ratio of the teacher model, the guidance ratio of the student model is updated based on the difference between the visual content output by the teacher model and the output of the student model, ensuring the quality of the teacher model output, and using the quality level of the assessed visual content to adjust the guidance ratio to improve the update effect of the student model.
While maintaining the high inference efficiency of the student model, the quality of the model output is improved, the problem of quality degradation of the teacher model output is avoided, and the effect of model distillation is improved.
Smart Images

Figure CN120671738A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and more particularly, to methods, apparatuses, electronic devices, computer-readable storage media, and computer program products for model updating. Background Art
[0002] With the development of machine learning technology, machine learning models can now be used to perform tasks in a variety of application environments. For example, trained machine learning models can be used to generate content, which can be any suitable content such as video, images, and text. To meet the ever-increasing task requirements, machine learning models are being constructed with increasingly complex structures and an increasing number of parameters to support these complex tasks. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for model updating is provided. The method includes: updating the student model based on the difference between the first visual content output by the teacher model and the second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information and a guidance ratio, the guidance ratio indicating the degree of influence of the positive cue word information and the negative cue word information on the first visual content, and the student model is configured to generate the second visual content based on the positive cue word information; in the updating process of the student model, using the teacher model to generate evaluation visual content corresponding to the evaluation cue word information based on the evaluation cue word information; in response to the quality level of the evaluation visual content being lower than the first threshold quality, updating the guidance ratio of the teacher model; and using the teacher model to perform an update on the student model based on the updated guidance ratio.
[0004] In a second aspect of the present disclosure, a device for model updating is provided. The device includes: a student model updating module, configured to update the student model based on the difference between the first visual content output by the teacher model and the second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive prompt word information, negative prompt word information and a guidance ratio, the guidance ratio indicates the degree of influence of the positive prompt word information and the negative prompt word information on the first visual content, and the student model is configured to generate the second visual content based on the positive prompt word information; an evaluation content generation module, configured to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information using the teacher model during the update process of the student model; a guidance ratio update module, configured to update the guidance ratio of the teacher model in response to the quality level of the evaluation visual content being lower than the first threshold quality; and an update execution module, configured to execute the update of the student model based on the updated guidance ratio using the teacher model.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the electronic device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, the method of the first aspect is implemented.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, wherein the computer-executable instructions implement the method according to the first aspect of the present disclosure when executed by a processor.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented;
[0011] Figure 2 shows an example architecture for model updating according to some embodiments of the present disclosure;
[0012] Figure 3 An example process for updating a guidance ratio according to some embodiments of the present disclosure is shown;
[0013] Figure 4 A flowchart of a method for model updating according to some embodiments of the present disclosure is shown;
[0014] Figure 5 An exemplary structural block diagram of an apparatus for model updating according to some embodiments of the present disclosure is shown; and
[0015] Figure 6 A block diagram of an electronic device in which one or more embodiments of the present disclosure may be implemented is shown. DETAILED DESCRIPTION
[0016] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0018] Herein, unless explicitly stated otherwise, executing a step “in response to A” does not mean executing the step immediately after “A” but may include one or more intermediate steps.
[0019] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0020] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0021] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested to be performed will require obtaining and using the user's personal information, so that the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium and other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.
[0022] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0023] It is understandable that the above notification and the process of obtaining user authorization are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0024] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training is completed, corresponding outputs can be generated for given inputs. The generation of the model can be based on machine learning technology. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this article, "model" may also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably in this article.
[0025] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs. It typically includes an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence so that the output of the previous layer is provided as input to the next layer, where the input layer receives the input of the neural network and the output of the output layer serves as the final output of the neural network. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each of which processes the input from the previous layer.
[0026] Generally speaking, machine learning can be roughly divided into three stages, namely the training stage, the testing stage, and the application stage (also called the inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are continuously updated iteratively until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association between input and output (also called input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained through training to determine the corresponding model output.
[0027] Diffusion models, also known as diffusion probability models, are a type of generative model in which the data generation process is based on a pair of Markov processes, namely the forward diffusion process and the backward denoising process. The forward diffusion process (expressed as: is the stepwise interference data x (0) ~q(x(0) ), through T stepwise noise addition steps x (1:T) =x1,…,x (t-1) ,x (t) ,…,x (T) , and obtain the static noise distribution x (T) ~q noise Through model training, the learned backward denoising process (expressed as: Perform the reverse process and gradually denoise the samples towards the data distribution to obtain data x (0) ~q(x (0) ). It can be seen that the backward denoising process can correspond to the desired data modeling process and finally obtain the desired data.
[0028] Figure 1 A schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, a model update system 120 can be deployed in an electronic device 110. The model update system 120 can also be referred to as a model distillation system. Model distillation is a technique for transferring knowledge from a complex model (also referred to as a teacher model) to a simpler model (also referred to as a student model). Its core goal is to significantly reduce the computational complexity and storage requirements of the model while maintaining model performance.
[0029] The model update system 120 can, for example, use the training samples 105 to perform model distillation on the teacher model 130 and the student model 140. The teacher model 130 is typically a large, complex model with high accuracy and generalization performance, while the student model 140 is a smaller, simpler model with lower computational and storage costs. The teacher model 130 and the student model 140 can be used to perform the same task. As an example, the teacher model 130 and the student model 140 can both be configured to perform a visual generation task. In a visual generation task, the model input can include text, images, and / or videos, and the model output is visual content generated based on the model input. The visual content includes images or videos. Both the teacher model 130 and the student model 140 can be based on any appropriate model structure, including but not limited to a Transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN), a deep neural network (DNN), and the like. In some embodiments, both the teacher model 130 and the student model 140 can be implemented based on a diffusion model.
[0030] During the training phase, a machine learning model based on a diffusion model can train the model by continuously adding noise to known data and predicting the added noise. During the application phase, a machine learning model based on a diffusion model can continuously add noise to the model input to determine the noise corresponding to the model input, and then continuously perform a denoising process on the noise to obtain the model output corresponding to the model input. For example, when generating a video, a machine learning model based on a diffusion model can generate a video that meets the conditions by continuously performing a denoising process on random noise under certain constraints.
[0031] The model updating system 120 may, for example, process the training samples 105 using the teacher model 130 and the student model 140, respectively, and obtain the model outputs of the teacher model 130 and the student model 140. The model updating system 120 may, for example, update the student model 140 based on the difference between the model outputs of the teacher model 130 and the model outputs of the student model 140.
[0032] The electronic device 110 may be any type of device with computing capabilities, including a terminal device or a server device. The terminal device may be any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof.
[0033] A server-side device can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. Server-side devices may include, for example, computing systems / servers such as mainframes, edge computing nodes, and computing devices in cloud environments.
[0034] It should be understood that the structure and function of the various elements in the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.
[0035] To further improve model capabilities, the use of Classifier-Free Guidance (CFG) technology to enhance the generation effect of machine learning models has also been proposed. The core of CFG technology is that the reason model determines the positive output corresponding to positive prompt word information and the negative output corresponding to negative prompt word information, and determines the final model output based on the guidance scale, positive output, and negative output. Positive prompt word information can guide the model to generate content that matches the positive prompt word, while negative prompt word information can guide the model to avoid generating content related to the negative prompt word. In addition, the guidance scale used can be used to control the degree of influence of positive prompt word information and negative prompt word information on the model output. The model output using CFG technology will be of higher quality than the model output obtained directly based on single prompt word information (for example, only positive prompt word information), which can effectively improve the quality of the model output and enhance the model's task execution capability.
[0036] However, the model using CFG technology needs to determine the corresponding output for the two prompt word information respectively, and then calculate the two outputs based on the guidance ratio to determine the final model output. This will result in more reasoning steps and longer reasoning time for the model, which will affect the efficiency of the model in performing tasks. In addition, as mentioned earlier, in order to meet the increasing task requirements, machine learning models are constructed with increasingly complex structures and an increasing number of parameters to support the requirements of complex tasks. In cases where the model itself has large parameters and a more complex structure, the use of CFG technology will further affect the efficiency of the model, and model distillation technology is required to help accelerate the model.
[0037] To conserve graphics memory, a weighted average (e.g., exponential moving average) of the student model is typically used as the teacher model, combined with CFG for distillation. This means the student model learns the CFG inference results of the teacher model. In model distillation, EMA (Exponential Moving Average) is a commonly used parameter update technique. EMA can be considered a smoothing method for model parameters. Unlike traditional optimization algorithms such as stochastic gradient descent (SGD), which simply use the current gradient to update model parameters, EMA in model distillation combines parameter values from previous update steps to calculate the current parameter update. Using EMA can make the student model's parameter updates more stable. During the distillation process, the student model needs to continuously adjust its parameters to approximate the output of the teacher model. Overly drastic parameter updates can lead to unstable learning processes, such as oscillations. EMA, by smoothing updates, allows the student model to more smoothly learn the knowledge of the teacher model, reducing training instability caused by overly rapid parameter updates.
[0038] By performing model refinement based on CFG reasoning, it is possible to ensure high reasoning efficiency for the student model while allowing the learning model to perform reasoning based solely on positive cues, rather than on negative cues, while still achieving the superior generation results achieved through CFG reasoning. In this model distillation process, with the introduction of CFG technology, the guidance ratios for the student and teacher models are often initially set to the same value. As the number of model training iterations increases, the guidance ratio for the student model during reasoning needs to gradually reduce the impact of negative cues on the model output to avoid quality issues such as oversaturation. However, since the teacher model is determined by the student model through EMA during the model distillation process, the teacher model also needs to gradually reduce the impact of negative cues on its output to ensure the quality of its output. This means that if the guidance ratio of the teacher model is not changed, it may not be possible to ensure that the teacher model's output is always sufficient to meet expectations and thus sufficiently guide the parameter updates of the student model.
[0039] In view of this, according to an embodiment of the present disclosure, an improved solution for model updating is provided, which proposes updating the guidance ratio of the CFG of the teacher model to ensure the quality of the model output of the teacher model. This can help improve the effect of model distillation. According to the solution of an embodiment of the present disclosure, the student model is updated based on the difference between the first visual content output by the teacher model and the second visual content output by the student model. The teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information and the guidance ratio. The guidance ratio indicates the degree of influence of the positive cue word information and the negative cue word information on the first visual content. The student model is configured to generate the second visual content based on the positive cue word information. During the update process of the student model, the teacher model is used to generate evaluation visual content corresponding to the evaluation cue word information based on the evaluation cue word information. In response to the quality level of the evaluation visual content being lower than the first threshold quality, the guidance ratio of the teacher model is updated. The teacher model is used to perform an update on the student model based on the updated guidance ratio.
[0040] In this way, the guidance ratio of the CFG-introduced teacher model can be dynamically adjusted during the student model update process to dynamically adjust the impact of negative cue word information on the visual content output by the teacher model. This can avoid the problem of the quality of the visual content output by the teacher model degrading as the number of model distillation iterations increases. This ensures the quality of the teacher model's model output and helps improve the update effect of the student model.
[0041] The following will continue to describe some example embodiments of the present disclosure with reference to the accompanying drawings. The model updating method involved in the present disclosure can be implemented at the electronic device 110, and specifically can be implemented at the model updating system 120. The present disclosure uses the teacher model 130 and the student model 140 as an example to describe the example of performing a visual content generation task (that is, the model outputs of the teacher model 130 and the student model 140 are both visual content). The visual content includes images and / or videos.
[0042] Figure 2 An example architecture 200 for model updating according to some embodiments of the present disclosure is shown. The example architecture 200 can be implemented at the model updating system 120. Box 210 in the example architecture 200 shows an example process of performing a model update on the student model 140 (which can be referred to as performing a model iteration). The teacher model 130 is configured to generate visual content 205 (for example, it can be referred to as first visual content) based on positive cue word information 201, negative cue word information 202 and a guidance ratio 203, where the guidance ratio 203 indicates the degree of influence of the positive cue word information 201 and the negative cue word information 202 on the visual content 205. As an example, the teacher model 130 can generate visual content 205 based on the following formula:
[0043] t i =scale×t i c -(scale-1)×t i uc (1)
[0044] where t i , t i c and t i uc Respectively represent the final model output (i.e., visual content 205) of the teacher model 130 for the i-th positive prompt word information, the positive output determined based on the positive prompt word information 201, and the negative output determined based on the negative prompt word information 202, and scale indicates the guidance ratio 203. In formula (1), the smaller the guidance ratio scale, the smaller the impact of the negative prompt word information on the model output. When the guidance ratio scale = 1, it means that the negative prompt word has no effect on the model output, that is, the model does not need to perform reasoning based on the negative prompt word. It can be understood that formula (1) is only an example, and the CFG technology adopted by the teacher model 130 can actually be expressed based on other appropriate formulas. For example, the formula can be involved as the larger the value of scale, the higher the impact of the negative prompt word on the model output.
[0045] In the initial stage of model distillation, the student model 140 is configured to be identical to the teacher model 130. The purpose of model distillation is to enable the student model 140 to generate visual content 204 (e.g., which can be referred to as second visual content) based only on the positive cue word information 201, such as the guidance scale used by the student model 140 in the CFG reasoning based on formula (1) = 1.
[0046] During the model distillation process, the model update system 120 can update the student model 140 based on the difference 206 between the visual content 205 output by the teacher model 130 and the visual content 204 output by the student model 140. The model update system 120 can use any appropriate loss function such as mean square error (MSE), cross-entropy loss (Cross-Entropy Loss), multi-classification cross-entropy loss (Categorical Cross-Entropy Loss) to calculate the difference 206 between the visual content 205 and the visual content 204. For example, in the process of performing inference based on the MSE loss, the loss function can be expressed as Ls=MSE(si,ti), where si represents the model output of the student model 140 for the i-th positive prompt word information. It should be understood that the present disclosure does not limit the loss function of model distillation. The ultimate goal of model updating is to reduce the difference 206 as much as possible.
[0047] The model updating system 120 can perform multiple updates (or multiple iterations) on the student model 140 by repeatedly executing the example process shown in block 201. It is understood that the positive prompt word information and negative prompt word information applied in different updates are different. After each execution of the example process shown in block 201, the model updating system 120 can use the evaluation prompt word information to determine (230) whether to update the guidance ratio 203. Figure 3 The detailed process of using the evaluation prompt word information to determine whether to update the guidance ratio 203 is described in detail, which is not repeated here. In some embodiments, to reduce the amount of calculation, the model system 120 can also determine (220) whether the evaluation condition of the guidance ratio 203 is met each time the example process shown in box 201 is completed. If it is determined that the evaluation condition is met, it determines whether to update the guidance ratio 203.
[0048] The evaluation condition may indicate that the number of updates to the student model 140 reaches a predetermined number of intervals (Interval), that is, whether to update the guidance ratio 203 is determined every certain number of times. For example, if the current number of updates is the i-th time, and the predetermined number of intervals is Interval val , then i can be Inter val In the case of integer divisibility (for example, i%Inter val ==0), which means that after passing Inter val If the number of updates to the student model 140 exceeds a predetermined number of updates, the model update system 120 may determine that the evaluation condition is satisfied. For example, if the predetermined number of updates to the student model 140 exceeds 50, the model update system 120 may determine that the evaluation condition is satisfied.
[0049] The evaluation condition may indicate that the quality of the visual content 204 output by the updated student model 140 reaches a threshold quality (e.g., a second threshold quality). That is, in this case, the model update system 120 may determine that the evaluation condition is satisfied in response to the visual content 204 output by the student model 140 having a high quality. In some embodiments, the second threshold quality may include multiple different thresholds, and the model update system 120 may determine that the evaluation condition is satisfied each time the visual content output by the student model 140 reaches a threshold, and perform an evaluation on the guidance ratio 203. For example, the second threshold quality may include 50, and the quality of the visual content 204 output by the student model 140 may be scored, and the evaluation condition is determined to be satisfied in response to the visual content 204 having a score of 50.
[0050] The model update system 120 may, in response to the evaluation condition being not satisfied, execute block 250. At block 250, the model update system 120 may continue to update the student model 140 based on the current guide ratio 203. For example, if N updates have been performed on the student model 140, in this case, the model update system 120 may maintain the current guide ratio and perform the N+1th model update on the student model 140 based on the guide ratio. In response to the evaluation condition being satisfied, the model update system 120 may determine (230) whether to update the guide ratio 203 using the evaluation prompt word information. In response to determining not to update the guide ratio 203, the model update system 120 may execute block 250.
[0051] In some embodiments, the model updating system 120 may also determine whether to update the guidance ratio of the teacher model 130 in each round of updating of the student model 140, or the model updating system 120 may depend on other update conditions to determine whether to update the guidance ratio of the teacher model 130. The embodiments of the present disclosure are not limited in this respect.
[0052] In response to determining that the guidance ratio 203 needs to be updated, the model updating system 120 can update the guidance ratio 203 in any appropriate manner. In some embodiments, the negative prompt word information 202 indicated by the updated guidance ratio 203 should have a lower impact on the visual content 205 output by the teacher model 130 than the negative prompt word information 202 indicated by the guidance ratio 203 before the update. In some embodiments, the model updating system 120 can update the guidance ratio 203 based on (240) a predetermined scaling factor to obtain a first guidance ratio. Taking the predetermined scaling factor as an example, the model updating system 120 can determine that the updated guidance ratio 203 is ratio*scale.
[0053] The value of the predetermined scaling factor is associated with the formula applied by the teacher model 130. For example, if the teacher model 130 determines the model output based on the above formula (1), and the degree of influence of the negative prompt word information 202 on the visual content 205 decreases as the guidance ratio 203 decreases, then the predetermined scaling factor should be a positive number less than 1 (i.e., the predetermined scaling factor is between (0, 1)). If the teacher model 130 determines the model output based on other formulas, and the degree of influence of the negative prompt word information 202 on the visual content 205 decreases as the guidance ratio 203 increases, then the predetermined scaling factor should be a positive number greater than 1.
[0054] For example, the model updating system 120 can directly determine the first guiding ratio as the updated guiding ratio 203. In some embodiments, the model updating system 120 can select the updated guiding ratio from the first guiding ratio and the predetermined guiding ratio based on the comparison between the first guiding ratio and the predetermined guiding ratio. Taking the case where the teacher model 130 determines the model output based on the above formula (1) and the predetermined guiding ratio is 1 as an example, the model updating system 120 can determine the larger value of the first guiding ratio and 1 as the updated guiding ratio 203. That is, the updated guiding ratio 203 is max(1, ratio*scale). It can be understood that if the teacher model 130 determines the model output based on other formulas, the model updating system 120 may also determine the smaller value of the first guiding ratio and the predetermined guiding ratio as the updated guiding ratio.
[0055] In some embodiments, the training goal of the student model 140 can be determined based on this predetermined guidance ratio. Still taking the example of the teacher model 130 determining the model output based on the above formula (1) and the predetermined guidance ratio being 1, the training goal of the student model 140 can be to be trained to be able to perform CFG reasoning with a guidance ratio of 1, that is, to perform reasoning based only on positive prompt word information without considering negative prompt word information, thereby achieving the same effect as the teacher model 140.
[0056] After the guidance ratio 203 is updated, the model updating system 120 can use the teacher model 130 to update the student model 140 based on the updated guidance ratio 203. For example, if the student model 140 has been updated N times, the model updating system 120 can perform the N+1th model update on the student model 140 based on the updated guidance ratio 203.
[0057] In some embodiments, the model updating system 120 may also determine (260) whether the model update target corresponding to the student model 140 is satisfied each time the example process shown in block 201 is completed. The model update target may be, for example, based on whether the number of updates to the student model 140 reaches a predetermined number (e.g., reaching a maximum number of iterations S). max ). Taking the predetermined number of times as 500 as an example, the model update system 120 can also determine that the model update target is met in response to the number of updates of the student model 140 reaching 500. The model update target can be, for example, based on the content quality of the visual content output by the updated student model 140 reaching a threshold quality (for example, it can be called a third threshold quality). That is, the model update system 120 can determine that the model update target is met in response to the content quality of the visual content output by the updated student model 140 being high. Currently, the model update target can also be based on any other appropriate content, and the present disclosure is not limited to this.
[0058] The model update system 120 may stop (270) updating the student model 140 in response to the model update target being met. At this point, the model distillation process of the learning model 140 may be considered complete. It is understood that when the model update target is met, the model update system 120 does not need to determine whether the evaluation conditions are met, determine whether to update the guidance ratio, etc., and the model update system 120 does not need to update the guidance ratio. The model update system 120 may continue (280) updating the student model 140 in response to the model update target not being met. The model update system 120 may, for example, determine whether the evaluation conditions are met, determine whether to update the guidance ratio, etc. in response to the model update target not being met.
[0059] Figure 3 An example process 300 for updating a guidance scale is shown, according to some embodiments of the present disclosure.
[0060] At block 310, the model update system 120 utilizes the teacher model to generate evaluation visual content corresponding to the evaluation cue word information based on the evaluation cue word information. Taking the example above where the positive and negative cue word information are derived from a training sample set, the evaluation cue word information can be derived from another evaluation sample set. The evaluation sample set can include multiple evaluation cue word information and the corresponding labeled visual content.
[0061] In box 320, the model update system 120 can determine whether the quality level of the evaluated visual content is lower than the first threshold quality. The model update system 120 can determine the quality level of the evaluated visual content in any appropriate manner. The model update system 120 can, for example, determine the quality level of the evaluated visual content based on predetermined rules, algorithms, or using a specific evaluation model. In some embodiments, since the size of the guidance ratio will affect the saturation, content diversity, content generation efficiency, fineness, aesthetic quality, etc. of the visual content generated by the model, the model update system 120 can determine the quality level of the evaluated visual content based on one or more of the saturation score of the evaluated visual content, the content diversity score of the evaluated visual content, the content generation efficiency score of the evaluated visual content, the fineness score of the evaluated visual content, and the aesthetic quality score of the evaluated visual content. The saturation score, content diversity score, content generation efficiency score, etc. here can be determined, for example, using a specific scoring model.
[0062] It should be noted that, in some embodiments, according to the above formula (1), the impact of the guidance ratio on the content generation efficiency of the visual content is only related to whether the value of the guidance ratio is greater than 1 (because when the guidance ratio is equal to 1, there is no need to consider the negative prompt word information). When the guidance ratio is 1, the content generation efficiency score can be a higher value. When the guidance ratio is greater than 1, no matter how large the guidance ratio is (for example, 3 or 5), the content generation efficiency score can be a fixed, smaller value. For example, the content generation efficiency score when the guidance ratio is 1 can be twice the content generation efficiency score when the guidance ratio is greater than 1. Of course, in actual applications, the impact of the guidance ratio on the content generation efficiency can also be affected by many factors such as the model itself and the hardware. In this case, when the guidance ratio is greater than 1, the content generation efficiency score may decrease appropriately as the guidance ratio increases.
[0063] The model updating system 120 may, for example, determine that the quality level of the evaluated visual content is high in response to the saturation score, content diversity score, content generation efficiency score, fineness score, similarity score, aesthetic quality score, etc. corresponding to the evaluated visual content being within their respective corresponding predetermined ranges. For example, if the saturation score corresponding to the evaluated visual content is too high or too low, meaning that there is a problem of oversaturation or undersaturation, the quality level of the evaluated visual content is low; conversely, the quality level of the evaluated visual content is high. For another example, the higher the similarity score with the labeled visual content (e.g., greater than 90), the higher the quality level of the evaluated visual content.
[0064] It should be noted that, in actual applications, the quality level of the assessed visual content may also be determined based on other content, and this disclosure does not limit this. For example, in some embodiments, the model updating system 120 may determine the quality level of the assessed visual content based on a similarity score between the assessed visual content and the tagged visual content corresponding to the assessment prompt word information. The similarity score may, for example, be determined based on a specific algorithm or using a specific scoring model. It will be understood that the higher the similarity score, the higher the quality level.
[0065] It should be understood that only a few factors that can be used to evaluate content quality are given here. These factors can be used alone or in any combination to determine the quality level of visual content. In actual applications, other factors can also be selected to evaluate content quality.
[0066] At block 330 , the model updating system 120 may determine that the guidance scale currently applied by the teacher model 130 is inappropriate in response to determining that the quality level of the assessed visual content is below a first quality threshold, and may update the guidance scale based on a predetermined scaling factor. At block 340 , the model updating system 120 may determine that the guidance scale currently applied by the teacher model 130 is appropriate in response to determining that the quality level of the assessed visual content is not below the first quality threshold, and may maintain the current guidance scale.
[0067] In summary, according to the embodiments of the present disclosure, the guidance ratio of the teacher model that incorporates CFG can be dynamically adjusted during the student model update process to adjust the impact of negative prompt word information on the visual content output by the teacher model. This can avoid the problem of the quality of the visual content output by the teacher model decreasing with increasing iterations. This can ensure the quality of the model output of the teacher model and help improve the update effect of the student model.
[0068] Figure 4 4 shows a flow chart of a method 400 for model updating according to some embodiments of the present disclosure. The method 400 may be implemented in Figure 1 The electronic device 110 will refer to Figure 1 The method 400 is described with reference to the environment 100 of FIG.
[0069] In box 410, the electronic device 110 updates the student model based on the difference between the first visual content output by the teacher model and the second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information and a guidance ratio, the guidance ratio indicates the degree of influence of the positive cue word information and the negative cue word information on the first visual content, and the student model is configured to generate the second visual content based on the positive cue word information.
[0070] In block 420 , the electronic device 110 utilizes the teacher model to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information during the updating process of the student model.
[0071] At block 430 , the electronic device 110 updates the guidance ratio of the teacher model in response to evaluating the quality level of the visual content as being below the first threshold quality.
[0072] In block 440 , the electronic device 110 performs updating of the student model based on the updated guidance ratio using the teacher model.
[0073] In some embodiments, the degree of influence of the negative prompt word information indicated by the updated guidance ratio on the visual content output by the teacher model is lower than the degree of influence of the negative prompt word information indicated by the guidance ratio before the update on the visual content output by the teacher model.
[0074] In some embodiments, the quality level is determined based on at least one of: a saturation score of the evaluated visual content, a content diversity score of the evaluated visual content, a content generation efficiency score of the evaluated visual content, a fineness score of the evaluated visual content, a similarity score between the evaluated visual content and the tagged visual content corresponding to the evaluation prompt word information, or an aesthetic quality score of the evaluated visual content.
[0075] In some embodiments, updating the guidance scale of the teacher model includes: updating the guidance scale based on a predetermined scaling factor to obtain a first guidance scale; and selecting an updated guidance scale from the first guidance scale and the predetermined guidance scale based on a comparison of the first guidance scale and the predetermined guidance scale.
[0076] In some embodiments, the training objective of the student model is determined based on a predetermined guidance ratio.
[0077] In some embodiments, during the update process of the student model, the teacher model is used to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information, including: during the update process of the student model, in response to the evaluation conditions for the guidance ratio being met, the teacher model is used to generate evaluation visual content based on the evaluation prompt word information.
[0078] In some embodiments, the evaluation condition indicates at least one of the following: the number of updates to the student model reaches a predetermined number interval, and the content quality of the visual content output by the updated student model reaches a second threshold quality.
[0079] In some embodiments, method 400 further includes: stopping updating the student model in response to a model update target being met, wherein the model update target is based on at least one of the following: the number of updates of the student model reaches a predetermined number, or the content quality of the visual content output by the updated student model reaches a third threshold quality.
[0080] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 5 FIG2 shows an exemplary structural block diagram of an apparatus 500 for model updating according to some embodiments of the present disclosure. The apparatus 500 may be implemented as or included in the electronic device 110. Each module / component in the apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0081] like Figure 5As shown, the device 500 includes a student model updating module 510, which is configured to update the student model based on the difference between the first visual content output by the teacher model and the second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information and a guidance ratio, the guidance ratio indicating the degree of influence of the positive cue word information and the negative cue word information on the first visual content, and the student model is configured to generate the second visual content based on the positive cue word information. The device 500 also includes an evaluation content generation module 520, which is configured to generate an evaluation visual content corresponding to the evaluation cue word information based on the evaluation cue word information using the teacher model during the update process of the student model. The device 500 also includes a guidance ratio updating module 530, which is configured to update the guidance ratio of the teacher model in response to the quality level of the evaluation visual content being lower than the first threshold quality. The device 500 also includes an update execution module 540, which is configured to execute the update of the student model based on the updated guidance ratio using the teacher model.
[0082] In some embodiments, the degree of influence of the negative prompt word information indicated by the updated guidance ratio on the visual content output by the teacher model is lower than the degree of influence of the negative prompt word information indicated by the guidance ratio before the update on the visual content output by the teacher model.
[0083] In some embodiments, the quality level is determined based on at least one of: a saturation score of the evaluated visual content, a content diversity score of the evaluated visual content, a content generation efficiency score of the evaluated visual content, a fineness score of the evaluated visual content, a similarity score between the evaluated visual content and the tagged visual content corresponding to the evaluation prompt word information, or an aesthetic quality score of the evaluated visual content.
[0084] In some embodiments, the guidance scale update module 530 is further configured to: update the guidance scale based on a predetermined scaling factor to obtain a first guidance scale; and select an updated guidance scale from the first guidance scale and the predetermined guidance scale based on a comparison of the first guidance scale and the predetermined guidance scale.
[0085] In some embodiments, the training objective of the student model is determined based on a predetermined guidance ratio.
[0086] In some embodiments, the evaluation content generation module 520 is further configured to: during the updating process of the student model, in response to the evaluation condition for the guidance ratio being satisfied, generate evaluation visual content based on the evaluation prompt word information using the teacher model.
[0087] In some embodiments, the evaluation condition indicates at least one of the following: the number of updates to the student model reaches a predetermined number interval, and the content quality of the visual content output by the updated student model reaches a second threshold quality.
[0088] In some embodiments, the device 500 also includes: a stop update module, configured to stop updating the student model in response to the model update target being met, wherein the model update target is based on at least one of the following: the number of updates of the student model reaches a predetermined number of times, or the content quality of the visual content output by the updated student model reaches a third threshold quality.
[0089] The units and / or modules included in the device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine executable instructions stored on a storage medium. In addition to or as an alternative to machine executable instructions, some or all of the units and / or modules in the device 500 can be implemented at least in part by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0090] It should be understood that one or more steps in the above method can be performed by a suitable electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may include, for example, Figure 1 The electronic device 110 in.
[0091] Figure 6 1 shows a block diagram of an electronic device 600 in which one or more embodiments of the present disclosure may be implemented. It should be understood that Figure 6 The illustrated electronic device 600 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein. Figure 6 The electronic device 600 shown can be used to implement Figure 1 electronic device 110 or Figure 5 device 500.
[0092] like Figure 6As shown, electronic device 600 is in the form of a general electronic device. Components of electronic device 600 may include, but are not limited to, one or more processing units or processors 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processor 610 may be a real or virtual processor and is capable of performing various processes according to programs stored in memory 620. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 600.
[0093] The electronic device 600 typically includes a plurality of computer storage media. Such media can be any accessible media that can be obtained by the electronic device 600, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 600.
[0094] The electronic device 600 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 6 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 620 may include a computer program product 625 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0095] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0096] The input device 650 may be one or more input devices, such as a mouse, keyboard, or trackball. The output device 660 may be one or more output devices, such as a display, a speaker, or a printer. The electronic device 600 may also communicate with one or more external devices (not shown) through the communication unit 640 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with the electronic device 600, or with any device that allows the electronic device 600 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0097] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0098] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0099] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0100] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0101] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some updated implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0102] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for model updating, comprising: updating the student model based on a difference between first visual content output by the teacher model and second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information, and a guidance ratio, the guidance ratio indicating a degree of influence of the positive cue word information and the negative cue word information on the first visual content, and the student model is configured to generate the second visual content based on the positive cue word information; During the updating process of the student model, the teacher model is used to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information; in response to the assessed visual content having a quality level below a first threshold quality, updating the guidance ratio of the teacher model; as well as Updating the student model is performed using the teacher model based on the updated guidance ratio.
2. The method according to claim 1, wherein the degree of influence of the negative prompt word information indicated by the updated guidance ratio on the visual content output by the teacher model is lower than the degree of influence of the negative prompt word information indicated by the guidance ratio before the update on the visual content output by the teacher model.
3. The method of claim 1 , wherein the quality level is determined based on at least one of: The saturation score of the assessed visual content, The content diversity score for evaluating visual content, The content generation efficiency score for evaluating visual content, The score for evaluating the fineness of visual content, The similarity score between the evaluation visual content and the label visual content corresponding to the evaluation prompt word information, or The described evaluation evaluates the aesthetic quality score of the visual content.
4. The method of claim 1 , wherein updating the guidance ratio of the teacher model comprises: updating the guidance scale based on a predetermined scaling factor to obtain a first guidance scale; as well as The updated guidance ratio is selected from the first guidance ratio and the predetermined guidance ratio based on a comparison of the first guidance ratio and the predetermined guidance ratio. The method according to claim 4 , wherein a training target of the student model is determined based on the predetermined guidance ratio.
6. The method according to claim 1, wherein during the updating process of the student model, using the teacher model to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information comprises: In the updating process of the student model, in response to the evaluation condition for the guidance ratio being satisfied, the evaluation visual content is generated based on the evaluation prompt word information using the teacher model.
7. The method according to claim 6, wherein the evaluation condition indicates at least one of the following: The number of updates to the student model reaches a predetermined number of intervals, The content quality of the visual content output by the updated student model reaches a second threshold quality.
8. The method according to claim 1, further comprising: In response to a model update target being met, stopping updating the student model, wherein the model update target is based on at least one of the following: the number of updates of the student model reaches a predetermined number, or the content quality of the visual content output by the updated student model reaches a third threshold quality.
9. A device for model updating, comprising: a student model updating module configured to update the student model based on a difference between a first visual content output by a teacher model and a second visual content output by the student model, wherein the teacher model is configured to generate the first visual content based on positive cue word information, negative cue word information, and a guidance ratio, the guidance ratio indicating a degree of influence of the positive cue word information and the negative cue word information on the first visual content, and the student model is configured to generate the second visual content based on the positive cue word information; an evaluation content generating module configured to generate evaluation visual content corresponding to the evaluation prompt word information based on the evaluation prompt word information by using the teacher model during the updating process of the student model; a guidance ratio updating module configured to update the guidance ratio of the teacher model in response to the quality level of the assessed visual content being below a first threshold quality; as well as An update execution module is configured to execute an update of the student model based on the updated guidance ratio using the teacher model.
10. An electronic device comprising: at least one processor; as well as At least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 8 when executed by the at least one processor.
11. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.
12. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.