Methods, apparatuses, devices, and media for updating a decoder model

By adjusting the reference features of the decoder model and combining noise data and color templates, the structure of the decoder model is optimized, solving the problems of poor performance and high resource consumption in the existing technology, and achieving more efficient image processing results.

CN122137973APending Publication Date: 2026-06-02BEIJING ZITIAO NETWORK TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-12-02
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing decoder models have poor performance and accuracy, and their complex structure and large number of parameters lead to excessive resource consumption during the decoding process.

Method used

By acquiring the first reference feature of the reference image, adjusting the second reference feature of the reference image using noise data and color templates, generating a predicted reference image by combining the decoder model, updating the decoder model based on the differences, and then performing pruning to optimize the model.

Benefits of technology

It improves the performance of the decoder model, reduces the time and computational resource overhead of the decoding process, and enhances the saturation and clarity of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122137973A_ABST
    Figure CN122137973A_ABST
Patent Text Reader

Abstract

Methods, apparatus, devices, and media for updating a decoder model are provided. In one method, a reference image is acquired, having first reference features determined by an encoder model. Based on noisy data and the reference image, second reference features of the reference image are determined, the second reference features being different from the first reference features. A predicted reference image associated with the reference image is determined by a decoder model corresponding to the encoder model, based on the second reference features. The decoder model is updated based on the difference between the reference image and the predicted reference image. Using some implementations of this disclosure, adding noise information to the second reference features can reduce the quality of the second reference features. In this way, the decoder model can be updated in a direction that enables the decoder model to generate higher-quality images from lower-quality features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Implementations of this disclosure generally relate to machine learning, and in particular to methods, apparatus, devices, and computer-readable storage media for updating decoder models. Background Technology

[0002] Machine learning techniques have been widely used to handle various types of tasks. For example, encoders can be used to transform raw data from the original data space to a latent representation, and data transformations can be performed in the latent representation space to determine the target representation. Furthermore, decoder models can be used to transform the target representation back into the original data space to determine the desired data output. However, the performance and accuracy of existing decoder models are not satisfactory, and the complex structure and large number of parameters of decoder models can lead to significant resource consumption during the decoding process. Therefore, it is desirable to optimize decoder models to improve their performance and reduce the time and computational overhead of the decoding process. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for updating a decoder model is provided. In this method, a reference image is acquired, having a first reference feature determined by an encoder model. Based on noisy data and the reference image, a second reference feature of the reference image is determined, the second reference feature being different from the first reference feature. A predicted reference image associated with the reference image is determined by a decoder model corresponding to the encoder model, based on the second reference feature. The decoder model is updated based on the difference between the reference image and the predicted reference image.

[0004] In a second aspect of this disclosure, an apparatus for updating a decoder model is provided. The apparatus includes: an acquisition module configured to acquire a reference image having a first reference feature determined by an encoder model; a determination module configured to determine a second reference feature of the reference image based on noise data and the reference image, the second reference feature being different from the first reference feature; a prediction module configured to determine a predicted reference image associated with the reference image based on the second reference feature using a decoder model corresponding to the encoder model; and an update module configured to update the decoder model based on the difference between the reference image and the predicted reference image.

[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0007] In a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.

[0008] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent in the following detailed description, taken in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0010] Figure 1 A block diagram of an application environment according to one implementation of this disclosure is shown;

[0011] Figure 2 A block diagram of an update decoder model according to some implementations of this disclosure is shown;

[0012] Figure 3 A block diagram illustrating the process of determining a second reference feature according to some implementations of this disclosure is shown;

[0013] Figure 4 A block diagram illustrating the process of determining a second reference feature according to some implementations of this disclosure is shown;

[0014] Figure 5 A block diagram is shown illustrating the use of color templates to perform color degradation according to some implementations of this disclosure;

[0015] Figure 6 A block diagram is shown illustrating the selection of a color template from a plurality of predetermined color templates according to some implementations of this disclosure;

[0016] Figure 7 A block diagram is shown illustrating a pixel region-based method for determining a set of pixels according to some implementations of this disclosure;

[0017] Figure 8 A block diagram of an update decoder model according to some implementations of this disclosure is shown;

[0018] Figure 9A block diagram showing the output effect of the decoder model before and after optimization according to some implementations of this disclosure;

[0019] Figure 10 A flowchart is shown of a method for updating a decoder model according to some implementations of this disclosure;

[0020] Figure 11 A block diagram of an apparatus for updating a decoder model according to some implementations of this disclosure is shown; and

[0021] Figure 12 A block diagram of a device capable of implementing various implementations of the present disclosure is shown. Detailed Implementation

[0022] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0023] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.

[0024] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0025] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0026] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0027] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0028] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0029] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.

[0030] Example Environment

[0031] Machine learning techniques have been widely used to handle a variety of tasks. For example, encoders can be used to transform raw data from a raw data space to a latent representation, and data transformations can be performed in the space of the latent representation to determine the target representation. See also Figure 1 This describes how machine learning models work. Figure 1FIG. 100 is a block diagram showing an application environment according to an implementation of the present disclosure. In the context of the present disclosure, specific details of various embodiments will be described by taking a variational autoencoder (VAE) and a corresponding decoder as an example. A variational autoencoder is a generative model that probabilistically models the latent representation of data and generates new data similar to the training data. VAE combines the concepts of deep neural networks and Bayesian inference. The main idea of VAE is to assume that there is a latent variable that can generate the observed data, and new data can be generated by learning the distribution of this latent variable. The core of VAE is an encoder model and a decoder model.

[0032] As Figure 1 shown, the encoder 110 can be denoted as E, and the decoder 120 can be denoted as D. In the original data space R N (N is a positive integer representing the dimension of the original data space), there is original data X. The encoder 110 can be used to encode the original data to generate encoded data E(X) in an encoded data space (denoted as R M , M is a positive integer, M < N and represents the dimension of the original data space). The decoder 120 can be used to decode the encoded data to generate decoded data D(E(X)). At this time, the data space of D(E(X)) is the same as the original data space and is denoted as R N . Taking an image processing task as an example, R N can represent an image space, and R M can represent a vector space, also called an embedding space.

[0033] Currently, various encoder models and decoder models have been proposed. However, the performance and accuracy of existing decoder models are not satisfactory. For example, the visual effect of the images output by the decoder model is poor, such as the saturation and contrast may be low, etc. Further, the complex structure and a large number of parameters of the decoder model may cause the decoding process to consume a large amount of resources. At this time, it is desirable to optimize the decoder model to improve the performance of the decoder model and reduce the time overhead and computational resource overhead in the decoding process.

[0034] Summary of updating decoder model

[0035] To at least partially address the deficiencies in the prior art, according to an implementation of the present disclosure, a method for updating a decoder model is proposed. The technical solution of the present disclosure can be executed in the context of image generation and image processing. Refer to Figure 2 for a summary of an implementation of the present disclosure. The Figure 2A block diagram 200 for updating a decoder model is shown, according to some implementations of this disclosure. For example... Figure 2 As shown, a reference image (also called a training image) 230 can be obtained, which has a first reference feature 210 determined by the encoder model 250. In the context of this disclosure, the focus is primarily on optimizing the decoder model 260, and the parameters of the encoder model 250 can be fixed.

[0036] To improve the performance of the decoder model, a second reference feature 220, different from the first reference feature 210, can be determined for the reference image 230 based on the noisy data 240 and the reference image 230. In this case, the second reference feature 220 may include noise information, and therefore its quality is lower than that of the first reference feature 210. Further, a predicted reference image 232 associated with the reference image 230 can be determined by the decoder model 220 corresponding to the encoder model 210 based on the second reference feature 220. The decoder model 260 can be updated based on the difference 234 between the reference image 230 and the predicted reference image 232. According to some implementations of this disclosure, pruning can be further performed on the updated decoder model to reduce various resource overheads of the decoder model.

[0037] By utilizing some implementations of this disclosure, the quality of the second reference feature 220 can be reduced by adding noise information. In this way, the decoder model can be updated to improve its performance, thereby enabling it to generate higher-quality images from lower-quality features.

[0038] Detailed process of updating the decoder model

[0039] Having outlined some implementations according to this disclosure, further details of updating the decoder model will be described below. According to some implementations of this disclosure, the noise data may include adjustment parameters for adjusting the first reference feature. In this case, during the process of determining a second reference feature of the reference image based on the noise data and the reference image, the adjustment parameters can be applied to various dimensions of the first reference feature to determine the second reference feature. See also Figure 3 To describe more details, the Figure 3 A block diagram 300 illustrates a process for determining a second reference feature according to some implementations of this disclosure.

[0040] like Figure 3As shown, noise data 240 can be added to the first reference feature 210. Initially, the reference image 230 can be encoded using the encoder model 250 to determine the first reference feature 210. Noise data 240 can be added to the first reference feature 210 to obtain the second reference feature 220. Adjustment parameters may include, for example, […]. Figure 3 The degradation parameter 310 is used to reduce the quality of the first reference feature 210, that is, to make the second reference feature 220 deviate from the original first reference feature 210 to a certain extent. The degradation parameter can include various types, such as increasing (or decreasing) the value of one or more dimensions of the first reference feature 210 (e.g., offset parameter Δ). Assuming the first reference feature 210 is represented as f1 and the second reference feature 220 as f2, then f2 = f1 + Δ. In this way, the second reference feature 220 can be appropriately deviated from the original first reference feature 210, and the quality of the features input to the decoder model can be reduced. In this way, the second reference feature 220 includes both relevant information from the reference image 230 and introduced noise, thus guiding the model to learn the degraded parts of the image in a more aggressive manner, thereby improving the image's saturation and contrast.

[0041] Alternatively and / or additionally, the adjustment parameters may include, for example, scaling parameters for adjusting the first reference feature. In applying the adjustment parameters to the dimensions of the first reference feature to determine the second reference feature, the second reference feature can be determined based on the product of the scaling parameters and the dimensions of the first reference feature. Assuming the scaling parameter is denoted as α, then f2 = f1·α. α can be set to a small positive number, for example, in the range (0, 0.15). Alternatively and / or additionally, other ranges of values ​​can be set. In this way, the second reference feature 220 can be reduced and deviated from the original first reference feature 210, thereby reducing the quality of the features to be input into the decoder model. It should be understood that the above is only an example of a formula for reducing the quality of the first reference feature; alternatively and / or additionally, other formulas can be used to determine the second reference feature.

[0042] According to some implementations of this disclosure, noisy data can be used with a predetermined probability, for example, during training, noisy data can be used with a probability of 10%, 15% (or other values). The techniques described above can be executed in multiple training epochs. For example, in one epoch, noisy data can be added to the features of one or more images in a batch of training images with a predetermined probability. As another example, in different epochs, noisy data can be added to the features of all training images in a certain epoch with a predetermined probability, and so on. Using some implementations of this disclosure, the saturation and sharpness of the output image of the decoder model can be improved.

[0043] It should be understood that, despite Figure 3 and Figure 4 The process of modifying the first reference feature is illustrated. Alternatively and / or additionally, the encoder model 250 can be adjusted, and the aforementioned noise addition process can be integrated into the encoder model 250. At this point, the encoder model 250 can directly output the second reference feature 220. By introducing noise into the second reference feature 220, the color quality of the training samples can be reduced. After inputting the reduced-quality features into the decoder model, a reduced-quality predicted reference image can be determined. Furthermore, by calculating the loss between this reduced-quality predicted reference image and a normal-quality reference image (i.e., the original ground truth), the model can be guided to learn the degraded portions of the image in a more aggressive manner, thereby improving the image's saturation and contrast.

[0044] According to some implementations of this disclosure, noise information can be introduced into the second reference feature 220 by adjusting the reference image. In this case, the noise data may include a color template used to adjust the color of the reference image. During the process of determining the second reference feature of the reference image based on the noise data and the reference image, the reference image can be updated using the color template; and the updated reference image can be processed using an encoder model to determine the second reference feature. See also Figure 4 To describe more details, the Figure 4 A block diagram 400 illustrates the process of determining a second reference feature according to some implementations of this disclosure.

[0045] like Figure 4 As shown, noise parameter 240 may include color template 410, and the color template can be used to adjust the color of reference image 230 to obtain an adjusted reference image. Encoder model 250 can be used to perform the encoding process to determine second reference feature 220. Using some implementations of this disclosure, noise information can be introduced by adjusting the color of reference image 230, thereby reducing the quality of the second reference feature. Furthermore, in later update processes, the decoder model can be updated in a way that enables the decoder model to generate higher-quality images from lower-quality features.

[0046] In this way, the performance of the decoder model can be improved. Specifically, modifying the image color and reducing the image quality can incentivize the decoder model to learn the degraded parts of the image more aggressively, without needing to adjust the encoder model; instead, only the color of the reference image itself is degraded. See also Figure 5 To describe more details, the Figure 5A block diagram 500 illustrating color degradation using a color template according to some implementations of this disclosure is shown. A color template 530 (e.g., a green color template) can be overlaid on the original reference image 510 to degrade the original reference image 510 and obtain a degraded reference image 520. Specifically, brightness, saturation, contrast, etc., can be reduced. The degraded reference image 520 can then be input to an encoder model and a decoder model to determine a prediction image. Further, the prediction image can be compared with the initial ground truth to calculate the loss, enabling the decoder model to learn the color information before degradation more aggressively.

[0047] Based on some implementations of this disclosure, multiple predefined color templates can be provided, see [link / reference]. Figure 6 Provide more information. Figure 6 A block diagram 600 illustrates the selection of a color template from a plurality of predetermined color templates according to some implementations of this disclosure. For example... Figure 6 As shown, multiple color templates 620, 622, ..., 530, etc., can be provided. Here, the number of color templates and the specific colors can be predefined; for example, red, green, and blue color templates can be provided. Alternatively and / or additionally, red, orange, yellow, green, cyan, blue, and purple color templates can be provided, etc.

[0048] According to some implementations of this disclosure, the decoder model can be updated iteratively across multiple rounds. To utilize color templates to update the reference image, a color template can be selected from multiple predetermined color templates in one of the multiple rounds. In that round, the color template can be overlaid onto the reference image to update it. For example, color template 620 can be applied to reference image 230, color template 622 can be applied to reference image 610, color template 530 can be applied to reference image 612, color template 620 can be applied to reference image 614, and so on. According to some implementations of this disclosure, the color template can be used with a predetermined probability, for example, during training, the color template can be used with a probability of 10%, 15% (or other values). Utilizing some implementations of this disclosure, the saturation and sharpness of the decoder model's output image can be improved.

[0049] It should be understood that a many-to-many relationship can exist between the reference image and the color template. For example, multiple different color templates can be applied to a single reference image to generate multiple different training images. Alternatively and / or additionally, the same color template can be applied to multiple different reference images to generate multiple different training images. The training process can then be performed using these multiple different training images so that the decoder model can learn multiple color information before degradation.

[0050] According to some implementations of this disclosure, in determining the loss used to update the decoder model, the loss can be determined based on the difference between all pixels in the reference image and all pixels in the predicted reference image. Alternatively and / or additionally, a subset of pixels can be filtered from the reference image, and the loss can be determined only based on the difference between this subset of pixels; this method can be referred to as a pixel filtering method. Specifically, in updating the decoder model based on the difference between the reference image and the predicted reference image, multiple pixel losses for multiple pixels in the reference image can be determined based on the difference between the reference image and the predicted reference image; an image loss corresponding to the reference image can be determined based on the multiple pixel losses; and the decoder model can be updated based on the image loss.

[0051] See Figure 7 To describe more details, the Figure 7 A block diagram 700 is shown illustrating a pixel region-based method for determining a set of pixels according to some implementations of this disclosure. For example... Figure 7 As shown, assuming the width and height of the reference image 710 are W and H respectively, the reference image 710 includes W*H pixels. The implicit features of the reference image can be determined using the encoder model, and these implicit features are input into the decoder model to determine the corresponding predicted reference image. In this case, the predicted reference image also includes W*H pixels. For each pixel position (i,j), the pixel loss between the reference image and the predicted reference image can be determined, that is, the difference D between the two corresponding pixel values. i,j Furthermore, the determined individual Ds can be... i,j Sorting (e.g., in descending order). A predetermined number or proportion of pixels can be determined from the sorted pixels, and the correlation D of the set of pixels can be used. i,j To determine the loss (i.e., image loss) used to update the decoder model.

[0052] According to some implementations of this disclosure, in the process of determining the image loss corresponding to the reference image based on multiple pixel losses, a set of pixels can be selected from multiple pixels based on the ranking of the multiple pixel losses; and the image loss can be determined based on a set of pixels and the prediction reference image. Generally speaking, models have a poor ability to learn complex image content, and therefore their prediction ability for parts of complex image content is also poor. For example... Figure 7As shown, the correlated pixel loss of pixels within pixel region 720 is typically large, thus placing them near the top of the sorting hierarchy. For other regions outside pixel region 720, since the image content in these regions is relatively simple and the predictions made by the decoder model are usually quite accurate, the image loss can be determined solely based on the correlated loss of pixels within pixel region 720. Using some implementations of this disclosure, the decoder model can be updated to improve its ability to predict complex image content.

[0053] According to some implementations of this disclosure, in the process of determining image loss based on a set of pixels and a prediction reference image, for each pixel in the set of pixels, the image loss can be determined based on the difference between the predicted pixel value of the corresponding pixel in the prediction reference image and the predicted pixel value of the same pixel. Specifically, the predicted pixel value of each pixel within the pixel region 720 can be compared with the true value and the difference can be determined. Then, the image loss can be determined based on the summation of the differences corresponding to each pixel within the pixel region 720. Utilizing some implementations of this disclosure can improve the performance of the decoder model in outputting complex image content.

[0054] According to some implementations of this disclosure, pixels corresponding to higher pixel losses can be preferentially selected; that is, a group of pixels can be determined from the head position of the sorting. In this case, for the first pixel in the group and the second pixel outside the group, the first pixel loss corresponding to the first pixel is higher than the second pixel loss corresponding to the second pixel. In this way, the decoder model can focus more on improving the quality of the output complex image content.

[0055] According to some implementations of this disclosure, the ratio between a set of pixels and all pixels can be determined in various ways. For example, the ratio can be set to a predetermined value (e.g., 50%, 60%, or other values). Alternatively and / or additionally, the ratio can be determined based on the complexity of the reference image; the greater the proportion of complex parts included in the reference image, the higher the ratio. For example, when the reference image includes a foreground image and a background image, the ratio can be determined based on the proportion of the foreground image in the reference image. Using some implementations of this disclosure, the specific formula for updating the decoder model can be adaptively adjusted based on the complexity of the reference image, thereby making the predicted image output by the updated decoder model more closely match the ground truth.

[0056] According to some implementations of this disclosure, the decoder model can be trained progressively over multiple rounds. For example, in the T-th round of training, the loss can be determined based on the prediction in the (T-1)-th round, and the positions of a set of pixels with a loss ranking in the top 50% can be found. Data for the set of pixels corresponding to these positions can be retained, and data for other pixels can be set to predetermined values ​​(e.g., 0, 1, or other values) to update the decoder model. Alternatively and / or additionally, the above process can be performed according to predetermined probabilities; for example, the pixel filtering process can be performed every 3 rounds during training. Using some implementations of this disclosure, the saturation and sharpness of the decoder model's output image can be improved.

[0057] The preceding sections have described several steps for updating the decoder model. Alternatively and / or additionally, pruning can be performed on the network structure in the updated decoder model. It should be understood that pruning a machine learning model aims to reduce its complexity by decreasing redundant parameters and connections, thereby improving its efficiency and accuracy. For example, weight pruning can be performed to remove individual weights with absolute values ​​less than a certain threshold, significantly reducing the number of model parameters. Alternatively and / or additionally, structured pruning can be performed to remove convolutional kernels, neurons, or channels from the model, resulting in a more regular sparse structure, which is beneficial for hardware acceleration.

[0058] According to some implementations of this disclosure, pruning can be performed with predetermined probabilities across multiple training epochs. For example, pruning can be performed after each epoch, alternatively and / or additionally, after multiple epochs. Utilizing some implementations of this disclosure, the amount of data in the decoder model can be reduced, and correspondingly, the computational and time resource overhead of running the decoder model can be reduced.

[0059] According to some implementations of this disclosure, the decoder model described above can be run on various computing devices, and the decision to perform pruning can be based on the type of computing device. For decoder models running on client devices, pruning can be performed on these models due to the limited computing power of the client devices. For decoder models running on server devices, pruning is unnecessary because the computing power of server devices is typically stronger. In this way, suitable decoder models can be provided according to the hardware capabilities of different computing devices.

[0060] According to some implementations of this disclosure, a discriminator can be set based on a Generative Adversarial Network (GAN) and used to evaluate the quality of the output image of the decoder model. According to some implementations of this disclosure, a replay caching technique can be used. In GAN training, replay caching helps maintain the stability of model training. Specifically, a fixed-capacity cache can be provided, and past training samples can be stored in this cache and the decoder model can be iteratively updated continuously. The discriminator not only judges the performance of the current training samples but also the performance of past training samples in the cache. In this way, the convergence process of the discriminator can be prevented from being too volatile, and the decoder model can be prevented from fluctuating too much, thereby optimizing the decoder model in a more stable manner.

[0061] During pruning, VAEs can undergo multiple structured pruning operations. Specifically, unnecessary parameters can be physically removed based on L1 loss, while other dependent layers are automatically pruned. After each pruning, the pruned model can be optimized using GAN training – a process known as re-pruning (progressive pruning). This process can be repeated multiple times. In one example, three pruning operations can be performed, with pruning rates of 10%, 15%, and 15% respectively. The number of model parameters can be significantly reduced after pruning; for example, the final result can reduce the amount of data by approximately 50%.

[0062] The specific implementation of each step in updating the decoder model has been described, and see below. Figure 8 Describe the overall workflow. Figure 8 A block diagram 800 for updating a decoder model is shown, according to some implementations of this disclosure. For example... Figure 8 As shown, multiple reference images 230 can be acquired, and the method described above can be executed in multiple rounds. It should be understood that the individual steps of the above method can be executed according to predetermined probabilities. Figure 8 As shown in box 810, multiple fine-tuning steps can be used with different probabilities in each training epoch, such as using downgrade parameters 310, color targets 410, pixel filtering 820, replay buffer 840, etc. The replay buffer 840 is used to enable the discriminator 842 to simultaneously determine the effect of the decoder model 260's output image, as well as the related effect of previous samples in the replay buffer 840. Furthermore, pruning 830 can be performed.

[0063] Based on some implementations of this disclosure, a generative adversarial network can be applied to the VAE decoder model after each pruning during model fine-tuning. In other words, a trainable discriminator network can be introduced to evaluate the performance of the VAE decoder model in real time. Regarding the learning rate, a cosine annealing learning rate decay method can be used for training, meaning that the more training epochs, the smaller the model update amplitude to maintain model stability. Currently available public datasets can be used to perform the training process, and data degradation can be performed on the training data with a certain probability. Assuming the original training samples are H0, the following process can be performed.

[0064] Step 1: Degrade the entire training sample color with a 10% probability (or other value), obtaining training sample H00 or H01 (where H00 is the color-rich sample and H01 is the degraded sample). Step 2: Filter the sample (H0 or H01) obtained in Step 1 pixel-wise with a 25% probability (or other value), obtaining sample H0, H01, H02, or H012. Step 3: Degrade the sample obtained in Step 2 using latent features, obtaining H03, H013, H023, or H0123. These samples can be input into the decoder model. Since there are four possible inputs, there are four possible outputs: R03, R013, R023, or R0123.

[0065] Step 4: When the number of training epochs is less than 5000, only the discriminator can be trained and a replay cache can be applied; when the number of training epochs is 7N (or other values), the VAE decoder is trained; when the number of training epochs is 7N+1 (or other values), the VAE encoder is trained; in other cases, the discriminator can be trained and a replay cache can be applied. Further, loss calculation can be performed on the above samples, and the loss calculation can be based on various loss functions. Specifically, the loss between R03-H0 can provide implicit feature degradation; the loss between R013-H00 can provide color enhancement; the loss between R023-H02 can provide implicit feature degradation and pixel filtering; and the loss between R0123-H012 can provide color enhancement, latent variable degradation, and pixel filtering. Step 5: The gradient determined by the loss can be accumulated until gradient update is performed after 2 epochs. In this way, the batch size can be physically increased. Step 6: New samples can be introduced, and the loop can start from step 1.

[0066] According to some implementations of this disclosure, the methods described above can be used to update the decoder model during the execution of various image processing tasks. The appropriate decoder model can be selected based on the specific task type. Furthermore, these image processing tasks may include, but are not limited to, image generation tasks, image transformation tasks, and so on.

[0067] The methods described above can be applied to a known decoder model to obtain an optimized decoder model. In one example, for a decoder running on a server, the number of parameters in the unoptimized model was approximately 80M, and the number of parameters in the optimized model was reduced to approximately 60M. The techniques described above can be used to optimize the decoder model, resulting in a reduction of approximately 60% in inference time and approximately 40% in memory usage.

[0068] See Figure 9 Describe the output effect of the decoder model before and after optimization. Figure 9 A block diagram 900 illustrates the output effects of a decoder model before and after optimization according to some implementations of this disclosure. Image 910 represents the image output by the decoder model before optimization, and image 920 represents the image output by the decoder model after optimization. By comparison, it can be found that image 920 has higher saturation and contrast, and provides better visual effects while reducing various resource overheads.

[0069] Within the context of this disclosure, a complete set of methods for lightweight VAE model fine-tuning and pruning is proposed. Utilizing some implementations of this disclosure, the quality of the second reference features can be reduced by adding noise information. In this way, the decoder model can be updated in a manner that enables it to generate higher-quality images from lower-quality features, thereby improving the performance of the decoder model.

[0070] Example process

[0071] Figure 10 A flowchart of a method 1000 for updating a decoder model according to some implementations of this disclosure is shown. At block 1010, a reference image is acquired, having a first reference feature determined by an encoder model. At block 1020, a second reference feature of the reference image is determined based on noise data and the reference image, the second reference feature being different from the first reference feature. At block 1030, a predicted reference image associated with the reference image is determined by a decoder model corresponding to the encoder model based on the second reference feature. At block 1040, the decoder model is updated based on the difference between the reference image and the predicted reference image.

[0072] According to some implementations of this disclosure, the noise data is an adjustment parameter used to adjust the first reference feature, and determining the second reference feature of the reference image based on the noise data and the reference image includes: applying the adjustment parameter to each dimension of the first reference feature to determine the second reference feature.

[0073] According to some implementations of this disclosure, the adjustment parameter is a scaling parameter used to adjust the first reference feature, and applying the adjustment parameter to each dimension of the first reference feature to determine the second reference feature includes: determining the second reference feature based on the product of the scaling parameter and each dimension of the first reference feature.

[0074] According to some implementations of this disclosure, the noise data is a color template for adjusting the color of a reference image, and determining a second reference feature of the reference image based on the noise data and the reference image includes: updating the reference image using the color template; and processing the updated reference image using an encoder model to determine the second reference feature.

[0075] According to some implementations of this disclosure, updating a reference image using a color template includes: selecting a color template from a plurality of predetermined color templates in one of a plurality of rounds; and overlaying the color template onto the reference image in the round to update the reference image.

[0076] According to some implementations of this disclosure, updating the decoder model based on the difference between a reference image and a predicted reference image includes: determining a multiple pixel loss for multiple pixels in the reference image based on the difference between the reference image and the predicted reference image; determining an image loss corresponding to the reference image based on the multiple pixel loss; and updating the decoder model based on the image loss.

[0077] According to some implementations of this disclosure, determining the image loss corresponding to a reference image based on multiple pixel losses includes: selecting a set of pixels from multiple pixels based on the sorting of multiple pixel losses; and determining the image loss based on the set of pixels and the predicted reference image.

[0078] According to some implementations of this disclosure, for a first pixel in a set of pixels and a second pixel outside the set of pixels, the loss of the first pixel corresponding to the first pixel is higher than the loss of the second pixel corresponding to the second pixel.

[0079] According to some implementations of this disclosure, determining image loss based on a set of pixels and a prediction reference image includes: for a pixel in the set of pixels, determining image loss based on the difference between the pixel and the prediction pixel corresponding to the pixel in the prediction reference image.

[0080] According to some implementations of this disclosure, the method is executed in multiple rounds and the method is executed according to a predetermined probability.

[0081] According to some implementations of this disclosure, the method further includes performing pruning on the network structure in the updated decoder model.

[0082] Example devices and equipment

[0083] Figure 11 A block diagram 1100 of an apparatus 800 for updating a decoder model according to some implementations of the present disclosure is shown. The apparatus includes: an acquisition module 1110 configured to acquire a reference image having a first reference feature determined by an encoder model; a determination module 1120 configured to determine a second reference feature of the reference image based on noise data and the reference image, the second reference feature being different from the first reference feature; a prediction module 1130 configured to determine a predicted reference image associated with the reference image based on the second reference feature using a decoder model corresponding to the encoder model; and an update module 1140 configured to update the decoder model based on the difference between the reference image and the predicted reference image.

[0084] According to some implementations of this disclosure, the noise data is an adjustment parameter used to adjust the first reference feature, and determining the second reference feature of the reference image based on the noise data and the reference image includes: applying the adjustment parameter to each dimension of the first reference feature to determine the second reference feature.

[0085] According to some implementations of this disclosure, the adjustment parameter is a scaling parameter used to adjust the first reference feature, and applying the adjustment parameter to each dimension of the first reference feature to determine the second reference feature includes: determining the second reference feature based on the product of the scaling parameter and each dimension of the first reference feature.

[0086] According to some implementations of this disclosure, the noise data is a color template for adjusting the color of a reference image, and determining a second reference feature of the reference image based on the noise data and the reference image includes: updating the reference image using the color template; and processing the updated reference image using an encoder model to determine the second reference feature.

[0087] According to some implementations of this disclosure, updating a reference image using a color template includes: selecting a color template from a plurality of predetermined color templates in one of a plurality of rounds; and overlaying the color template onto the reference image in the round to update the reference image.

[0088] According to some implementations of this disclosure, updating the decoder model based on the difference between a reference image and a predicted reference image includes: determining a multiple pixel loss for multiple pixels in the reference image based on the difference between the reference image and the predicted reference image; determining an image loss corresponding to the reference image based on the multiple pixel loss; and updating the decoder model based on the image loss.

[0089] According to some implementations of this disclosure, determining the image loss corresponding to a reference image based on multiple pixel losses includes: selecting a set of pixels from multiple pixels based on the sorting of multiple pixel losses; and determining the image loss based on the set of pixels and the predicted reference image.

[0090] According to some implementations of this disclosure, for a first pixel in a set of pixels and a second pixel outside the set of pixels, the loss of the first pixel corresponding to the first pixel is higher than the loss of the second pixel corresponding to the second pixel.

[0091] According to some implementations of this disclosure, determining image loss based on a set of pixels and a prediction reference image includes: for a pixel in the set of pixels, determining image loss based on the difference between the pixel and the prediction pixel corresponding to the pixel in the prediction reference image.

[0092] According to some implementations of this disclosure, the method is executed in multiple rounds and the method is executed according to a predetermined probability.

[0093] According to some implementations of this disclosure, the apparatus further includes a pruning module configured to perform pruning processing on the network structure in the updated decoder model.

[0094] Figure 12 A block diagram of a device 1200 capable of implementing various implementations of the present disclosure is shown. It should be understood that... Figure 12 The computing device 1200 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementation described herein. Figure 12 The computing device 1200 shown can be used to implement the method described above.

[0095] like Figure 12 As shown, computing device 1200 is in the form of a general-purpose computing device. Components of computing device 1200 may include, but are not limited to, one or more processors or processing units 1210, memory 1220, storage devices 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260. Processing unit 1210 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1200.

[0096] Computing device 1200 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1230 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1200.

[0097] The computing device 1200 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 12 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1220 may include computer program product 1225 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.

[0098] The communication unit 1240 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1200 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1200 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0099] Input device 1250 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1260 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1200 can also communicate with one or more external devices (not shown) via communication unit 1240 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1200, or with any device (e.g., network card, modem, etc.) that enables computing device 1200 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).

[0100] According to an implementation of this disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the method described above. According to an implementation of this disclosure, a computer program product is provided, on which a computer program is stored, which, when executed by a processor, implements the method described above.

[0101] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0103] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0105] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for updating a decoder model, comprising: Acquire a reference image, the reference image having a first reference feature determined by an encoder model; Based on the noise data and the reference image, a second reference feature of the reference image is determined, the second reference feature being different from the first reference feature; A predicted reference image associated with the reference image is determined by a decoder model corresponding to the encoder model, based on the second reference feature; as well as The decoder model is updated based on the difference between the reference image and the predicted reference image.

2. The method of claim 1, wherein the noise data is an adjustment parameter for adjusting the first reference feature, and determining the second reference feature of the reference image based on the noise data and the reference image comprises: The adjustment parameters are applied to each dimension of the first reference feature to determine the second reference feature.

3. The method of claim 2, wherein the adjustment parameter is a scaling parameter for adjusting the first reference feature, and applying the adjustment parameter to each dimension of the first reference feature to determine the second reference feature comprises: The second reference feature is determined based on the product of the scaling parameter and each dimension of the first reference feature.

4. The method of claim 1, wherein the noise data is a color template for adjusting the color of the reference image, and determining the second reference feature of the reference image based on the noise data and the reference image comprises: The reference image is updated using the color template. as well as The updated reference image is processed using the encoder model to determine the second reference feature.

5. The method of claim 4, wherein updating the reference image using the color template comprises: In one of a series of rounds, the color template is selected from a series of predetermined color templates; as well as In the round, the color template is overlaid on the reference image to update the reference image.

6. The method of claim 1, wherein updating the decoder model based on the difference between the reference image and the predicted reference image comprises: Based on the difference between the reference image and the predicted reference image, a multiple pixel loss is determined for multiple pixels in the reference image; Based on the multiple pixel losses, determine the image loss corresponding to the reference image; as well as The decoder model is updated based on the image loss.

7. The method of claim 6, wherein determining the image loss corresponding to the reference image based on the plurality of pixel losses comprises: Based on the sorting of the multiple pixel losses, a group of pixels is selected from the multiple pixels; as well as The image loss is determined based on the set of pixels and the predicted reference image.

8. The method of claim 7, wherein for a first pixel in the set of pixels and a second pixel outside the set of pixels, the loss of the first pixel corresponding to the first pixel is higher than the loss of the second pixel corresponding to the second pixel.

9. The method of claim 7, wherein determining the image loss based on the set of pixels and the prediction reference image comprises: For each pixel in the set of pixels, the image loss is determined based on the difference between the pixel and the corresponding predicted pixel in the prediction reference image.

10. The method of claim 1, wherein the method is performed in a plurality of rounds and the method is performed according to a predetermined probability.

11. The method of claim 1, further comprising: Pruning is performed on the network structure in the updated decoder model.

12. An apparatus for updating a decoder model, comprising: An acquisition module is configured to acquire a reference image having a first reference feature determined by an encoder model; The determining module is configured to determine a second reference feature of the reference image based on noise data and the reference image, wherein the second reference feature is different from the first reference feature; A prediction module is configured to determine a predicted reference image associated with the reference image based on the second reference features, using a decoder model corresponding to the encoder model. as well as An update module is configured to update the decoder model based on the difference between the reference image and the predicted reference image.

13. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 11.

14. A computer-readable storage medium having stored thereon computer instructions that, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 11.

15. A computer instruction product comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 11.