Method and apparatus for generating a model

By performing diffusion and component attenuation processing on the original image, the generative model training method solves the problem of low training efficiency for high-dimensional data, and achieves fast and efficient generative model training and high-precision generation.

CN115641485BActive Publication Date: 2026-08-04ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA (CHINA) CO LTD
Filing Date
2022-11-02
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing generative models are costly and difficult to train on high-dimensional data, with long training times and significant fitting challenges. This is especially true in image generation tasks, where the increased data dimensionality leads to low model training efficiency.

Method used

By performing diffusion processing on the original image and attenuating image components during the diffusion process, a noisy image set is generated and combined with the original image for training until a target generative model that meets the training stopping condition is obtained.

Benefits of technology

It enables fast and efficient generative model training, reduces the difficulty of model fitting, improves training speed, and enhances the accuracy of the generative model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641485B_ABST
    Figure CN115641485B_ABST
Patent Text Reader

Abstract

Embodiments of the present specification provide a generation model training method and device, wherein the generation model training method comprises: obtaining an original image; performing diffusion processing on the original image, and performing attenuation processing on an image component of the original image to obtain a set of noisy images; determining a noisy image according to the original image and the set of noisy images, and inputting the noisy image into an initial generation model for processing to obtain a restored image; and adjusting the initial generation model based on the restored image and the original image until a target generation model that satisfies a training stop condition is obtained. During the diffusion and inverse diffusion processing, the attenuation processing of the image component is combined, so that the generation model can learn the image change process in different dimensions, thereby effectively improving the model training precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of machine learning technology, and in particular to generative model training methods and apparatus. Background Technology

[0002] With the development of internet technology, diffusion processes have been widely used to build generative models. By teaching the generative model the inverse of the diffusion process, it can generate samples from a given data distribution from Gaussian noise. The diffusion process can be understood as an iterative process. In each step of the iteration, Gaussian noise is further added to the current noisy data. This iterative process ensures that the dimensionality of the noisy data remains unchanged before and after each addition, always maintaining the same dimensionality as the original data. Similarly, the inverse process of diffusion is also an iterative process, and its dimensionality remains unchanged during this process. In other words, in both the diffusion and its inverse processes, the dimensionality of the noisy data remains the same as the original data. When the data dimensionality is high, this leads to high model training costs and learning difficulties. Furthermore, due to this characteristic of maintaining data dimensionality, generative models based on diffusion and its inverse process have long training times and are difficult to fit on high-dimensional data. Therefore, an effective solution is urgently needed to address these issues. Summary of the Invention

[0003] In view of the above, embodiments of this specification provide a generative model training method. One or more embodiments of this specification also relate to a generative model training apparatus, an image processing method, an image processing apparatus, another image processing method, another image processing apparatus, another generative model training method, another generative model training apparatus, a computing device, a computer-readable storage medium, and a computer program, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a generative model training method is provided, which, given a training dataset, independently hypothesizes and learns the relationship between input and output based on preset algorithm parameters, and then learns a model based on this relationship. For given input features, the model can process and output features that meet usage requirements. The method includes:

[0005] Obtain the original image;

[0006] The original image is subjected to diffusion processing, and the image components of the original image are subjected to attenuation processing to obtain a noisy image set;

[0007] The noisy image is determined based on the original image and the set of noisy images, and the noisy image is input into the initial generation model for processing to obtain the restored image;

[0008] The initial generative model is tuned based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0009] According to a second aspect of the embodiments of this specification, a generative model training apparatus is provided, comprising:

[0010] The acquisition module is configured to acquire the raw image;

[0011] The processing module is configured to perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set;

[0012] The determination module is configured to determine a noisy image based on the original image and the set of noisy images, and input the noisy image into the initial generation model for processing to obtain the restored image;

[0013] The training module is configured to tune the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0014] According to a third aspect of the embodiments of this specification, an image processing method is provided, comprising:

[0015] Acquire the image to be processed uploaded by the user terminal;

[0016] The image to be processed is input into the target generation model in the above method for processing to obtain the target image;

[0017] The target object is determined based on the target image, and the object information corresponding to the target object is loaded.

[0018] The object information is sent to the user terminal.

[0019] According to a fourth aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:

[0020] The image acquisition module is configured to acquire images to be processed uploaded by the user terminal;

[0021] The model processing module is configured to input the image to be processed into the target generation model in the above method for processing, so as to obtain the target image;

[0022] The object determination module is configured to determine the target object based on the target image and load the object information corresponding to the target object;

[0023] The information sending module is configured to send the object information to the user terminal.

[0024] According to a fifth aspect of the embodiments of this specification, another image processing method is provided, comprising:

[0025] Obtain the initial shopping search image uploaded by the user's terminal;

[0026] The initial shopping search image is input into the target generation model in the above method for processing to obtain the target shopping search image corresponding to the initial shopping search image;

[0027] Based on the target shopping search image, related products are determined, and the product information corresponding to the related products is loaded;

[0028] The product information is sent to the user terminal, whereby the user terminal generates and displays a product recommendation interface based on the product information.

[0029] According to a sixth aspect of the embodiments of this specification, another image processing apparatus is provided, comprising:

[0030] The image acquisition module is configured to acquire the initial shopping search image uploaded by the user terminal;

[0031] The input model module is configured to input the initial shopping search image into the target generation model in the above method for processing, so as to obtain the target shopping search image corresponding to the initial shopping search image;

[0032] The information loading module is configured to determine associated products based on the target shopping search image and load the product information corresponding to the associated products;

[0033] The information sending module is configured to send the product information to the user terminal, wherein the user terminal generates and displays a product recommendation interface based on the product information.

[0034] According to a seventh aspect of the embodiments of this specification, another generative model training method is provided, applied on a server, wherein the generative model is a machine learning model, comprising:

[0035] Receive the original images uploaded by the model requester;

[0036] The original image is subjected to diffusion processing, and the image components of the original image are subjected to attenuation processing to obtain a noisy image set;

[0037] The noisy image is determined based on the original image and the set of noisy images, and the noisy image is input into the initial generation model for processing to obtain the restored image;

[0038] The initial generative model is tuned based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0039] Determine the model parameters corresponding to the target generation model, and feed the model parameters back to the model demand side.

[0040] According to an eighth aspect of the embodiments of this specification, another generative model training apparatus is provided, applied on a server, wherein the generative model is a machine learning model, comprising:

[0041] The image receiving module is configured to receive raw images uploaded by the model requesting end;

[0042] The image processing module is configured to perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set;

[0043] The image determination module is configured to determine a noisy image based on the original image and the noisy image set, and input the noisy image into the initial generation model for processing to obtain the restored image;

[0044] The model training module is configured to tune the parameters of the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0045] The parameter sending module is configured to determine the model parameters corresponding to the target generated model and feed the model parameters back to the model demand side.

[0046] According to a ninth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0047] Memory and processor;

[0048] The memory is used to store computer-executable instructions, and the processor is used to implement the steps of the above method when executing the computer-executable instructions.

[0049] According to a tenth aspect of an embodiment of this specification, a computer-readable storage medium is provided that stores computer-executable instructions that, when executed by a processor, implement the steps of the method described above.

[0050] According to an eleventh aspect of the embodiments of this specification, a computer program is provided, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described method.

[0051] The generative model training method provided in this specification aims to achieve rapid and efficient model training while reducing model fitting difficulty. After acquiring the original image, a diffusion process is performed, simultaneously attenuating the image components of the original image during diffusion. This results in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original image and the noisy image set and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original image and the restored image until the target generative model that meets the training stopping condition is obtained. This method achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of a generative model training method provided in one embodiment of this specification;

[0053] Figure 2 This is a flowchart illustrating a generative model training method provided in one embodiment of this specification;

[0054] Figure 3 This is a schematic diagram of the structure of a generative model training device provided in one embodiment of this specification;

[0055] Figure 4 This is a flowchart illustrating an image processing method provided in one embodiment of this specification;

[0056] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification;

[0057] Figure 6 This is a flowchart of another image processing method provided in one embodiment of this specification;

[0058] Figure 7 This is a schematic diagram of the structure of another image processing apparatus provided in one embodiment of this specification;

[0059] Figure 8 This is a flowchart of another generative model training method provided in one embodiment of this specification;

[0060] Figure 9 This is a schematic diagram of another generative model training device provided in one embodiment of this specification;

[0061] Figure 10 This is a flowchart illustrating the processing steps of a generative model training method provided in one embodiment of this specification.

[0062] Figure 11 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0063] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0064] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0065] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0066] First, the terms and concepts used in one or more embodiments of this specification will be explained.

[0067] Generative models are models that can sample data from a given data distribution.

[0068] Diffusion process: A stochastic process that gradually adds noise to data and has an inverse process to recover the data distribution. It can be used to build generative models.

[0069] Machine learning models are models that, given a training dataset, independently learn the relationship between inputs and outputs based on predefined algorithm parameters. Based on this learned relationship, a model can process given input features and output features that meet specific usage requirements. For example, a probabilistic model, given a training dataset, can learn the probability distribution of inputs and outputs based on conditional assumptions about features. Based on this model, for a given input x, it can output y with the highest probability.

[0070] This specification provides a generative model training method. One or more embodiments of this specification also relate to a generative model training apparatus, an image processing method, an image processing apparatus, another image processing method, another image processing apparatus, another generative model training method, another generative model training apparatus, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail in the following embodiments.

[0071] In existing technologies, during the iterations of the diffusion and inverse processes, the dimension of the noisy data remains the same as the dimension of the original data. This results in the model's input and output dimensions being identical to the original data dimensions during training. As the data dimension increases, the training iteration speed and the model's fitting difficulty also increase. This is especially true for common image generation tasks, where the data dimension increases quadratically with image size, significantly increasing the difficulty of fitting large image datasets. Currently, to avoid the impact of maintaining the same dimension, an additional inverse process of the diffusion process with a lower dimension is trained. This diffusion process is a diffusion process of a low-dimensional principal component of the original data. When sampling using the inverse process, this low-dimensional process is iterated for a certain number of steps, then the low-dimensional noisy data is upsized and subjected to noise compensation. Finally, the inverse process with the original dimension is iterated again to generate the final sample. However, in this method, the distribution obtained after upsizing and noise compensation is not the same as the distribution of the original diffusion process at that iteration step, and the difference can be significant, potentially reducing the quality of the final generated sample. Due to this limitation, the method is also difficult to perform multiple dimensional changes and is not effective in reducing dimensionality for higher-dimensional data.

[0072] In view of this, see Figure 1 The schematic diagram illustrates the generative model training method provided in this embodiment. To achieve rapid and efficient model training while reducing model fitting difficulty, the original image undergoes diffusion processing after acquisition. Simultaneously, during diffusion, image components of the original image are attenuated to obtain a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original image and the noisy image set and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original image and the restored image until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0073] It should be noted that the user characteristic information or user data involved in this application are all information and data authorized by the user or fully authorized by all parties. The user characteristic information includes, but is not limited to, user personal information and user preference information. The user data includes, but is not limited to, data used for analysis, data stored, and data displayed, such as images. Furthermore, the collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0074] Figure 2 A flowchart of a generative model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0075] Step S202: Obtain the original image.

[0076] The generative model training method provided in this embodiment can be applied to image processing scenarios, that is, to train a generative model that can be used to process images. The trained generative model can be a generative model for image sharpness processing, or a generative model for image size adjustment, or a generative model for image color adjustment, etc. This embodiment takes training a generative model that can be used for image sharpness processing as an example to describe the training method of the generative model. For other scenarios with the same or corresponding descriptions, please refer to this embodiment. This embodiment will not elaborate further here.

[0077] Specifically, the original image refers to the unprocessed image. Processing the original image can generate sample pairs for training the generative model. In the sample pairs, the original image can be used as the label, and the noisy image obtained after processing the original image can be used as the sample.

[0078] Based on this, in order to improve training speed and reduce fitting difficulty when training a generative model for image processing, we can first obtain the original image without any processing. This makes it easier to use the original image as a basis for diffusion to achieve dimensionality reduction. The processed result is then combined with the original image to form sample pairs, which enables the generative model to learn the inverse process of diffusion, thereby achieving image dimensionality upscaling. This allows the model to complete training quickly while ensuring prediction accuracy.

[0079] Step S204: Perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set.

[0080] Specifically, after obtaining the original image, it needs to be diffused to convert it into a noisy image set, which facilitates subsequent training of the generative model. In this process, in order to reduce the impact of the same dimensions between the images before and after diffusion, the image components of the original image can be attenuated in each iteration of the diffusion process. This is done to reduce the dimensionality of the image when the image components are attenuated to the minimum. This process continues until the dimensionality reduction and the noisy images meet the requirements. Then, the images obtained in the diffusion process can be combined into a noisy image set for subsequent training of the generative model.

[0081] Specifically, diffusion processing refers to the process of adding Gaussian noise to the original image, transforming it into a noisy image with a different representation from the original. Multiple noisy images are obtained during the diffusion iteration process. Correspondingly, image components refer to the orthogonal components of the original image after orthogonal decomposition, such as low-frequency and high-frequency components. The noisy image set refers to the collection of noisy images obtained after diffusion processing and attenuation of image components. By attenuating image components in each diffusion iteration, dimensionality reduction can be achieved, transforming the original image into a noisy image set. This facilitates subsequent model training using images of different dimensions. The image dimension of the images in the noisy image set can be less than or equal to the image dimension of the original image.

[0082] Therefore, after obtaining the original image, in order to train a generative model that meets the usage requirements, the original image needs to be diffused first. To reduce the impact of the unchanged image dimensionality, the components can be attenuated synchronously in each iteration of the diffusion process until the image components are attenuated to a sufficiently small value, thus obtaining the original image with reduced dimensionality. This process can be repeated in subsequent iterations, using a lower-dimensional diffusion process to approximate the result, thereby achieving the goal of image dimensionality reduction. By repeating the above process, multiple dimensionality reductions can be achieved, resulting in a noisy image set. This noisy image set can then be used to train the corresponding generative model for the inverse process.

[0083] In practical applications, when attenuating image components, the attenuation of different components can be controlled by preset hyperparameters. This allows for separate attenuation processing of each image component during different iterations, achieving the goal of reducing image dimensionality. During the attenuation process, the hyperparameter for image component attenuation decreases from 1 to near 0 before the required number of iterations for dimensionality reduction, indicating that the attenuation is complete and thus dimensionality reduction is achieved. In other words, attenuation of the original image will result in the loss of some components. The solution provided in this embodiment controls the attenuation of this portion of the components, thus the error introduced by dimensionality reduction is controllable. This allows for dimensionality reduction to be completed with sufficiently small errors, resulting in a noisy image that meets the usage requirements. For example, if the original image is 256-dimensional, attenuating the image components during diffusion can yield a 128-dimensional image, and so on, until the final noisy image is a 64-dimensional image.

[0084] In practice, the image components are attenuated while the original image is diffused, which can be achieved using the following formula (1):

[0085]

[0086] Where the subscript t represents the current iteration step of the diffusion process, the subscript k represents the number of dimensionality reductions, and y k,t D represents the noisy image after the original image undergoes k-fold dimensionality reduction and t-fold noise addition during the diffusion process. k Denotes the dimensionality reduction operator, x k,t With y k,t Correspondingly, v represents the original image without dimensionality reduction but with attenuation of its image components. i λ represents the i-th component of the original image. i,t This indicates the condition for v after t iterations. i The degree of attenuation, z i This indicates standard Gaussian noise at v i The projection of the subspace, σ i,t This indicates the condition for v after t iterations. i The standard deviation of the added noise. In formula (1), the subscripts of the summation start from i = k, excluding i = 0, 1, ..., k-1. Here, these λ values ​​are considered. i,t In iteration to T k The previous time was close enough to 0, meaning that these v values ​​had already been... i The component attenuation is small enough to enable low-dimensional approximation.

[0087] Formula (1) above allows us to obtain a noisy image after obtaining the original image and performing t diffusion iterations and k dimensionality reduction iterations. This makes it convenient to use in subsequent noisy image sets. During the diffusion and attenuation processes, some orthogonal components of the added Gaussian noise are lost during dimensionality reduction. Therefore, it is necessary to record the image noise parameters during the diffusion process so that when using the trained target generation model for image generation, the lost noise can be re-added to the image to obtain a more accurate restored image. Therefore, σ in formula (1) can be used. i,t A Gaussian noise is resampled and added to the noisy image during image generation by the generative model to achieve noise compensation.

[0088] For example, first, a 256-dimensional original image is obtained. Then, the 256-dimensional original image is normalized to obtain an intermediate image. Simultaneously, orthogonal decomposition is performed on the intermediate image to obtain its corresponding orthogonal components. During the diffusion process, Gaussian noise is added to the 256-dimensional intermediate image through the calculation process shown in formula (1). Simultaneously, in each iteration of noise addition, different orthogonal components are attenuated until the attenuation value of the orthogonal components approaches 0, thus achieving dimensionality reduction of the 256-dimensional image to obtain a 128-dimensional image. This process continues, through continuous diffusion and component attenuation, finally resulting in a 16-dimensional noisy image. This facilitates subsequent training of the generative model.

[0089] Furthermore, when performing diffusion processing on the original image and attenuation processing on the image components, in order to ensure successful dimensionality reduction of the original image, it is necessary to first normalize the original image and determine its corresponding image components for diffusion and attenuation processing. In this embodiment, the specific implementation method is as follows:

[0090] The pixel values ​​corresponding to the pixels in the original image are normalized to obtain an intermediate image, and the intermediate image is orthogonally decomposed to obtain image components; the intermediate image is diffused and the image components are attenuated to obtain the noisy image set.

[0091] Specifically, the intermediate image refers to the vector representation of the original image in high-dimensional Euclidean space, and the image component refers to the orthogonal component obtained after orthogonal decomposition of the intermediate image. There are multiple image components, which are used to attenuate in each iteration process, so that the original image can be reduced in dimensionality multiple times.

[0092] Therefore, after obtaining the original image, in order to add noise to the original image and attenuate different image components during the noise addition process to achieve dimensionality reduction, we can first normalize the pixel values ​​corresponding to each pixel in the original image to obtain a vector representation in high-dimensional Euclidean space, i.e., an intermediate image. Simultaneously, we perform orthogonal decomposition on the intermediate image to obtain multiple image components, such as high-frequency or low-frequency components. Then, we perform diffusion processing on the intermediate image, and during the iterative diffusion process, we attenuate the image components to achieve dimensionality reduction. This results in a set of noisy images with dimensions less than or equal to the dimensions of the original image, which can be used for subsequent model training.

[0093] In summary, to change the dimension of the image during the diffusion process, image component attenuation can be used. This involves attenuating the image components in each iteration of the diffusion process to convert the high-dimensional original image into a low-dimensional noisy image with added Gaussian noise. This noisy image set is then used to facilitate the subsequent training of the generative model, enabling the model to learn noise compensation capabilities and thus reducing the difficulty of model fitting.

[0094] Furthermore, when performing diffusion processing on the intermediate image and attenuation processing on the image components, the images and components processed in different diffusion periods are different. This allows the original image to be continuously noise-added and dimensionality-reduced to obtain a noisy image set that meets the usage requirements. In this embodiment, the specific implementation method is as follows:

[0095] Determine the intermediate image corresponding to the i-th diffusion cycle, and the i-th image component corresponding to the intermediate image, where i starts from 1 and is a positive integer;

[0096] Add i-th noise to the intermediate image and perform attenuation processing on the i-th image component;

[0097] If the attenuation result of the i-th image component is not less than the component threshold, a first noisy image with the same image dimension as the intermediate image is determined according to the diffusion processing result. The first noisy image is used as the intermediate image, i is incremented by 1, and the steps of determining the intermediate image corresponding to the i-th diffusion period and the i-th image component corresponding to the intermediate image are executed.

[0098] If the attenuation result of the i-th image component is less than the component threshold, a second noisy image with an image dimension smaller than the intermediate image is determined according to the diffusion processing result. The second noisy image is used as the intermediate image, i is incremented by 1, and the steps of determining the intermediate image corresponding to the i-th diffusion period and the i-th image component corresponding to the intermediate image are executed.

[0099] The noisy image set is formed based on the first noisy image and the second noisy image until the diffusion processing and attenuation processing meet the iteration stopping condition.

[0100] Specifically, the diffusion period refers to the period during which the intermediate image undergoes diffusion processing, with each period corresponding to one diffusion iteration. Correspondingly, the i-th image component refers to the image component that needs attenuation processing in the i-th diffusion period. Correspondingly, the i-th noise refers to the Gaussian noise that needs to be added to the intermediate image in the i-th diffusion period. Correspondingly, the first noisy image refers to the image that has undergone diffusion processing but whose image dimensions have not changed; correspondingly, the second noisy image refers to the image that has undergone diffusion processing and whose image dimensions have been reduced. Correspondingly, the iteration stopping condition refers to the stopping condition for diffusion processing of the intermediate image and image component attenuation. When the stopping condition is met, all the obtained first and second noisy images can be combined into a noisy image set, which is convenient for subsequent image sampling based on this set, and for training the model after obtaining the noisy images.

[0101] Based on this, firstly, the intermediate image that needs to be diffused in the i-th diffusion cycle and the i-th image component that needs to be attenuated in the current i-th diffusion cycle are determined. Secondly, i-th noise is added to the intermediate image in the i-th diffusion cycle, and the i-th image component is attenuated. If the attenuation result of the i-th image component is not less than the component threshold, it means that the current i-th diffusion cycle does not meet the image dimensionality reduction condition. In this case, the intermediate image after diffusion processing in the i-th diffusion cycle can be used as the first noisy image, and the current first noisy image has the same image dimension as the intermediate image. Subsequently, the first noisy image is used as the intermediate image in the (i+1)-th diffusion cycle, and the diffusion and component attenuation processing is performed again. During this process, if the image dimension of the noisy image obtained in each diffusion cycle is the same as the image dimension of the original image, it is used as the first noisy image.

[0102] When the attenuation result of an image component is less than the component threshold in a certain diffusion cycle, it indicates that the current diffusion cycle meets the image dimensionality reduction condition. At this point, the intermediate image after diffusion processing can be used as the second noisy image, and the image dimension of the obtained second noisy image is smaller than that of the intermediate image. If further dimensionality reduction processing is needed, the second noisy image can be used as the intermediate image for the next diffusion cycle, and the diffusion and component attenuation processing can be repeated. During this process, if the image dimension of the noisy image obtained in each diffusion cycle is smaller than that of the original image, it is used as the second noisy image.

[0103] This embodiment describes the process of generating a noisy image set by taking a maximum of 4 iterations in the diffusion process and changing the dimension once as an example. The diffusion process in actual applications can be referred to the same or corresponding descriptions in this embodiment, and will not be elaborated on in this embodiment.

[0104] First, the normalized original image x0 is obtained. In the first iteration cycle, noise is added to the original image x0, and the image components are attenuated. After processing, it is determined that the image with noise and attenuated components does not meet the dimensionality reduction condition, so a noisy image x1 is obtained based on the diffusion result. Second, noise is added to the noisy image x1, and the image components are attenuated again. It is determined that the image with noise and attenuated components does not meet the dimensionality reduction condition, so a noisy image x2 is obtained based on the diffusion result. Third, noise is added to the noisy image x2, and the image components are attenuated again. It is determined that the image with noise and attenuated components meets the dimensionality reduction condition, so a noisy image x3 is obtained based on the diffusion result. Finally, noise is added to the noisy image x3, and the image components are attenuated again. It is determined that the image with noise and attenuated components does not meet the dimensionality reduction condition, so a noisy image x4 is obtained based on the diffusion result. Among them, the image dimensions of noisy images x1 and x2 are the same as those of the original image x0, the image dimensions of noisy images x3 and x4 are smaller than those of the original image, and the image dimensions of noisy images x3 and x4 are the same.

[0105] At this point, the noisy images x1, x2, x3, and x4 can be combined to form the noisy image set corresponding to the original image x0, which can then be used to determine the noisy image in conjunction with the original image x0 for training the generative model.

[0106] In summary, by combining noisy images with altered dimensions and noisy images without altered dimensions to generate a noisy image set, it is convenient to train the generation model with noisy images of different dimensions during subsequent sampling. This allows the generation model to learn the ability to recover the original dimensions, thereby improving the model's prediction accuracy.

[0107] Step S206: Determine the noisy image based on the original image and the noisy image set, and input the noisy image into the initial generation model for processing to obtain the restored image.

[0108] Specifically, after obtaining the set of noisy images corresponding to the original image, in order to train a generative model that meets the usage requirements and has the ability to increase dimensionality and compensate for noise, the noisy images can be determined by combining the original image and the set of noisy images. The noisy images are used as samples and the original images are used as labels to train the initial generative model until the target generative model that meets the usage requirements is trained.

[0109] In this context, the noisy image refers to an image obtained by combining a set of noisy images and the original image, which can be used as a training sample for the model. It can be any image from the noisy image set or an image obtained by processing the original image. The initial generative model refers to a model capable of processing images, but which has not yet been fully trained. While it possesses some predictive ability, its accuracy is low. In other words, the initial generative model is a pre-trained generative model. After determining the noisy image by combining the noisy image set and the original image, the generative model processes the noisy image to obtain a reconstructed image that approximates the original image. For example, an original image with a sharpness of 1, after being noisily processed, becomes a noisy image with a sharpness of 8. Inputting this into a trained generative model yields a reconstructed image with a sharpness of 2, which is closer to the original image with a sharpness of 1. Correspondingly, the reconstructed image refers to the image obtained by inverting the diffusion process of the noisy image using the initial generative model; it has the same image dimensions as the original image.

[0110] In other words, in order to train a generative model with the ability to increase dimensionality and compensate for noise, we can combine the original image and the set of noisy images to determine the noisy image. The noisy image is then input into the initial generative model for processing to obtain the restored image. By combining the restored image and the original image, we can analyze the inverse processing accuracy of the diffusion process of the initial generative model, which will facilitate subsequent parameter adjustments to obtain a more accurate target generative model.

[0111] Furthermore, in order to reduce the difficulty of model fitting and improve the model prediction accuracy when determining the noisy image, a noisy image of any dimension can be determined by random sampling, and then combined with the original image to form a noisy image. In this embodiment, the specific implementation method is as follows:

[0112] A first target denoised image is randomly sampled from the set of denoised images, and the original image is subjected to diffusion processing based on the first target denoised image to obtain a second target denoised image; the first target denoised image and the second target denoised image are used as the denoised image.

[0113] Specifically, the first target denoised image refers to the denoised image obtained after random sampling from the denoised image set. The first target denoised image obtained at this time can be a denoised image with the same image dimension as the original image, or it can be a denoised image with a smaller image dimension than the original image. Correspondingly, the second target denoised image refers to the denoised image obtained after diffusion processing of the original image, and the diffusion is obtained by combining the diffusion of the relevant Gaussian noise of the first target denoised image.

[0114] Based on this, after obtaining the set of noisy images, in order to train a generative model with higher prediction accuracy and stronger prediction ability, the set of noisy images can be randomly sampled first, so that a first target noisy image can be sampled from the set. In order to train the model, the original image can be diffused based on the first target noisy image to obtain a diffused second target noisy image. Then, the first target noisy image and the second target noisy image can be used as noisy images to input into the model for training.

[0115] Following the previous example, after obtaining the noisy images x1, x2, x3, and x4, random sampling can be performed on these images. If the noisy image x1 or x2 is sampled, it means that the image dimension of the noisy image obtained at this time is the same as the image dimension of the original image x0. Then, the original image x0 is diffused according to the noisy image x1 or x2 to obtain the diffused image xi. After that, the diffused image xi and the noisy image x1 or x2 are input into the generative model for processing to predict the original image x0, thereby obtaining a restored image that is close to the original image x0. Then, the subsequent model parameter tuning can be performed.

[0116] If a noisy image x4 or x3 is sampled, it means that the image dimension of the noisy image obtained at this time is smaller than the image dimension of the original image x0. Then, the original image x0 is diffused according to the noisy image x3 or x4 to obtain the diffused image xi. After that, the diffused image xi and the noisy image x4 or x3 are input into the generative model for processing to predict the original image x0 and obtain a restored image that is close to the original image x0. Then, the subsequent model parameter tuning can be performed.

[0117] In summary, by obtaining noisy images with different dimensions through random sampling and using them in conjunction with the original image for model training, the model can learn to recover the dimensions, thereby effectively improving the model's prediction accuracy.

[0118] Step S208: Based on the restored image and the original image, the parameters of the initial generative model are adjusted until a target generative model that meets the training stopping condition is obtained.

[0119] Specifically, after obtaining the restored image output by the generative model, the predictive ability of the initial generative model can be analyzed by comparing the restored image with the original image. If the predictive ability does not meet the requirements, the parameters can be adjusted until the target generative model that meets the training stopping condition is obtained.

[0120] The training stopping condition can be a loss value comparison condition, an iteration count condition, or a validation set verification condition. In practical applications, the appropriate condition can be selected based on the requirements. This embodiment does not impose any limitations on this condition.

[0121] Furthermore, model parameter adjustments can be made by calculating the loss value. In this embodiment, the specific implementation method is as follows:

[0122] The model loss value corresponding to the initial generated model is calculated based on the restored image and the original image; if the model loss value is less than a preset loss value threshold, the initial generated model is determined to meet the training stopping condition, and the initial generated model is used as the target generated model.

[0123] Based on this, after obtaining the reconstructed image, a preset loss function can be used to calculate the model loss value corresponding to the reconstructed image and the original image. Then, the loss value is compared with a preset loss threshold. If the loss value is greater than the threshold, it indicates that the current model's prediction accuracy is insufficient, and new images need to be acquired for further training and parameter tuning. If the loss value is less than or equal to the threshold, it indicates that the current model's prediction accuracy meets the requirements, and it can be used as the target generation model.

[0124] In practical applications, the loss value can be calculated using cross-entropy loss function, absolute value loss function, squared loss function, etc., and this embodiment does not impose any limitations on it.

[0125] The generative model training method provided in this specification aims to achieve rapid and efficient model training while reducing model fitting difficulty. After acquiring the original image, a diffusion process is performed, simultaneously attenuating the image components of the original image during diffusion. This results in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original image and the noisy image set and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original image and the restored image until the target generative model that meets the training stopping condition is obtained. This method achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0126] Corresponding to the above method embodiments, this specification also provides embodiments of a generative model training apparatus. Figure 3 A schematic diagram of a generative model training apparatus according to one embodiment of this specification is shown. Figure 3 As shown, the device includes:

[0127] The acquisition module 302 is configured to acquire the raw image;

[0128] The processing module 304 is configured to perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set;

[0129] The determining module 306 is configured to determine a noisy image based on the original image and the noisy image set, and input the noisy image into the initial generation model for processing to obtain the restored image;

[0130] Training module 308 is configured to tune the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0131] In an optional embodiment, the processing module 304 is further configured to:

[0132] The pixel values ​​corresponding to the pixels in the original image are normalized to obtain an intermediate image, and the intermediate image is orthogonally decomposed to obtain image components; the intermediate image is diffused and the image components are attenuated to obtain the noisy image set.

[0133] In an optional embodiment, the processing module 304 is further configured to:

[0134] The process involves: determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; adding i-th noise to the intermediate image and performing attenuation processing on the i-th image component; if the attenuation result of the i-th image component is not less than a component threshold, determining a first noisy image with the same image dimension as the intermediate image based on the diffusion processing result, using the first noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; if the attenuation result of the i-th image component is less than a component threshold, determining a second noisy image with a smaller image dimension than the intermediate image based on the diffusion processing result, using the second noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; and finally, forming the noisy image set based on the first noisy image and the second noisy image when the diffusion processing and attenuation processing meet the iteration stopping condition.

[0135] In an optional embodiment, the determining module 306 is further configured to:

[0136] A first target denoised image is randomly sampled from the set of denoised images, and the original image is subjected to diffusion processing based on the first target denoised image to obtain a second target denoised image; the first target denoised image and the second target denoised image are used as the denoised image.

[0137] In an optional embodiment, the training module 308 is further configured to:

[0138] The model loss value corresponding to the initial generated model is calculated based on the restored image and the original image; if the model loss value is less than a preset loss value threshold, the initial generated model is determined to meet the training stopping condition, and the initial generated model is used as the target generated model.

[0139] In an optional embodiment, the diffusion process and the attenuation process are calculated using the following formula:

[0140]

[0141] Where the subscript t represents the current iteration step of the diffusion process, the subscript k represents the number of dimensionality reductions, and y k,t D represents the noisy image after the original image undergoes k-fold dimensionality reduction and t-fold noise addition during the diffusion process. k Denotes the dimensionality reduction operator, x k,t With y k,t Correspondingly, v represents the original image without dimensionality reduction but with attenuation of its image components. i λ represents the i-th component of the original image. i,t This indicates the condition for v after t iterations. i The degree of attenuation, z i This indicates standard Gaussian noise at v i The projection of the subspace, σ i,t This indicates the condition for v after t iterations. i The standard deviation of the added noise.

[0142] The generative model training apparatus provided in this specification, in order to complete model training quickly and efficiently while reducing the difficulty of model fitting, can perform diffusion processing on the original image after acquisition. Simultaneously, during the diffusion process, the image components of the original image are attenuated to obtain a noisy image set that has undergone diffusion and dimensionality reduction. Then, based on the original image and the noisy image set, noisy images suitable for model training are selected and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original image and the restored image until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing the difficulty of model fitting, and thus enabling the rapid and efficient training of a high-precision generative model.

[0143] The above is an illustrative scheme of a generative model training device according to this embodiment. It should be noted that the technical solution of this generative model training device and the technical solution of the generative model training method described above belong to the same concept. For details not described in detail in the technical solution of the generative model training device, please refer to the description of the technical solution of the generative model training method described above.

[0144] Figure 4 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0145] Step S402: Obtain the image to be processed uploaded by the user terminal;

[0146] Step S404: Input the image to be processed into the target generation model in the above method for processing to obtain the target image;

[0147] Step S406: Determine the target object based on the target image, and load the object information corresponding to the target object;

[0148] Step S408: Send the object information to the user terminal.

[0149] Specifically, the user terminal refers to the terminal held by a user with a need to purchase goods, including but not limited to mobile phones, computers, or tablets. Correspondingly, the image to be processed refers to the image that needs to be processed by the target generation model, including but not limited to images requiring sharpness adjustment, size adjustment, or color adjustment; this embodiment does not impose any limitations. Correspondingly, the target object refers to the identifiable object in the target image, which can be a face, item, or product. Correspondingly, the object information is the relevant information of the target object, including but not limited to descriptive information and links.

[0150] Therefore, after obtaining the image to be processed uploaded by the user, in order to use a clearer, larger, or more color-accurate image for downstream business processing, the image to be processed can first be input into the target generation model to obtain the target image. Then, the target object is determined by combining the target image, and its corresponding object information is loaded and sent to the user terminal for use.

[0151] It should be noted that the processing of the generative model is the inverse of the diffusion process in the training process of the above model. That is, iterates back step by step in reverse. When iterates to the number of steps corresponding to the dimensionality reduction in the diffusion process, it is necessary to perform corresponding dimensionality increase and noise compensation, so as to process the image to be processed into the target image.

[0152] Furthermore, after the image to be processed is input into the target generation model, when the generation model performs the inverse processing of the diffusion process, considering that the image to be processed may not reflect the actual content captured, the restoration can be completed by adding loss noise and dimensionality increase. In this embodiment, the specific implementation method is as follows:

[0153] The image to be processed is input into the target generation model; the image to be processed is subjected to dimensionality upscaling to obtain an intermediate noisy image; the intermediate noisy image is subjected to noise compensation processing to obtain the target image and output the target generation model.

[0154] Specifically, dimensionality upscaling refers to the process of upscaling a low-dimensional image to the same dimension as the real scene; correspondingly, noise compensation refers to the process of adding noise to the image to be processed to compensate for any potential loss.

[0155] Based on this, after obtaining a lower-dimensional image to be processed that carries Gaussian noise, in order to facilitate downstream business use, the image to be processed can be input into the target generation model. The target generation model will then perform dimensionality upscaling on the image to obtain an intermediate noisy image with the same image dimension as the real scene. Subsequently, noise compensation processing is required on the intermediate noisy image. This allows the dimensionality upscaling and noise addition to be completed in the inverse processing, and finally the target image is obtained and the generation model is output for the convenience of downstream business use.

[0156] In practical applications, when using a target generation model to upscale an image, it can be understood as upsampling the image, such as upsampling an image into a larger one. The components lost during diffusion can be gradually compensated for by the target generation model in the inverse processing iterations. Furthermore, if the image cannot be upscaled, noise compensation can be performed directly.

[0157] Furthermore, in order to make the image processed by the model more closely resemble the real image during noise compensation, compensation can be performed only for the noise that causes image loss. In this embodiment, the specific implementation method is as follows:

[0158] Determine the image loss information corresponding to the intermediate noisy image; perform noise compensation processing on the intermediate noisy image based on the image loss information to obtain the target image.

[0159] Specifically, image loss information refers to the information related to noise that may be lost in the image corresponding to the real scene. Based on this, we can first determine the image loss information corresponding to the intermediate noisy image, and then determine the noise lost in the diffusion process. Then, we can perform noise compensation processing on the intermediate noisy image based on the image loss information to obtain the target image and output the model.

[0160] In practical applications, when the generative model performs noise compensation on the intermediate noisy image, it's because some noise is lost during the dimensionality reduction operation of the forward diffusion process. The noise added during diffusion is Gaussian noise, and its standard deviation is controlled by a preset hyperparameter. Therefore, this parameter can be used to re-add Gaussian noise to the intermediate noisy image to achieve noise compensation and restore it to the original image as closely as possible.

[0161] For example, after obtaining a 16-dimensional noisy image, an image generation model with noise compensation and dimensionality upscaling capabilities can be acquired first. Then, the 16-dimensional noisy image is input into the image generation model for processing. After noise compensation and dimensionality upscaling by the model, a 256-dimensional restored image is obtained. The dimensionality upscaling and noise compensation are actually the inverse of the diffusion process, i.e., initializing t=T. k k = K-1, initialize y k,1 The noise is Gaussian, followed by a single-step diffusion inverse process t := t-1; dimensionality increase to determine if t equals T. k If not, it indicates that further noise compensation and dimensionality increase are needed, and then a judgment on whether the inverse processing has been completed can be made; if yes, it indicates that the conditions for dimensionality increase are met, and then y can be... k,1 The image is then upgraded in dimensionality, and missing noise is compensated for. Afterward, k := k-1, and then it is checked whether t equals T0. If they do, it means the upgraded image is identical to the original image, and noise compensation is complete; the restored image can then be output. If they do not equal, it means the upgraded image has a different dimension than the original image, and noise compensation and dimensionality upgrade need to continue. In this case, the process returns to the single-step diffusion inverse process t := t-1 until the restored image is output.

[0162] In summary, by incorporating image noise parameters into the generative model and performing inverse diffusion processing on the noisy image, dimensionality enhancement and noise compensation can be achieved during the inverse processing. This allows the model to have higher predictive power and enables the rapid training of generative models with lower fitting difficulty.

[0163] In summary, by using a generative model with variable dimensions to process images, images can be processed to resemble real-world scenes, which can then be used in downstream business applications to better meet business needs.

[0164] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 5 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 5 As shown, the device includes:

[0165] Image acquisition module 502 is configured to acquire images to be processed uploaded by the user terminal;

[0166] Model processing module 504 is configured to input the image to be processed into the target generation model in the above method for processing, and obtain the target image;

[0167] The object determination module 506 is configured to determine the target object based on the target image and load the object information corresponding to the target object;

[0168] The information sending module 508 is configured to send the object information to the user terminal.

[0169] In an optional embodiment, the model processing module 504 is further configured to:

[0170] The image to be processed is input into the target generation model; the image to be processed is subjected to dimensionality upscaling to obtain an intermediate noisy image; the intermediate noisy image is subjected to noise compensation processing to obtain the target image and output the target generation model.

[0171] In an optional embodiment, the model processing module 504 is further configured to:

[0172] Determine the image loss information corresponding to the intermediate noisy image; perform noise compensation processing on the intermediate noisy image based on the image loss information to obtain the target image.

[0173] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0174] Figure 6 A flowchart of another image processing method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0175] Step S602: Obtain the initial shopping search image uploaded by the user terminal;

[0176] Step S604: Input the initial shopping search image into the target generation model in the above method for processing to obtain the target shopping search image corresponding to the initial shopping search image;

[0177] Step S606: Determine associated products based on the target shopping search image, and load the product information corresponding to the associated products;

[0178] Step S608: The product information is sent to the user terminal, wherein the user terminal generates and displays a product recommendation interface based on the product information.

[0179] Specifically, the user terminal refers to the terminal held by a user with a need to purchase goods, including but not limited to mobile phones, computers, or tablets. Correspondingly, the initial shopping search image refers to the image submitted by the user when searching for goods. The target shopping search image refers to the image obtained after processing by the target generation model; it is clearer than the initial shopping search image and meets business search requirements. Related products refer to the products that can be searched using the related images. Product information refers to the relevant information for the corresponding related products, used to display the product details interface on the user terminal.

[0180] Based on this, after obtaining the initial shopping search image uploaded by the user terminal, the initial shopping search image can be input into the target generation model in the above method for processing to obtain the target shopping search image corresponding to the initial shopping search image. Then, associated products can be determined based on the target shopping search image, and the product information corresponding to the associated products can be loaded. Finally, the product information is sent to the user terminal, and the user terminal generates a product recommendation interface based on the product information and displays it.

[0181] For example, after user A takes a picture of item 1 and uploads it, the image can be input into an image generation model for processing to obtain a target image with higher clarity that is suitable for the product search scenario. Then, based on the target image, product A and product B are determined, and the information corresponding to product A and product B is loaded. This information is then sent to user A's client to display a product recommendation interface that includes at least product A and product B.

[0182] In summary, by using a generative model with variable dimensions to process images, images can be processed to resemble real-world scenes, which can then be used in downstream business applications to better meet business needs.

[0183] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 7 A schematic diagram of another image processing apparatus provided in one embodiment of this specification is shown. Figure 7 As shown, the device includes:

[0184] Image acquisition module 702 is configured to acquire the initial shopping search image uploaded by the user terminal;

[0185] The input model module 704 is configured to input the initial shopping search image into the target generation model in the above method for processing, so as to obtain the target shopping search image corresponding to the initial shopping search image;

[0186] The information loading module 706 is configured to determine associated products based on the target shopping search image and load the product information corresponding to the associated products;

[0187] The information sending module 708 is configured to send the product information to the user terminal, wherein the user terminal generates and displays a product recommendation interface based on the product information.

[0188] In summary, by using a generative model with variable dimensions to process images, images can be processed to resemble real-world scenes, which can then be used in downstream business applications to better meet business needs.

[0189] The above is an illustrative scheme of another image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.

[0190] Figure 8 A flowchart of another generative model training method according to an embodiment of this specification is shown, applied to a server, and specifically includes the following steps.

[0191] Step S802: Receive the original image uploaded by the model requester.

[0192] Step S804: Perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set.

[0193] Step S806: Determine the noisy image based on the original image and the noisy image set, and input the noisy image into the initial generation model for processing to obtain the restored image.

[0194] Step S808: Based on the restored image and the original image, the parameters of the initial generative model are adjusted until a target generative model that meets the training stopping condition is obtained.

[0195] Step S810: Determine the model parameters corresponding to the target generation model and feed the model parameters back to the model demand side.

[0196] Specifically, the model demand side refers to the demand side with the training target to generate a model. Correspondingly, the server side refers to the server side that provides the model training requirement. It provides a pre-trained generative model. When the model demand side has a need to use the model, it can send a model training request to the server side. At this time, the server side can use the pre-trained initial generative model and combine it with the samples provided by the model demand side to further train it. In this way, the computing resources of the server side can be used to assist the model demand side in completing the training of the target generative model, so as to save resources and improve the model training accuracy.

[0197] Optionally, the step of performing diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set includes:

[0198] The pixel values ​​corresponding to the pixels in the original image are normalized to obtain an intermediate image, and the intermediate image is orthogonally decomposed to obtain image components; the intermediate image is diffused and the image components are attenuated to obtain the noisy image set.

[0199] Optionally, the step of performing diffusion processing on the intermediate image and attenuation processing on the image components to obtain the noisy image set includes:

[0200] The process involves: determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; adding i-th noise to the intermediate image and performing attenuation processing on the i-th image component; if the attenuation result of the i-th image component is not less than a component threshold, determining a first noisy image with the same image dimension as the intermediate image based on the diffusion processing result, using the first noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; if the attenuation result of the i-th image component is less than a component threshold, determining a second noisy image with a smaller image dimension than the intermediate image based on the diffusion processing result, using the second noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; and finally, forming the noisy image set based on the first noisy image and the second noisy image when the diffusion processing and attenuation processing meet the iteration stopping condition.

[0201] Optionally, determining the noisy image based on the original image and the set of noisy images includes:

[0202] A first target denoised image is randomly sampled from the set of denoised images, and the original image is subjected to diffusion processing based on the first target denoised image to obtain a second target denoised image; the first target denoised image and the second target denoised image are used as the denoised image.

[0203] Optionally, the step of tuning the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained includes:

[0204] The model loss value corresponding to the initial generated model is calculated based on the restored image and the original image; if the model loss value is less than a preset loss value threshold, the initial generated model is determined to meet the training stopping condition, and the initial generated model is used as the target generated model.

[0205] Optionally, the diffusion treatment and the attenuation treatment are calculated using the following formula:

[0206]

[0207] Where the subscript t represents the current iteration step of the diffusion process, the subscript k represents the number of dimensionality reductions, and y k,t D represents the noisy image after the original image undergoes k-fold dimensionality reduction and t-fold noise addition during the diffusion process. k Denotes the dimensionality reduction operator, x k,t With y k , tCorrespondingly, v represents the original image without dimensionality reduction but with attenuation of its image components. i λ represents the i-th component of the original image. i,t This indicates the condition for v after t iterations. i The degree of attenuation, z i This indicates standard Gaussian noise at v i The projection of the subspace, σ i,t This indicates the condition for v after t iterations. i The standard deviation of the added noise.

[0208] In summary, to achieve rapid and efficient model training while reducing model fitting difficulty, a diffusion process can be performed on the original image after acquisition. Simultaneously, the image components of the original image are attenuated during the diffusion process, resulting in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original and noisy image sets and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original and restored images until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0209] The alternative generative model training method provided in this embodiment is an illustrative scheme of another generative model training method in this embodiment. It should be noted that the technical solution of the alternative generative model training method belongs to the same concept as the technical solution of the first generative model training method described above. For details not described in detail in the technical solution of the alternative generative model training method, please refer to the description of the technical solution of the first generative model training method described above.

[0210] Corresponding to the above method embodiments, this specification also provides another embodiment of a generative model training device. Figure 9 A schematic diagram of another generative model training apparatus provided in one embodiment of this specification is shown. Figure 9 As shown, the device includes:

[0211] The image receiving module 902 is configured to receive the original image uploaded by the model request end;

[0212] Image processing module 904 is configured to perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set;

[0213] The image determination module 906 is configured to determine a noisy image based on the original image and the noisy image set, and input the noisy image into the initial generation model for processing to obtain the restored image;

[0214] The model training module 908 is configured to tune the parameters of the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0215] The parameter sending module 910 is configured to determine the model parameters corresponding to the target generated model and feed the model parameters back to the model demand end.

[0216] Optionally, the step of performing diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set includes:

[0217] The pixel values ​​corresponding to the pixels in the original image are normalized to obtain an intermediate image, and the intermediate image is orthogonally decomposed to obtain image components; the intermediate image is diffused and the image components are attenuated to obtain the noisy image set.

[0218] Optionally, the step of performing diffusion processing on the intermediate image and attenuation processing on the image components to obtain the noisy image set includes:

[0219] The process involves: determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; adding i-th noise to the intermediate image and performing attenuation processing on the i-th image component; if the attenuation result of the i-th image component is not less than a component threshold, determining a first noisy image with the same image dimension as the intermediate image based on the diffusion processing result, using the first noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; if the attenuation result of the i-th image component is less than a component threshold, determining a second noisy image with a smaller image dimension than the intermediate image based on the diffusion processing result, using the second noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; and finally, forming the noisy image set based on the first noisy image and the second noisy image when the diffusion processing and attenuation processing meet the iteration stopping condition.

[0220] Optionally, determining the noisy image based on the original image and the set of noisy images includes:

[0221] A first target denoised image is randomly sampled from the set of denoised images, and the original image is subjected to diffusion processing based on the first target denoised image to obtain a second target denoised image; the first target denoised image and the second target denoised image are used as the denoised image.

[0222] Optionally, the step of tuning the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained includes:

[0223] The model loss value corresponding to the initial generated model is calculated based on the restored image and the original image; if the model loss value is less than a preset loss value threshold, the initial generated model is determined to meet the training stopping condition, and the initial generated model is used as the target generated model.

[0224] Optionally, the diffusion treatment and the attenuation treatment are calculated using the following formula:

[0225]

[0226] Where the subscript t represents the current iteration step of the diffusion process, the subscript k represents the number of dimensionality reductions, and y k,t D represents the noisy image after the original image undergoes k-fold dimensionality reduction and t-fold noise addition during the diffusion process. k Denotes the dimensionality reduction operator, x k,t With y k,t Correspondingly, v represents the original image without dimensionality reduction but with attenuation of its image components. i λ represents the i-th component of the original image. i,t This indicates the condition for v after t iterations. i The degree of attenuation, z i This indicates standard Gaussian noise at v i The projection of the subspace, σ i,t This indicates the condition for v after t iterations. i The standard deviation of the added noise.

[0227] In summary, to achieve rapid and efficient model training while reducing model fitting difficulty, a diffusion process can be performed on the original image after acquisition. Simultaneously, the image components of the original image are attenuated during the diffusion process, resulting in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original and noisy image sets and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original and restored images until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0228] In summary, to achieve rapid and efficient model training while reducing model fitting difficulty, a diffusion process can be performed on the original image after acquisition. Simultaneously, the image components of the original image are attenuated during the diffusion process, resulting in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original and noisy image sets and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original and restored images until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0229] The above is an illustrative scheme of another generative model training device in this embodiment. It should be noted that the technical solution of this generative model training device and the technical solution of the generative model training method described above belong to the same concept. For details not described in detail in the technical solution of the generative model training device, please refer to the description of the technical solution of the generative model training method described above.

[0230] The following is in conjunction with the appendix Figure 10 Taking the application of the generative model training method provided in this specification in a practical application scenario as an example, the generative model training method will be further explained. Among them, Figure 10 The flowchart of a generative model training method according to an embodiment of this specification is shown, which specifically includes the following steps.

[0231] Step S1002: Obtain the original image.

[0232] Step S1004: Normalize the pixel values ​​corresponding to the pixels in the original image to obtain an intermediate image, and perform orthogonal decomposition on the intermediate image to obtain image components.

[0233] Step S1006: Perform diffusion processing on the intermediate image and attenuation processing on the image components to obtain a set of noisy images.

[0234] The process involves: determining an intermediate image corresponding to the i-th diffusion cycle and an i-th image component corresponding to the intermediate image; adding i-th noise to the intermediate image and performing attenuation processing on the i-th image component; if the attenuation result of the i-th image component is not less than a component threshold, determining a first noisy image with the same image dimension as the intermediate image based on the diffusion processing result, using the first noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; if the attenuation result of the i-th image component is less than a component threshold, determining a second noisy image with a smaller image dimension than the intermediate image based on the diffusion processing result, using the second noisy image as the intermediate image, incrementing i by 1, and executing the step of determining the intermediate image corresponding to the i-th diffusion cycle and the i-th image component corresponding to the intermediate image; and finally, when the diffusion processing and attenuation processing meet the iteration stopping condition, forming the noisy image set based on the first noisy image and the second noisy image.

[0235] Step S1008: Randomly sample a first target noisy image from the noisy image set, and perform diffusion processing on the original image based on the first target noisy image to obtain a second target noisy image.

[0236] Step S1010: Use the first target noisy image and the second target noisy image as the noisy images.

[0237] Step S1012: Obtain the initial generated model.

[0238] Step S1014: Input the noisy image into the initial generation model for processing to obtain the restored image.

[0239] Step S1016: Adjust the parameters of the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

[0240] In summary, to achieve rapid and efficient model training while reducing model fitting difficulty, a diffusion process can be performed on the original image after acquisition. Simultaneously, the image components of the original image are attenuated during the diffusion process, resulting in a noisy image set that has undergone diffusion and dimensionality reduction. Subsequently, noisy images suitable for model training are selected from the original and noisy image sets and input into the initial generative model for processing to obtain the restored image. Finally, the initial generative model is parameter-tuned based on the original and restored images until the target generative model that meets the training stopping condition is obtained. This achieves the goal of training a generative model through a dimensionality-variable diffusion process, thereby improving training speed while reducing model fitting difficulty, and ultimately enabling rapid and efficient training of a high-precision generative model.

[0241] Figure 11 A structural block diagram of a computing device 1100 according to one embodiment of this specification is shown. The components of the computing device 1100 include, but are not limited to, a memory 1110 and a processor 1120. The processor 1120 is connected to the memory 1110 via a bus 1130, and a database 1150 is used to store data.

[0242] The computing device 1100 also includes an access device 1140, which enables the computing device 1100 to communicate via one or more networks 1160. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1140 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0243] In one embodiment of this application, the aforementioned components of the computing device 1100 and Figure 11 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 11 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0244] The computing device 1100 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1100 can also be a mobile or stationary server.

[0245] The processor 1120 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-described generative model training method and image processing method.

[0246] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solutions of the generative model training method and the image processing method described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the generative model training method and the image processing method described above.

[0247] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described generative model training method and image processing method.

[0248] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solutions of the generative model training method and the image processing method described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the generative model training method and the image processing method described above.

[0249] An embodiment of this specification also provides a computer program, wherein when the computer program is executed in a computer, it causes the computer to perform the steps of the above-described generative model training method and image processing method.

[0250] The above is an illustrative scheme of a computer program according to this embodiment. It should be noted that the technical solution of this computer program belongs to the same concept as the technical solutions of the above-described generative model training method and image processing method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the above-described generative model training method and image processing method.

[0251] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0252] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added to or subtracted according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0253] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0254] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0255] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A generative model training method, wherein the generative model is a machine learning model, comprising: Obtain the original image; The original image is subjected to diffusion processing, and the image components of the original image are subjected to attenuation processing to obtain a noisy image set, wherein the image components are obtained by orthogonal decomposition processing of the original image, and the noisy image set is composed of the processing results of the original image after diffusion processing and attenuation processing in each diffusion cycle; The noisy image is determined based on the original image and the set of noisy images, and the noisy image is input into the initial generation model for processing to obtain the restored image; The initial generative model is tuned based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

2. The method according to claim 1, wherein the diffusion processing of the original image and the attenuation processing of the image components of the original image to obtain a noisy image set comprises: The pixel values ​​corresponding to the pixels in the original image are normalized to obtain an intermediate image, and the intermediate image is orthogonally decomposed to obtain image components. The intermediate image is subjected to diffusion processing, and the image components are subjected to attenuation processing to obtain the noisy image set.

3. The method according to claim 2, wherein performing diffusion processing on the intermediate image and attenuation processing on the image components to obtain the noisy image set comprises: Determine the intermediate image corresponding to the i-th diffusion cycle, and the i-th image component corresponding to the intermediate image; Add i-th noise to the intermediate image and perform attenuation processing on the i-th image component; If the attenuation result of the i-th image component is not less than the component threshold, a first noisy image with the same image dimension as the intermediate image is determined according to the diffusion processing result. The first noisy image is used as the intermediate image, i is incremented by 1, and the steps of determining the intermediate image corresponding to the i-th diffusion period and the i-th image component corresponding to the intermediate image are executed. If the attenuation result of the i-th image component is less than the component threshold, a second noisy image with an image dimension smaller than the intermediate image is determined according to the diffusion processing result. The second noisy image is used as the intermediate image, i is incremented by 1, and the steps of determining the intermediate image corresponding to the i-th diffusion period and the i-th image component corresponding to the intermediate image are executed. The noisy image set is formed based on the first noisy image and the second noisy image until the diffusion processing and attenuation processing meet the iteration stopping condition.

4. The method according to any one of claims 1-3, wherein determining the noisy image based on the original image and the set of noisy images comprises: A first target noisy image is randomly sampled from the set of noisy images, and the original image is subjected to diffusion processing based on the first target noisy image to obtain a second target noisy image; The first target denoised image and the second target denoised image are used as the denoised images.

5. The method according to any one of claims 1-3, wherein tuning the initial generative model based on the restored image and the original image until a target generative model satisfying the training stopping condition is obtained, comprises: Calculate the model loss value corresponding to the initial generation model based on the restored image and the original image; If the model loss value is less than a preset loss value threshold, the initial generated model is determined to meet the training stopping condition, and the initial generated model is used as the target generated model.

6. A generative model training apparatus, comprising: The acquisition module is configured to acquire the raw image; The processing module is configured to perform diffusion processing on the original image and attenuation processing on the image components of the original image to obtain a noisy image set, wherein the image components are obtained by orthogonal decomposition processing of the original image, and the noisy image set is composed of the processing results of the original image after diffusion processing and attenuation processing in each diffusion cycle. The determination module is configured to determine a noisy image based on the original image and the set of noisy images, and input the noisy image into the initial generation model for processing to obtain the restored image; The training module is configured to tune the initial generative model based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained.

7. An image processing method, comprising: Acquire the image to be processed uploaded by the user terminal; The image to be processed is input into the target generation model in any one of claims 1-5 for processing to obtain the target image; The target object is determined based on the target image, and the object information corresponding to the target object is loaded. The object information is sent to the user terminal.

8. The method according to claim 7, wherein inputting the image to be processed into a target generation model for processing to obtain a target image includes: The image to be processed is input into the target generation model; The image to be processed is subjected to dimensionality upscaling to obtain an intermediate noisy image; The intermediate noisy image is subjected to noise compensation processing to obtain the target image and output the target generation model.

9. The method according to claim 8, wherein performing noise compensation processing on the intermediate noisy image to obtain the target image comprises: Determine the image loss information corresponding to the intermediate noisy image; The intermediate noisy image is subjected to noise compensation processing based on the image loss information to obtain the target image.

10. An image processing method, comprising: Obtain the initial shopping search image uploaded by the user's terminal; The initial shopping search image is input into the target generation model in any one of claims 1-5 for processing to obtain the target shopping search image corresponding to the initial shopping search image; Based on the target shopping search image, related products are determined, and the product information corresponding to the related products is loaded; The product information is sent to the user terminal, whereby the user terminal generates and displays a product recommendation interface based on the product information.

11. A model generation training method applied to a server, wherein, The generative model is a machine learning model, including: Receive the original images uploaded by the model requester; The original image is subjected to diffusion processing, and the image components of the original image are subjected to attenuation processing to obtain a noisy image set, wherein the image components are obtained by orthogonal decomposition processing of the original image, and the noisy image set is composed of the processing results of the original image after diffusion processing and attenuation processing in each diffusion cycle; The noisy image is determined based on the original image and the set of noisy images, and the noisy image is input into the initial generation model for processing to obtain the restored image; The initial generative model is tuned based on the restored image and the original image until a target generative model that meets the training stopping condition is obtained. Determine the model parameters corresponding to the target generation model and feed the model parameters back to the model demand side.

12. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 5, 7 to 9, 10, or 11.

13. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 5, 7 to 9, 10, or 11.