Image processing method and device, equipment, storage medium and program product

By combining image processing networks and gating networks, and utilizing the combination of differential training and gating networks, we have achieved collaborative processing of multiple beautification functions, overcoming the limitations of existing technologies in terms of ease of use, personalized effects, and naturalness, and improving the user experience.

CN121353098APending Publication Date: 2026-01-16JINGDONG CITY BEIJING DIGITS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511493123.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing portrait enhancement technologies have limitations in terms of ease of use, personalized effects, naturalness, overall harmony, and user control. Users either need to invest a lot of time and energy to learn complex operations, or they can only accept automated processing results that are monotonous, lack personality, or even distorted.

Method used

An image processing method is adopted, which utilizes a network model including an image processing network, a gating network, and a set of fine-tuning parameter information. The parameters of the image processing network are adjusted based on prompts input by the user, so as to achieve the coordinated processing of multiple beautification functions, including whitening, eye enlargement, and skin smoothing. By combining differential training and gating networks, a target image that meets the user's needs is generated.

Benefits of technology

It achieves easy operation, highly personalized and natural image enhancement, improves user experience, solves the conflict and degradation problems of multiple beauty effects, and provides a good user control experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353098A_ABST
    Figure CN121353098A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and apparatus, a device, a storage medium and a program product. The method comprises the steps of obtaining first prompt information and a first image input by a user; the first prompt information is used for describing a processing mode for the first image; inputting the first prompt information and the first image into a first network model, and processing the first image by the first network model according to the first prompt information to obtain a target image corresponding to the first prompt information; wherein the first network model comprises an image processing network, a gating network and a fine tuning parameter information set, the image processing network is used for processing the first image according to the first prompt information, and the gating network is used for selecting fine tuning parameter information corresponding to the first prompt information from the fine tuning parameter information set, the fine tuning parameter information is used for adjusting parameters of the image processing network.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and particularly relates to an image processing method and device, equipment, a storage medium and a program product. BACKGROUND

[0002] In the field of digital image processing, portrait beautification has become one of the core requirements of image editing software. The existing technology mainly focuses on two types of solutions, namely, a manual fine adjustment tool and an automatic filter / preset filter. However, users need to invest a lot of time and effort to learn complex operations, or can only accept the automatic processing results which are single, lack of personality and even distorted. SUMMARY

[0003] To solve the above technical problems, the present application provides an image processing method and device, equipment, a storage medium and a program product.

[0004] The image processing method provided by the present application comprises the following steps: obtaining first prompt information input by a user and a first image; the first prompt information is used to describe a processing mode for the first image; inputting the first prompt information and the first image into a first network model; the first network model processes the first image according to the first prompt information to obtain a target image corresponding to the first prompt information; The first network model comprises an image processing network, a gating network and a set of fine-tuning parameter information, the image processing network is used to process the first image according to the first prompt information, the gating network is used to select fine-tuning parameter information corresponding to the first prompt information from the set of fine-tuning parameter information, and the fine-tuning parameter information is used to adjust parameters of the image processing network.

[0005] The image processing device provided by the present application comprises the following steps: an acquisition unit, configured to acquire first prompt information input by a user and a first image; the first prompt information is used to describe a processing mode for the first image; a processing unit, configured to input the first prompt information and the first image into a first network model; the first network model processes the first image according to the first prompt information to obtain a target image corresponding to the first prompt information; The first network model comprises an image processing network, a gating network and a set of fine-tuning parameter information, the image processing network is used to process the first image according to the first prompt information, the gating network is used to select fine-tuning parameter information corresponding to the first prompt information from the set of fine-tuning parameter information, and the fine-tuning parameter information is used to adjust parameters of the image processing network.

[0006] The image processing apparatus provided in this application includes: a processor and a memory for storing a computer program capable of running on the processor, wherein the processor executes the image processing method described above when running the computer program.

[0007] The computer-readable storage medium provided in this application is used to store a computer program that causes a computer to perform the above-described image processing method.

[0008] This application provides a computer program product, comprising: a computer program that, when executed by a processor, implements the above-described image processing method.

[0009] In the technical solution of this application, a first prompt message and a first image input by the user are obtained; the first prompt message describes the processing method for the first image; the first prompt message and the first image are input into a first network model, and the first network model processes the first image according to the first prompt message to obtain a target image corresponding to the first prompt message; wherein, the first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information, the image processing network is used to process the first image according to the first prompt message, the gating network is used to select fine-tuning parameter information corresponding to the first prompt message from the set of fine-tuning parameter information, and the fine-tuning parameter information is used to adjust the parameters of the image processing network. Thus, the gating network in the first network model can select the corresponding fine-tuning parameter information from the set of fine-tuning parameter information according to the first prompt message of the first user, adjust the parameters of the first network model, and then process the first image to generate a target image corresponding to the first prompt message, thereby meeting the user's needs and improving the user experience. Attached Figure Description

[0010] The accompanying drawings, which are provided to further illustrate this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application.

[0011] Figure 1 This is a schematic flowchart of the image processing method provided in the embodiments of this application; Figure 2 This is a schematic diagram of the overall structure provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structural composition of the image processing apparatus provided in the embodiments of this application; Figure 4 This is a schematic structural diagram of an image processing device provided in an embodiment of this application; Figure 5 This is a schematic structural diagram of the chip according to an embodiment of this application. Detailed Implementation

[0012] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0013] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0014] It should also be noted that the terms "first," "second," and "third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first," "second," and "third" can be interchanged in a specific order or sequence where permissible, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship. It should also be understood that the "instruction" mentioned in the embodiments of this application can be a direct instruction, an indirect instruction, or an indication of an association relationship. For example, A instructing B can mean that A directly instructs B, for example, B can be obtained through A; it can also mean that A indirectly instructs B, for example, A instructs C, and B can be obtained through C; or it can mean that there is an association relationship between A and B. It should also be understood that the term "correspondence" mentioned in the embodiments of this application may indicate a direct or indirect correspondence between the two, or an association between the two, or a relationship of instruction and being instructed, configuration and being configured, etc.

[0015] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.

[0016] In the field of digital image processing, portrait enhancement has become one of the core requirements of image editing software (such as Adobe Photoshop and similar applications, hereinafter referred to as "PS software"). Existing technologies mainly revolve around two main categories of solutions: manual fine-tuning tools and automated filters / presets. Manual fine-tuning tools: Existing solutions generally offer basic color and tone adjustment tools such as Curves, Levels, Selective Color, and Hue / Saturation, as well as tools for localized blemish repair and facial contouring, such as the Spot Healing Brush, Patch Tool, and Liquify. Theoretically, users can achieve highly customized beautification effects by combining these tools. However, these tools have a high barrier to entry, often requiring users to have professional color theory knowledge and proficient software skills. The process is tedious, time-consuming, and has a steep learning curve, making it extremely unfriendly to ordinary users.

[0017] Preset Filters and One-Click Enhancement: To lower the barrier to entry, existing technologies widely employ preset filters or one-click enhancement functions. These solutions are typically based on predefined parameter combinations, such as increasing brightness, enhancing contrast, increasing the saturation of specific tones, applying soft focus effects, or simple image statistical algorithms, such as global skin smoothing. However, while such methods significantly improve efficiency, their effects are often too simplistic and formulaic, lacking the ability to personalize for specific portrait features. For example, a preset "whitening" filter may over-brighten, leading to loss of detail or making skin tone appear unnatural; a generic "skin smoothing" algorithm may fail to effectively distinguish skin texture from important details, such as facial features and hair strands, resulting in an overly plastic-like appearance or blurred details in the skin area.

[0018] AI-based automated tools: Some advanced solutions incorporate deep learning-based AI technologies, such as for automatic skin smoothing (skin retouching), facial recognition and enhancement of specific areas (e.g., eyes, teeth), and background blurring. The mainstream approach uses Convolutional Neural Networks (CNNs), for example, combining an encoder-decoder structure with a U-Net architecture to first segment the skin region, then using a Generative Adversarial Network (GAN) to repair acne scars and dark circles while preserving details such as pores and hair. These technologies have improved the intelligence and naturalness of the beautification effect to some extent. Currently, solutions based on diffusion models fine-tune using Low-Rank Adaptation (LoRA), training specific LoRA models for different beautification tasks, such as whitening, stylization, and lighting. While these methods can generate good results, they rely on prompts and struggle to handle multiple beautification tasks simultaneously, easily leading to performance degradation. However, such methods suffer from the following problems: insufficient global and local coordination, weak control over the overall aesthetic style of the image (such as lighting consistency and color harmony), and a lack of coordination and unity between various local optimization effects (such as brightening the eyes, whitening teeth, and smoothing skin), resulting in an unnatural overall appearance of the final image. Furthermore, their single-function and multi-functional nature makes them prone to conflict; a single model can only process a single beautification task, and conflicts can easily arise when multiple functions are used in synergistic beautification, leading to a degradation in the overall effect.

[0019] In summary, existing portrait retouching technologies, whether traditional Photoshop tools, preset filters, or emerging AI tools, all have limitations in terms of ease of use, personalization of effects, naturalness, overall harmony, and user control. Users either need to invest a lot of time and energy to learn complex operations, or they can only accept automated processing results that are monotonous, lack personality, or even distorted. Considering these issues, while existing technologies can meet the needs of some portrait retouching tasks, there is still room for improvement in terms of ease of use, task coupling, workflow complexity, and overall effect.

[0020] Therefore, how to efficiently meet users' needs for image processing and improve user experience has become a problem that image processing needs to consider. To this end, the following technical solutions are proposed according to embodiments of this application.

[0021] To facilitate understanding of the technical solutions of the embodiments of this application, the technical solutions of this application are described in detail below through specific embodiments. The above-mentioned related technologies are optional solutions and can be arbitrarily combined with the technical solutions of the embodiments of this application, all of which fall within the protection scope of the embodiments of this application. The embodiments of this application include at least some of the following contents.

[0022] Figure 1 This is a schematic flowchart of the image processing method provided in the embodiments of this application, as shown below. Figure 1As shown, the image processing method includes the following steps: Step 101: Obtain the first prompt information and the first image input by the user; the first prompt information is used to describe the processing method for the first image.

[0023] Step 102: Input the first prompt information and the first image into the first network model. The first network model processes the first image according to the first prompt information to obtain the target image corresponding to the first prompt information.

[0024] In some implementations, a first network model is used to process a first image based on a first prompt message input by the user and a first image to obtain a target image corresponding to the first prompt message. The first prompt message can be a prompt, used to describe one or more processing methods for the first image.

[0025] In some implementations, the first network model is used for image beautification. The first prompt information can be the user's beautification request. For example, the first prompt information can be whitening. The first network model whitens the first image according to the user's whitening request in the first prompt information, that is, it performs whitening processing on the first image. Alternatively, the first prompt information can be whitening and eye enlargement. The first network model whitens the target part in the first image and enlarges the eyes in the first image according to the user's whitening and eye enlargement requests in the first prompt information, that is, it performs whitening processing and eye enlargement processing on the first image.

[0026] It is understood that when the first network is used for image beautification, the first image is an image including a human face. It should also be understood that the first prompt information can be a variety of beautification methods, such as individual beautification functions like whitening, eye enlargement, and skin smoothing, or a combination of multiple beautification functions. The specific beautification method can be determined according to the actual situation, and this application does not make specific limitations in this regard.

[0027] It should be noted that the first network model in this application embodiment can also be used in other application scenarios, such as adding filters to images. This application does not specifically limit the specific application scenarios.

[0028] In some implementations, the first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information. The image processing network is used to process a first image according to a first prompt, and the gating network is used to select fine-tuning parameter information corresponding to the first prompt from the set of fine-tuning parameter information. The fine-tuning parameter information is used to adjust the parameters of the image processing network.

[0029] In some implementations, the image processing network is a diffusion (DiT) model, and the fine-tuning parameter information set is the fine-tuning parameter information of each layer in the image processing network. The gating network is used to select the fine-tuning parameter information corresponding to the first prompt information from the fine-tuning parameter information set, and use the fine-tuning parameter information to adjust the parameters of each layer in the corresponding image processing network to meet the user's needs.

[0030] For example, when the first prompt is "whitening," the gating network selects whitening-related fine-tuning parameters from the set of fine-tuning parameters and adjusts the parameters in the image processing network accordingly to meet the user's whitening needs. When the first prompt is "whitening and eye enlargement," the gating network selects whitening-related and eye-enlargement-related fine-tuning parameters from the set of fine-tuning parameters and adjusts the parameters in the image processing network accordingly to meet the user's whitening and eye-enlargement needs.

[0031] It is understandable that the set of fine-tuning parameter information refers to the fine-tuning of parameters in each layer of the image processing network, or the fine-tuning of parameters in some layers. It should also be understood that the image processing network is a network model used for image processing in the current application scenario. For example, if the current application scenario is beautification, then the image processing network itself is an image processing network used to implement the beautification function.

[0032] In some implementations, before using the first network model, it is necessary to train the first network model. The method further includes: obtaining a first data set, training an initial network model based on the first data set, and obtaining the first network model. The initial network model includes an image processing network and an initial gating network. Here, the first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information. Therefore, in training the initial network model using the first data set, it is sufficient to obtain the set of fine-tuning parameter information and train the gating network. The parameters of the image processing network are not changed during the training process.

[0033] In some implementations, training an initial network model based on a first dataset to obtain the first network model includes: determining a set of fine-tuning parameter information based on the first dataset, an image processing network, and a first tool; the set of fine-tuning parameter information corresponds to N processing methods, where N processing methods include the processing methods described in the first prompt information, and N is a positive integer; determining a second dataset based on the first dataset, the set of fine-tuning parameter information, and the image processing network; and training an initial gating network based on the second dataset to obtain the gating network. Here, obtaining the first network model includes two stages: the first stage is used to obtain the set of fine-tuning parameter information, and the second stage is used to obtain the gating network.

[0034] In some implementations, when the first network model is used for beautification, the first dataset is a first image dataset. A set of fine-tuning parameter information is determined using the first image dataset, the image processing network, and the first tool. This set of fine-tuning parameter information is used for parameter fine-tuning in the beautification function. The set of fine-tuning parameter information corresponds to N processing methods; that is, it includes fine-tuning parameter information for each processing method, and the N processing methods include the processing method described in the first prompt message. For example, when the first prompt message indicates whitening of the first image, the N processing methods include a whitening method, and the set of fine-tuning parameter information includes the fine-tuning parameter information corresponding to the whitening method; when the first prompt message indicates whitening and eye enlargement of the first image, the N processing methods include a whitening method and an eye enlargement method, and the set of fine-tuning parameter information includes the fine-tuning parameter information corresponding to the whitening method and the fine-tuning parameter information corresponding to the eye enlargement method.

[0035] In some implementations, when fine-tuning the parameters of the image processing network, the initial weight information and the fine-tuning parameter information are used to weight the corresponding parameters. The specific preset weight information is not specifically limited in this application and can be determined according to the actual situation. For example, when the first prompt information includes one processing method, the initial weight can be set to 1. When the first prompt information includes multiple processing methods, the initial weights of all processing methods can be set to 1, or the sum of these initial weights can be set to 1, with equal proportions. For example, when there are two processing methods, the corresponding initial weight information can be set to 0.5.

[0036] In some implementations, when the first network model is used for beautification, it can perform N processing methods on the image; for example, N is 32. The number of processing methods can be determined based on actual circumstances and is not specifically limited here. It is understandable that N processing methods refer to N separate processing methods for the image.

[0037] In some implementations, the first data set includes a first image set and a second image set, where the second image set is the set of images obtained by processing the images in the first image set through N processing methods. Here, the first image set is the original image, and the second image set is the set of images obtained by processing the original images in the first image set through N processing methods.

[0038] In some implementations, if the first network model is used for beautification, then the N processing methods correspond to the N beautification functions. The first image processing set includes the original image for each beautification function, that is, the first image set includes N sets of original images, and the second image set includes N sets of beautified images after the N sets of original images have been processed by the corresponding N processing methods.

[0039] For example, the training set (equivalent to the aforementioned first dataset) consists of training pairs composed of the original image to be beautified (OriginalImage, equivalent to the aforementioned first image set) and the beautified image (EditedImage, equivalent to the aforementioned second image set). Taking functions A = whitening, B = enlarging eyes, and C = smoothing skin as examples (functions equivalent to the aforementioned processing methods), two sets of data are prepared for each function: the image before the change and the image after the change. The image pairs differ only in terms of the target function. For example, when adjusting for enlarging eyes, skin tone and other aspects are completely consistent, and the prompt word is "enlarging eyes". The number of images M in each group is greater than 100. For example, the dataset for function A: [original image, whitened skin tone image] * M; the dataset for function B: [original image, enlarging eye image] * M; the dataset for function C: [original image, smoothed skin image] * M. That is, the original image in the dataset belongs to the first image set, and the whitened skin tone image, enlarging eye image, and smoothed skin image in the dataset belong to the second image set.

[0040] It is understandable that when obtaining a beautified image from the original image, a corresponding prompt word needs to be entered. Specifically, DiT from this application can be used as the network model for obtaining the beautified image. This application does not make any specific limitations on this. It should also be understood that the original image corresponding to different processing methods is not the same.

[0041] In some implementations, determining a set of fine-tuning parameter information based on a first data set, an image processing network, and a first tool includes: determining a first set of fine-tuning parameter information based on a first image set, an image processing network, and a first tool; determining a second set of fine-tuning parameter information based on a second image set, an image processing network, and a first tool; and determining a set of fine-tuning parameter information based on the first set of fine-tuning parameter information and the second set of fine-tuning parameter information.

[0042] In some implementations, the fine-tuning parameter information is LoRA, and the first tool is a training optimizer used to obtain the set of fine-tuning parameter information. The first set of fine-tuning parameter information is determined based on the first image set, the image processing network, and the first tool; that is, the set of fine-tuning parameter information corresponding to the image before the change is obtained. The second fine-tuning parameter information set is determined based on the second image set, the image processing network, and the first tool, that is, the information corresponding to the changed image is obtained. The fine-tuning parameter information set is determined based on the first fine-tuning parameter information set and the second fine-tuning parameter information set. Here, the first fine-tuning parameter information set includes fine-tuning parameter information corresponding to N processing methods, and the second fine-tuning parameter information set includes fine-tuning parameter information corresponding to the corresponding N processing methods. For example, if the first network model is used for beautification of a first image, and if the N processing methods include whitening, then the first and second fine-tuning parameter information sets include LoRA for whitening.

[0043] It is understandable that for each processing method, corresponding fine-tuning parameter information will be obtained to fine-tune the parameters of the image processing network. Therefore, each LoRA corresponds to an array, which includes parameters for fine-tuning one or more layers of the image processing network. The specific number of network layers to be adjusted can be determined according to the actual situation, and this application does not impose a specific limitation on this.

[0044] In some implementations, determining the fine-tuning parameter information set based on the first fine-tuning parameter information set and the second fine-tuning parameter information set includes: subtracting the second fine-tuning parameter information set from the first fine-tuning parameter information set to determine the fine-tuning parameter information set.

[0045] Specifically, for each independent processing method, two overfitted LoRAs are trained, before the change. After the change After training is complete, the expert LoRA (equivalent to the fine-tuning parameter information in the aforementioned fine-tuning parameter information set) is obtained by subtracting the weights. The specific method of obtaining it is shown in equation (1): (1); In some implementations, the first tool is a training optimizer, using Adafactor, to obtain the learning rate of 5e-5, LoRA Rank of 32, and training epochs of 500 from the fine-tuning parameter information set. Specific parameter settings can be determined according to actual circumstances, and this application does not impose specific limitations on them.

[0046] In some implementations, after obtaining the set of fine-tuning parameter information, a second set of data for training the initial gating network can be determined using the first data set, the set of fine-tuning parameter information, and the image processing network; the initial gating network is then trained based on the second data set to obtain the gating network. The gating network is used to select the corresponding fine-tuning parameter information based on the user's first prompt.

[0047] In some implementations, determining a second data set based on a first data set, a fine-tuning parameter information set, and an image processing network includes: determining a third image set based on second prompt information, a first image set, the fine-tuning parameter information set, and the image processing network; the second prompt information includes one or more processing methods for the image; and determining the second data set based on the first image set, the third image set, and the second prompt information. Here, the dataset used to train the initial gating network includes the original image (equivalent to the aforementioned first image set), the image after applying one or more processing methods (equivalent to the third data set), and the user instruction (equivalent to the aforementioned second prompt information).

[0048] In some implementations, the first image set in the first dataset and the user's second prompt information are input into the image processing network. At the same time, the parameters of the image processing network are adjusted using the fine-tuning parameter information corresponding to the processing method included in the second prompt information to obtain a third image set corresponding to the second prompt information.

[0049] Specifically, if the first network model is used for image beautification, the first image set is the original image, and the second prompt information is the processing method for the first image set, which may include one or more beautification functions. The first image set and the second prompt information are input into the image processing network, and the fine-tuning parameter information corresponding to the processing method included in the second prompt information is selected from the fine-tuning parameter information set to adjust the image processing network, resulting in a third image set. For example, for the aforementioned training set, i.e., the first image set, a combination filter effect is randomly selected for each training original image, i.e., processing methods such as enlarging eyes, whitening, and skin smoothing are randomly selected. Usually, no more than three processing methods are used. The original image is then fed into the image processing network to generate the target image (equivalent to the third image set), i.e., the combined effect is applied. Then, the generated pair dataset (equivalent to the aforementioned second dataset) is fed into the model for training.

[0050] In some implementations, the initial gating network is a newly added network within the image processing model. Therefore, when training the initial gating network using the second dataset, the image processing model and the fine-tuning parameter information set are frozen. That is, during the training of the initial gating network, the parameters of the image processing model are not updated, and the fine-tuning parameter information set is also not updated. The training objective of the initial gating network is: selected fine-tuning parameter information is 1, and the rest are 0. For example, if the second prompt message indicates a whitening function, the gating network selects the fine-tuning parameter information corresponding to the whitening function, and the corresponding output is 1; the other unselected fine-tuning parameter information has a corresponding output of 0. Here, the model fed into the second dataset is the first network model, the foundation of which is the image processing network, and the gating network is a newly added part of the image processing network.

[0051] In some implementations, during the second stage, i.e., during the training of the initial gating network to obtain the gating network, the training loss includes the loss of the image processing network and the loss of the initial gating network.

[0052] Specifically, the training loss is calculated as shown in equation (2): (2); in, This is the flow-matching loss, also known as the DiT loss of the image processing network. w is a hyperparameter set to 0.01. The initial training loss for the gated network is given by equation (3): (3); in, It is an importance calculation that measures the uniformity of the gating output probability distribution of N processing methods, and the average probability. Specifically, as shown in equation (4): (4); in, In Let be the average importance score of each processing method i in each batch during training, which is the average gating probability of the samples in the batch. The gating probability is the probability of the output of the gating network. To measure the actual number of times each processing method is routed, i.e., the frequency of being selected as a Top-K expert, the goal is to ensure that the average load across all processing methods is roughly the same. The calculation formula is... Same, but at this time This represents the average number of times each processing method i is actually routed across the samples in the entire batch. The number of samples in each batch can be determined based on actual circumstances, and this application does not impose a specific limitation on this.

[0053] In some implementations, the first network model processes the first image according to the first prompt information to obtain a target image corresponding to the first prompt information, including: the gating network selects a third fine-tuning parameter information set from the fine-tuning parameter information set according to the first prompt information, the third fine-tuning parameter information set corresponding to the processing method in the first prompt information; and determines the target image according to the third fine-tuning parameter information set, the image processing network, and the first image.

[0054] In some implementations, if the first network model is used for beautification, then the processing method in the first prompt message is the beautification method. After the first image and the first prompt message are input into the first network model, the gating network selects the corresponding fine-tuning parameter information from the fine-tuning parameter information according to the processing method in the first prompt message, that is, the third fine-tuning parameter information set, and modifies the parameters in the image processing network accordingly, thereby outputting the target image corresponding to the first image.

[0055] In some implementations, the image processing network includes a P layer, for which a set of fine-tuning parameter information is trained on the Q layer. This Q layer also has a corresponding gating network. When the first image and the first prompt information are input, if the image passes through the Q layer, the gating network selects the corresponding fine-tuning parameter information to fine-tune the parameters of the current layer to meet the user's needs, ultimately outputting the target image. Here, P and Q are positive integers, and Q is less than or equal to P.

[0056] It is understandable that the fine-tuning parameter information corresponding to each processing method in the fine-tuning parameter information set corresponds to some or all layers in the image processing network. That is, the fine-tuning parameter information is not a single data point, but a collection of fine-tuning information of the parameters of each layer corresponding to the same processing method.

[0057] Specifically, the user inputs a prompt (equivalent to the first prompt message), which consists of specific beautification functions, such as "bigger eyes," "whitening," and "skin smoothing." The user inputs the original image to be processed (equivalent to the first image), and the user prompt is tokenized using a text tokenizer. The image is encoded using a VAE and tokenized into image tokens. The text and image latent are horizontally concatenated and fed into the first network model along with random noise for N rounds of denoising. The denoised latent is then decoded using a VAE to obtain the final image (equivalent to the target image).

[0058] In some embodiments, the method further includes: obtaining first weight information, wherein the first weight information is the weight information of the fine-tuning parameters corresponding to the processing method in the first prompt information; and updating the target image according to the first weight information and the fine-tuning parameters.

[0059] In some implementations, after acquiring the target image, if the user wants to adjust the initial weight information of a processing method in the first prompt information, the user can determine the first weight information based on the target image. After the first network model acquires the first weight information, it adjusts the initial weight information of the fine-tuning parameter information based on the first weight information and updates the target image.

[0060] In some implementations, when a user selects only one processing method, the default initial weight information is 1. The user can adjust this weight value, that is, obtain the first weight information, and change the functional strength of the current processing method. For example, the initial weight information 1 can be adjusted to the first weight information 2 or 0.5 to enhance or weaken the strength of the current processing method.

[0061] In some implementations, when a user selects multiple processing methods, the sum of the default initial weight information corresponding to each processing method is 1. The weights are adjusted using the first weight information to change the functional strength of the processing method. Here, the sum of the adjusted weights is still 1. For example, when the user selects two processing methods, the initial weight information for each method is 0.5, which can be adjusted to the first weight information, such as 0.25 and 0.75. Alternatively, when a user selects multiple processing methods, the default initial weight information for each method is 1. The functional strength of the processing method is changed by adjusting the weights using the first weight information. Here, the adjusted weights can be less than or greater than 1. For example, when the user selects two processing methods, the initial weight information for each method is 1, which can be adjusted to the first weight information, such as 0.5 and 2.

[0062] The technical solution of this application embodiment obtains a first prompt message and a first image input by a user; the first prompt message describes the processing method for the first image; the first prompt message and the first image are input into a first network model, and the first network model processes the first image according to the first prompt message to obtain a target image corresponding to the first prompt message; wherein, the first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information, the image processing network is used to process the first image according to the first prompt message, the gating network is used to select fine-tuning parameter information corresponding to the first prompt message from the set of fine-tuning parameter information, and the fine-tuning parameter information is used to adjust the parameters of the image processing network. In this way, the gating network in the first network model can select the corresponding fine-tuning parameter information from the set of fine-tuning parameter information according to the first prompt message of the first user, adjust the parameters of the first network model, and then process the first image to generate a target image corresponding to the first prompt message, thereby meeting the user's needs and improving the user experience.

[0063] The technical solutions of the embodiments of this application will be further explained below with reference to specific application examples.

[0064] Based on the foregoing embodiments, when the first network model is used for beautification function, the technical solution of the embodiments of this application will be further described.

[0065] Existing portrait beautification technologies (whether traditional Photoshop tools, preset filters, or emerging AI tools) all have limitations in terms of ease of use, personalized effects, naturalness, overall harmony, and user control. Users either need to invest a lot of time and energy to learn complex operations, or they can only accept automated processing results that are monotonous, lack personality, or even distorted. A new portrait beautification technology solution that can balance ease of operation, highly personalized effects, and natural harmony, while providing a good user control experience, is needed. This application aims to solve one or more of the above-mentioned technical problems. Addressing the problems of the above-mentioned solutions, this application proposes a combined reasoning-based portrait beautification method specifically for portrait beautification, aiming to solve: the difficulty of parameter adjustment in portrait beautification, the problem of global and local inconsistency, and the problem of effect conflict and degradation from multiple beautification effects.

[0066] Figure 2 This is a schematic diagram of the overall structure provided in an embodiment of this application. The training part of this method is divided into two stages, realizing the functions of multi-expert division of labor, hybrid expert routing, and dynamic combination during model inference. Figure 2 As shown, the first stage is differential LoRA training, which uses the original image and the beautified image to form a dataset and generates the image before and after the beautification. After the change The first stage involves obtaining expert LoRAs for each beautification function through subtraction. The second stage is the training of the gating network. Inputting the original image, the edited image, and user instructions, the LoRA injection layer is frozen (i.e., the LoRA and noise prediction layers are frozen), and only the gating network is trained. The gating network selects from N expert LoRAs and integrates them into the parameters of the DiT network. Simultaneously, the diagram shown in the second stage can also represent the processing procedure for the trained network; inputting the original image and user instructions, it can output the edited image. The specific meaning and details of each stage are as follows: Phase One: Figure 2 Stage 1. Basic expert LoRA training is performed on common beautification functions using differential training. The training set consists of original images (equivalent to the first image set mentioned above) and edited images (equivalent to the second image set mentioned above), forming training pairs. For example, functions A = whitening, B = enlarging eyes, and C = skin smoothing are used. Two sets of data are prepared for each function: data before and data after the change. The image pairs differ only in the target function; for example, when adjusting "enlarging eyes," skin tone and other aspects are completely consistent, and the prompt is "enlarging eyes." The number of images M in each group is greater than 100. For example, the dataset for function A: [original image, whitened skin tone image] * M; the dataset for function B: [original image, enlarging eyes image] * M. For each independent function, two overfitted LoRA models are trained based on the DiT model. The original images in the dataset are used to train and obtain the image before the change. The transformed images are trained using the beautified images in the dataset. After training is complete, the expert LoRA (equivalent to the aforementioned fine-tuning parameter information) is obtained by subtracting the weights, as shown in Equation (1). The training optimizer uses Adafactor with a learning rate of 5e-5. The LoRA Rank is 32, and the training rounds are 500. It can be understood that each beautification function will have an expert LoRA, which includes multiple data points and corresponds to some or all layers in the DiT network. That is, information for parameter fine-tuning will be trained for some or all layers. For example, if the DiT network includes 50 layers, then the LoRA for each beautification function can be trained for some or all layers. It should also be understood that the parameters of the DiT model will not be modified during the process of obtaining the expert LoRA. It should also be understood that the original images corresponding to each function are not the same in this stage.

[0067] Phase Two: Figure 2 Stage 2nd in the model trains a gating network that automatically activates experts, allowing the model to dynamically select beautification experts (i.e., beautification functions) based on user commands. First, corresponding training samples need to be automatically generated. For the training set from Stage 1, a combination of filter effects (i.e., randomly selecting effects like enlarging eyes, whitening, and skin smoothing) is randomly selected for each original training image, with a maximum of three filters. The image is then fed into the filter to generate the target image (applying the combined effects). Specifically, using the LoRA trained in Stage 1, all original images from Stage 1, along with user commands for no more than three beautification functions, are input into the DiT network. The LoRA corresponding to the beautification function in the user command is applied to the DiT network to generate the corresponding beautified image. Based on the applied filter effects, potential gating targets are determined: activating an expert = 1, others = 0. That is, the selected beautification expert outputs 1, and the unselected beautification expert outputs 0. The generated pair dataset is then fed into the model for training. Here, the pair dataset includes the original image, the edited image, and user instructions. The model is a DiT-based network with added gating networks and expert LoRA (equivalent to the aforementioned set of fine-tuning parameter information). The training result is the first network model mentioned above. It can be understood that the gating network is related to the layers containing fine-tuning parameter information. When a layer has corresponding fine-tuning parameter information, there will also be a corresponding gating network for selecting experts.

[0068] Phase 2 Training Details: The gating network structure is a 3-layer Multilayer Perceptron (MLP), where the input of the last layer is 32, meaning it supports a maximum of 32 experts. During training, the LoRA of the experts is frozen, and only the gating network is trained. The training loss is calculated as shown in Equation (2), where... For flow-matching loss, w is a hyperparameter set to 0.01. The training loss for Moe is shown in Equation (3). It is an importance calculation that measures the uniformity (average probability) of the gating output probability distribution of beauty experts. N is the number of experts, and its formula is shown in equation (4), where Let be the average importance score of expert i in the entire batch during training, which is the average gating probability of samples within the batch. To measure the actual number of times each expert is routed (i.e., the frequency of being selected as a Top-K expert), the goal is to ensure that the average load of all experts is roughly the same. The calculation formula is... Same, but at this time This represents the average number of times expert i was actually routed across all samples in the entire batch.

[0069] Inference phase: After the network is trained, (1) the user inputs a prompt: the prompt consists of specific beautification functions (e.g., "big eyes, whitening, skin smoothing"); the user inputs the original image to be processed; (2) the user prompt is tokenized by the text tokenizer. The image is encoded by VAE and tokenized into image tokens. (3) the text and image latent are horizontally spliced ​​together and input with random noise latent into the DiT network containing expert LoRA networks for N noise reduction. (4) the user can adjust the ratio of LoRA weights added to the original weights before noise reduction to adjust the intensity of specific beautification functions. When the user selects multiple experts, the LoRA ratio adjustment range for each expert is 0.1 to 1.0, that is, the sum of the adjusted weights is also 1, and the sum of the initial weights is also 1. When the user selects one expert, the LoRA weight of each expert is 1, which can be adjusted to a value greater than 1 or less than 1 to enhance or weaken the beautification function. (5) the latent obtained after noise reduction is decoded by VAE to obtain the final image.

[0070] This application presents a combined differential intelligent portrait beautification method that unifies the portrait beautification task. It employs a two-stage generative differential training method based on image pairs to train a loRA-MOE expert network (equivalent to the aforementioned first network model, i.e., a network based on the DiT network with the addition of expert LoRA and a gated network). This achieves adjustable combined portrait beautification functions based on text prompts. The technical solution of this application resolves the conflict between local and global adjustments in traditional beautification filters, as well as the degradation of effects from multiple filters. Furthermore, the unified approach simplifies filter settings and lowers the barrier to entry.

[0071] Figure 3 This is a schematic diagram of the structural composition of the image processing apparatus provided in the embodiments of this application, as shown below. Figure 3 As shown, the image processing apparatus includes: The acquisition unit 301 is used to acquire the first prompt information and the first image input by the user; the first prompt information is used to describe the processing method for the first image; Processing unit 302 is used to input the first prompt information and the first image into the first network model, and the first network model processes the first image according to the first prompt information to obtain the target image corresponding to the first prompt information; The first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information. The image processing network is used to process the first image according to the first prompt information. The gating network is used to select the fine-tuning parameter information corresponding to the first prompt information from the set of fine-tuning parameter information. The fine-tuning parameter information is used to adjust the parameters of the image processing network.

[0072] In some embodiments, the acquisition unit 301 is used to acquire a first data set; the processing unit 302 is used to train an initial network model based on the first data set to obtain a first network model, wherein the initial network model includes an image processing network and an initial gating network.

[0073] In some embodiments, the processing unit 302 is configured to determine a fine-tuning parameter information set based on a first data set, an image processing network, and a first tool; the fine-tuning parameter information set corresponds to N processing methods, the N processing methods include the processing methods described in the first prompt information, and N is a positive integer; determine a second data set based on the first data set, the fine-tuning parameter information set, and the image processing network; and train an initial gating network based on the second data set to obtain a gating network.

[0074] In some implementations, the first data set includes a first image set and a second image set, wherein the second image set is an image set obtained by processing the images in the first image set through N processing methods; the processing unit 302 is used to determine a first fine-tuning parameter information set based on the first image set, the image processing network, and the first tool; determine a second fine-tuning parameter information set based on the second image set, the image processing network, and the first tool; and determine a fine-tuning parameter information set based on the first fine-tuning parameter information set and the second fine-tuning parameter information set.

[0075] In some embodiments, the processing unit 302 is used to enable the gating network to select a third fine-tuning parameter information set from the fine-tuning parameter information set according to the first prompt information, wherein the third fine-tuning parameter information set corresponds to the processing method in the first prompt information; and to determine the target image according to the third fine-tuning parameter information set, the image processing network, and the first image.

[0076] In some embodiments, the acquisition unit 301 is used to acquire first weight information, which is the weight information of the fine-tuning parameters corresponding to the processing method in the first prompt information; the processing unit 302 is used to update the target image according to the first weight information and the fine-tuning parameters.

[0077] Those skilled in the art should understand that Figure 3 The functions of each unit in the image processing apparatus shown can be understood by referring to the relevant description of the aforementioned method. Figure 3 The functions of each unit in the image processing device shown can be implemented by a program running on a processor or by specific logic circuits.

[0078] Figure 4 This is a schematic structural diagram of an image processing device 400 provided in an embodiment of this application. Figure 4 The image processing device 400 shown includes a processor 410, which can call and run computer programs from memory to implement the methods in the embodiments of this application.

[0079] Optionally, such as Figure 4 As shown, the image processing device 400 may further include a memory 420. The processor 410 can retrieve and run computer programs from the memory 420 to implement the methods described in the embodiments of this application.

[0080] The memory 420 can be a separate device independent of the processor 410, or it can be integrated into the processor 410.

[0081] Optionally, such as Figure 4 As shown, the image processing device 400 may also include a transceiver 430, which the processor 410 can control to communicate with other devices. Specifically, it can send information or data to other devices or receive information or data sent by other devices.

[0082] The transceiver 430 may include a transmitter and a receiver. The transceiver 430 may further include an antenna, and the number of antennas may be one or more.

[0083] The image processing device 400 can implement the corresponding processes implemented by the image processing apparatus in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0084] Figure 5 This is a schematic structural diagram of the chip according to an embodiment of this application. Figure 5 The chip 500 shown includes a processor 510, which can call and run computer programs from memory to implement the methods in the embodiments of this application.

[0085] Optionally, such as Figure 5 As shown, chip 500 may further include memory 520. Processor 510 can retrieve and run computer programs from memory 520 to implement the methods described in this embodiment.

[0086] The memory 520 can be a separate device independent of the processor 510, or it can be integrated into the processor 510.

[0087] Optionally, the chip 500 may also include an input interface 530. The processor 510 can control the input interface 530 to communicate with other devices or chips; specifically, it can acquire information or data sent by other devices or chips.

[0088] Optionally, the chip 500 may also include an output interface 540. The processor 510 can control the output interface 540 to communicate with other devices or chips, specifically, to output information or data to other devices or chips.

[0089] This chip can implement the corresponding processes implemented by the image processing device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0090] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0091] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0092] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0093] It should be understood that the above-described memory is exemplary and not a limiting description. For example, the memory in the embodiments of this application may also be static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DR RAM), etc. That is to say, the memory in the embodiments of this application is intended to include, but is not limited to, these and any other suitable types of memory.

[0094] This application also provides a computer program product, including a computer program.

[0095] When executed by a processor, the computer program implements the corresponding processes implemented by the image processing device in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0096] This application also provides a computer-readable storage medium for storing computer programs.

[0097] The computer program causes the computer to execute the corresponding processes implemented by the image processing apparatus in the various methods of the embodiments of this application, which will not be described in detail here for the sake of brevity.

[0098] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0099] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0100] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0101] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0102] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0103] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. An image processing method, characterized in that, The method includes: Obtain the first prompt information and the first image input by the user; the first prompt information is used to describe the processing method for the first image; The first prompt information and the first image are input into the first network model. The first network model processes the first image according to the first prompt information to obtain the target image corresponding to the first prompt information. The first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information. The image processing network is used to process the first image according to the first prompt information. The gating network is used to select fine-tuning parameter information corresponding to the first prompt information from the set of fine-tuning parameter information. The fine-tuning parameter information is used to adjust the parameters of the image processing network.

2. The method according to claim 1, characterized in that, The method further includes: A first data set is obtained, and an initial network model is trained based on the first data set to obtain the first network model, wherein the initial network model includes the image processing network and the initial gating network.

3. The method according to claim 2, characterized in that, The step of training an initial network model based on the first dataset to obtain the first network model includes: The fine-tuning parameter information set is determined based on the first data set, the image processing network, and the first tool; the fine-tuning parameter information set corresponds to N processing methods, and the N processing methods include the processing methods described in the first prompt information, where N is a positive integer; A second data set is determined based on the first data set, the fine-tuning parameter information set, and the image processing network; The initial gating network is trained based on the second dataset to obtain the gating network.

4. The method according to claim 3, characterized in that, The first data set includes a first image set and a second image set, wherein the second image set is an image set obtained by processing the images in the first image set through N processing methods; Determining the fine-tuning parameter information set based on the first data set, the image processing network, and the first tool includes: A first set of fine-tuning parameter information is determined based on the first image set, the image processing network, and the first tool; A second set of fine-tuning parameter information is determined based on the second image set, the image processing network, and the first tool; The fine-tuning parameter information set is determined based on the first fine-tuning parameter information set and the second fine-tuning parameter information set.

5. The method according to claim 1, characterized in that, The first network model processes the first image according to the first prompt information to obtain a target image corresponding to the first prompt information, including: The gating network selects a third set of fine-tuning parameters from the set of fine-tuning parameters based on the first prompt information. The third set of fine-tuning parameters corresponds to the processing method in the first prompt information. The target image is determined based on the third set of fine-tuning parameters, the image processing network, and the first image.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: Obtain first weight information, which is the weight information of the fine-tuning parameter information corresponding to the processing method in the first prompt information; The target image is updated based on the first weight information and the fine-tuning parameters.

7. An image processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire first prompt information and a first image input by the user; the first prompt information is used to describe the processing method for the first image; The processing unit is used to input the first prompt information and the first image into the first network model, and the first network model processes the first image according to the first prompt information to obtain a target image corresponding to the first prompt information. The first network model includes an image processing network, a gating network, and a set of fine-tuning parameter information. The image processing network is used to process the first image according to the first prompt information. The gating network is used to select fine-tuning parameter information corresponding to the first prompt information from the set of fine-tuning parameter information. The fine-tuning parameter information is used to adjust the parameters of the image processing network.

8. An image processing device, characterized in that, include: The processor and the memory used to store computer programs that can run on the processor. When the processor is used to run the computer program, it performs the steps of the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Model fine tuning method and device, electronic equipment and nonvolatile storage medium

    CN120471129A

  • Image restoration method and device based on diffusion model and generative adversarial training

    CN120782676A