Model training methods, image processing methods and devices, media, and equipment

By fusing the mask and edge information of sample images to generate sample guidance images, and combining them with descriptive text to train the model, the problem of low accuracy of model training data is solved, and high-quality image generation and processing are achieved.

CN116935166BActive Publication Date: 2025-10-31GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311009300.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-10
Publication Date
2025-10-31
Estimated Expiration
2043-08-10

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of model training data is low, resulting in poor image quality and high limitations. Existing data factory links are inefficient and have poor practicality.

Method used

By fusing the sample mask image and sample edge image of the sample image, a sample guide image is generated. The image editing model is then trained by combining the sample description text to obtain the trained image editing model, which is used to generate the target image.

Benefits of technology

It improves the accuracy and comprehensiveness of model training, enhances the model's versatility and flexibility, improves the accuracy and quality of image processing, and increases the efficiency of data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116935166B_ABST
    Figure CN116935166B_ABST
Patent Text Reader

Abstract

This disclosure relates to a model training method and apparatus, an image processing method and apparatus, an electronic device, and a storage medium, belonging to the field of computer technology. The model training method includes: acquiring a sample image and a corresponding sample mask image; acquiring a sample edge image corresponding to the sample image, and fusing the sample edge image and the sample mask image to obtain a sample guidance image; training an image editing model using the sample image, the sample guidance image, and corresponding sample description text to obtain a trained image editing model. The technical solutions in this disclosure can improve the accuracy of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and more specifically, to a model training method and apparatus, an image processing method and apparatus, a computer-readable storage medium, and an electronic device. Background Technology

[0002] The performance of models used to process data usually requires a sufficient amount of training data for iterative training. However, in practical applications, the quality of training data is difficult to guarantee, so it is necessary to improve the accuracy of the generated data.

[0003] In related technologies, production networks and segmentation networks can be applied to data factory chains to generate large amounts of required image data based on images. However, the accuracy of the original images and segmentation data provided in these methods is relatively low, resulting in low accuracy of the trained models. Consequently, the richness of the generated images is limited, leading to both low accuracy and poor quality. Summary of the Invention

[0004] The purpose of this disclosure is to provide a model training method and apparatus, an image processing method and apparatus, a storage medium, and an electronic device, thereby overcoming, to at least a certain extent, the problem of low model accuracy caused by the limitations and defects of related technologies.

[0005] According to a first aspect of this disclosure, a model training method is provided, comprising: acquiring a sample image and a sample mask image corresponding to the sample image; acquiring a sample edge image corresponding to the sample image, and fusing the sample edge image and the sample mask image to obtain a sample guide image; and training an image editing model using the sample image, the sample guide image, and sample description text corresponding to the sample image to obtain a trained image editing model.

[0006] According to a second aspect of this disclosure, an image processing method is provided, comprising: acquiring an image to be processed and a mask image corresponding to the image to be processed; acquiring an edge image corresponding to the image to be processed, and fusing the edge image and the mask image to obtain a guide image; inputting the image to be processed, the guide image, and descriptive text corresponding to the image to be processed into a trained image editing model to obtain a target image; wherein the trained image editing model is trained according to any one of the model training methods described above.

[0007] According to a third aspect of this disclosure, a model training apparatus is provided, comprising: an image acquisition module for acquiring a sample image and a sample mask image corresponding to the sample image; a sample guidance image generation module for acquiring a sample edge image corresponding to the sample image and fusing the sample edge image with the sample mask image to obtain a sample guidance image; and a guidance training module for training an image editing model using the sample image, the sample guidance image, and sample description text corresponding to the sample image to obtain a trained image editing model.

[0008] According to a fourth aspect of this disclosure, an image processing apparatus is provided, comprising: an image acquisition module for acquiring an image to be processed and a mask image corresponding to the image to be processed; a guide image determination module for acquiring an edge image corresponding to the image to be processed and fusing the edge image with the mask image to obtain a guide image; and an image generation module for inputting the image to be processed, the guide image, and descriptive text corresponding to the image to be processed into a trained image editing model to obtain a target image; wherein the trained image editing model is trained according to the model training method described in any one of the above claims.

[0009] According to a fifth aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the model training method of the first aspect and the image processing method of the second aspect, and possible implementations thereof, by executing the executable instructions.

[0010] According to a sixth aspect of this disclosure, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the model training method of the first aspect and the image processing method of the second aspect, and their possible implementations.

[0011] In the technical solution provided in this disclosure, on the one hand, by fusing the sample edge image and the sample mask image to obtain the sample guidance image, the extraction and fusion of image details can be achieved, enhancing the level of image detail. Since the sample guidance image with detailed information is considered during model training, the image editing model is trained using the sample image, the sample guidance image, and the corresponding sample description text. This utilizes images with richer internal structures for model training, thus improving comprehensiveness and accuracy, and enhancing the model training effect. On the other hand, since the model is trained using both the sample description text and the sample guidance image, the problem of needing to retrain the model for each application scenario in related technologies is avoided, improving the model's versatility and flexibility, and expanding its application scope. Furthermore, during image processing based on the trained image editing model, the accuracy and quality of image processing can be improved.

[0012] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0014] Figure 1 A schematic diagram illustrates an application scenario where the model training method of the present disclosure embodiments can be applied.

[0015] Figure 2 The diagram illustrates a model training method according to an embodiment of the present disclosure.

[0016] Figure 3 The schematic diagram illustrates a process for determining a sample guide image in an embodiment of this disclosure.

[0017] Figure 4 The illustration shows a schematic diagram of a sample fusion image according to an embodiment of the present disclosure.

[0018] Figure 5 This illustration schematically shows a diagram of another sample fusion image in an embodiment of the present disclosure.

[0019] Figures 6A-6B A schematic diagram illustrating an adjusted sample fusion image according to an embodiment of the present disclosure is shown.

[0020] Figure 7The schematic diagram illustrates the process of training an image editing model according to an embodiment of the present disclosure.

[0021] Figure 8 The illustration shows a flowchart of an image processing method according to an embodiment of the present disclosure.

[0022] Figure 9 This illustration schematically shows a target image obtained from a data factory link according to an embodiment of the present disclosure.

[0023] Figure 10 This illustration shows a schematic diagram of generating an image based on a trained image editing model according to an embodiment of the present disclosure.

[0024] Figure 11 This schematic diagram illustrates the overall process of data production in an embodiment of the present disclosure.

[0025] Figure 12 This schematic diagram illustrates the overall process of data production via a data factory link in an embodiment of this disclosure.

[0026] Figure 13 This diagram illustrates a target image generated in an embodiment of the present disclosure.

[0027] Figure 14 A block diagram of a model training apparatus according to an embodiment of the present disclosure is shown schematically.

[0028] Figure 15 A block diagram of an image processing apparatus according to an embodiment of the present disclosure is shown schematically.

[0029] Figure 16 A block diagram of an electronic device according to an embodiment of the present disclosure is shown schematically. Detailed Implementation

[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0031] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0032] In current research on segmentation networks, achieving superior network performance typically requires a large amount of training data and long periods of iterative training. However, the quantity and quality of training data are often difficult to guarantee. Therefore, in practical applications of algorithms, data engineering can provide high-quality training data. Data engineering involves the process of collecting, cleaning, labeling, and preprocessing data to provide accurate, rich, and representative training samples for the algorithm. This includes collecting data from multiple sources, performing quality verification and screening, and labeling data to generate accurate labels. Therefore, it is necessary to address common problems such as data imbalance, noise, and labeling errors to ensure the quality and reliability of the training data.

[0033] In some embodiments of this disclosure, data production can be achieved through generative networks. For example, Stable Diffusion can be used. However, the output of the generative model is greatly affected by the prompt text and network hyperparameter configuration. In different application scenarios, it is necessary to retrain the model or fine-tune the hyperparameters to generate the required output, which lacks flexibility and has low training efficiency. To achieve the guideability of Diffusion output and obtain high-quality generated results in a short time, large models can be fine-tuned and retrained. This combines pre-trained generative networks with the needs of specific tasks, and improves the quality and controllability of generated output by fine-tuning the model parameters. However, when applying generative networks to a data factory pipeline, the original image-segmentation data pairs provided by the segmentation network contain errors, which may lead to poor quality of generated data. Therefore, it may be necessary to introduce prior information to augment the input data to supplement local texture details. Furthermore, existing data factory pipelines are inefficient and have poor practicality.

[0034] To address the aforementioned technical issues, this disclosure provides a model training method that can be applied to automated data generation scenarios within a data factory pipeline. Figure 1 A schematic diagram of a system architecture for model training methods and apparatus applicable to embodiments of the present disclosure is shown.

[0035] like Figure 1As shown, a sample image and its corresponding sample mask image can be obtained; then, the sample edge image corresponding to the sample image is obtained, and the sample edge image and the sample mask image are fused to obtain a sample guide image; next, the image editing model is trained using the sample image, the sample guide image, and the sample description text corresponding to the sample image to obtain the trained image editing model. Further, the image to be processed can be input into the trained image editing model to generate a target image corresponding to the image to be processed.

[0036] It should be noted that the above model training method can be executed and deployed on a server. For example, an image editing model can be trained on a server to obtain a trained image editing model. Furthermore, the trained image editing model can be used to execute subsequent model application stages to achieve batch data production.

[0037] This disclosure provides a model training method that can be applied to the model training phase during data production in a data factory chain. Next, refer to... Figure 2 The steps of the model training method in the embodiments of this disclosure are described in detail below.

[0038] In step S210, the sample image and the corresponding sample mask image are obtained.

[0039] In this embodiment, the sample image can be any type of image, such as a static image or a dynamic image, and the objects contained therein can be people, animals, still life, or any type of object. The mask image can be used to extract the objects contained in the sample image. The sample image can be segmented using an image segmentation algorithm to obtain the corresponding sample mask image. The image segmentation algorithm can be a semantic segmentation algorithm or an instance segmentation algorithm; no specific limitation is made here, as long as it can achieve foreground and background images. Specifically, for a single sample image, semantic segmentation can segment all targets (including the background), but it cannot distinguish different individuals within the same category. Instance segmentation can segment all targets in the sample image except for the background, and can distinguish different individuals within the same category.

[0040] Based on this, the foreground image corresponding to the sample image can be obtained as a sample mask image by performing semantic segmentation or instance segmentation on the sample image. The sample image and the sample mask image can be combined into an image pair, and this image pair can be used as input to train the image editing model in the data factory chain to achieve model training and data generation.

[0041] In step S220, the sample edge image corresponding to the sample image is obtained, and the sample edge image is fused with the sample mask image to obtain the sample guide image.

[0042] In this embodiment of the disclosure, the sample guide image is used to represent the conditions for generating images. Through this sample guide image, guided image data of the same type can be generated in batches based on a small number of images, realizing mass data production. The sample guide image may include any one of the following: sample edge image, sample mask image, sample edge image, sample fusion image obtained by fusing sample mask image, and other types of images, specifically determined according to the accuracy and requirements of the downstream data production task.

[0043] In some embodiments, sample images can be segmented, and the segmentation results represented by sample mask images can be used as sample guide images for generating images. This allows a data factory link capable of generating guided outputs in batches to be established using a small segmentation dataset, so that the output of this data factory link can be used as input for downstream tasks.

[0044] In other embodiments, to enable the network to learn more texture information from the image used as the guiding input and to preserve edge features, edge detection can be performed on the sample image to obtain a sample edge image. This sample edge image and the sample mask image are then combined to determine the sample guiding image. Specifically, the sample edge image and the sample mask image can be fused to obtain a sample fused image. The sample guiding image containing texture details is then determined based on this fused image. After obtaining the sample fused image, it can be directly used as the sample guiding image. Alternatively, to obtain a more refined edge segmentation result, the segmentation boundaries in the sample fused image can be finely adjusted to obtain an adjusted sample fused image, which is then used as the sample guiding image.

[0045] Figure 3 The flowchart illustrating the process of determining a sample guide image based on a sample fusion image is shown in the image below. Figure 3 As shown, the main steps include:

[0046] In step S310, edge detection is performed on the sample image to obtain the sample edge image.

[0047] In this step, edge detection can be performed on the sample image using an edge detection model. The edge detection model can be any type capable of extracting edges; here, the HED (Holistically-nested Edge Detection) model is used as an example. Based on this, the output HED image can be used as the sample edge image.

[0048] In some embodiments, the HED model includes a five-level feature extraction architecture. In each level, convolutional layers and pooling layers are used to extract hierarchical feature maps. Multiple hierarchical feature maps can be merged, for example, multiple hierarchical feature maps can be concatenated according to the channel dimension. Furthermore, 1*1 convolution operations are performed on multiple hierarchical feature maps to determine the sample edge image.

[0049] In step S320, the sample edge image and the sample mask image are fused to obtain a sample fused image.

[0050] In this step, the sample edge image and the sample fusion image are fused pixel by pixel. This involves fusing the pixel values ​​of each pixel in the sample edge image and the sample fusion image to obtain the fused pixel value for each pixel. The fused sample image is then determined based on these fused pixel values. It should be noted that the sample image can be a person, still life, etc. When the sample image is a still life, the effect of the resulting fused sample image can be referenced... Figure 4 As shown in the image. When the sample image is a person, the effect of the resulting fused image can be referenced. Figure 5 As shown in the image.

[0051] In step S330, the segmentation boundaries of the sample fusion image are adjusted to obtain the adjusted sample fusion image.

[0052] In this step, after obtaining the sample fusion image, in order to obtain more refined edge segmentation results, the segmentation boundaries of the sample fusion image can be finely adjusted to further improve the data quality. For example, the Matteformer model can be used to finely adjust the segmentation boundaries of the sample fusion image.

[0053] In some embodiments, a tri-image corresponding to the sample image can be obtained. Prior labels contained in the unknown regions of the tri-images are used as global labels to determine the grayscale values ​​of pixels in the unknown regions, thus obtaining an adjusted sample fusion image. Specifically, the Matteformer model can contain multiple encoding layers and one decoding layer. Each encoding layer can contain a Transformer module. Prior labels of the unknown regions in the tri-images can be used as the Key and Value values ​​of each Transformer module for self-attention operations, resulting in multiple implicit features as the output of the encoding layer. These implicit features are then input to the decoding layer for decoding to obtain the decoding result. The decoding result can be the grayscale value of each pixel in the unknown regions of the tri-images.

[0054] Specifically, the trimap image has prior labels, which contain information about each trimap region, such as foreground, background, and unknown regions. These prior labels are used as global prior labels and participate in the self-attention mechanism of each Transformer module. During the encoding stage, the trimap information in the Transformer block is fully utilized, enabling the network to learn more expressive implicit features.

[0055] In some embodiments, multiple self-attention operations can be performed on each encoding layer based on prior labels of the unknown region to obtain multiple implicit features. These implicit features can then be decoded to determine the grayscale values ​​of the pixels in the unknown region, resulting in an adjusted sample fusion image. For example, the prior labels are used to indicate global information for each Trimap region and can be grayscale values ​​of the unknown region, such as 0 or 1. Based on this, the grayscale values ​​of 0 or 1 in the unknown region can be converted into specific grayscale values ​​from 0 to 255, thereby obtaining the adjusted sample fusion image, which gives the unknown region more detailed features and improves detail. The adjusted sample fusion image can be referenced... Figure 6A and Figure 6B As shown in the image.

[0056] Based on this, in the process of fine-tuning the sample fusion image using prior labels, incorporating Trimap prior labels into the network allows for a better understanding and capture of subtle features of the segmentation boundaries, resulting in more refined and accurate edge segmentation results. This improves the quality of data processing and yields more precise and detailed sample fusion images.

[0057] In step S340, the adjusted sample fusion image is used as the sample guide image.

[0058] In this step, the finely adjusted sample fusion image containing texture details can be directly used as the sample guide image, which allows the output image to better retain the texture details of the sample image and improve the accuracy of details.

[0059] It should be added that the sample fusion image can also be used directly as the sample guide image without the need for fine-tuning of the sample fusion image.

[0060] In this embodiment of the disclosure, the sample guide image is determined by the sample fusion image or the adjusted sample fusion image, which enables the sample guide image to better preserve the texture details of the image, thereby making the generated image have detailed information and improving accuracy and detail.

[0061] Next, in step S230, the image editing model is trained using the sample image, the sample guide image, and the sample description text corresponding to the sample image to obtain the trained image editing model.

[0062] In this embodiment of the disclosure, the sample description text can include text of any type. For example, it can be a sentence, phrase, or other text containing one or more words, and can be used to describe the semantic information contained in the sample image. The sample description text can be used to characterize the content of the image to be generated, and is specifically determined based on the information of the sample image. The sample image can be input into the BLIP model to obtain the sample description text corresponding to the sample image. The sample guide image is used to characterize the region range of the image to be generated.

[0063] Image editing models can be diffusion models or any other type of image processing model; here, we will use the Stable Diffusion model as an example. The diffusion model employs the idea of ​​a diffusion process, enabling the generation of diverse and detailed data, such as images and text. Its principle is that since each image follows a certain regular distribution, this distribution information contained in the text is used as a guide to gradually denoise a purely noisy image, generating an image that matches the text information.

[0064] In some embodiments, the diffusion model includes a text encoder, an image information creator, and a decoder. The text encoder converts the input text into a spatial code that the image information creator can understand to obtain text features. Its input is a string of text, and its output is a series of semantic vectors containing textual information describing the text.

[0065] The image information creator contains multiple UNet networks. Each UNet simulates the distribution of noise, and its output is the predicted noise. The image information creator employs forward and backward diffusion processes. Forward diffusion progressively adds random noise to the sample image, eventually resulting in pure noise. Backward diffusion denoises the pure noise to generate the image. For example, the input to the image information creator can be the image latent vector of pure noise (i.e., the image features of the sample image superimposed with random noise), and the output can be the denoised image latent vector (denoised image features), which are low-dimensional image features. Specifically, text features and guiding features corresponding to the sample guiding image can be used as control conditions. Denoising is iteratively performed starting from the image features of pure noise, and text features and guiding features are incorporated into the image features to obtain denoised image features with semantic and guiding information. During the denoising process, text features and guiding features can be used as control condition inputs, along with the time step, to guide the denoising of low-dimensional features through simple connections or cross-attention.

[0066] The decoder is used to transform low-dimensional features into high-dimensional features to achieve image generation. Based on this, the denoised image features output by the image information creator can be input into the decoder for decoding to generate an image with semantic and guiding information.

[0067] To improve the accuracy of image generation, the image editing model can be trained to obtain a trained image editing model. Figure 7 The flowchart illustrating the training of the image editing model is shown in the image diagram below. Figure 7 As shown, the main steps include:

[0068] In step S710, the reference image editing model is trained based on the sample image, the sample guide image, and the sample description text to obtain the trained reference image editing model.

[0069] In this embodiment of the disclosure, the reference image editing model can be a model with the same type and parameters as the image editing model, such as the diffusion model Stable Diffusion.

[0070] This disclosure describes the training process of the diffusion model. When training the diffusion model, it is not necessary to train the entire network. Instead, a large model fine-tuning scheme using ControlNet is employed, employing a supernetwork design method to fine-tune the parameters of the diffusion model so that the generative network can adapt to specific data generation tasks. For example, during training, the gradient updates of the diffusion model are frozen, and a portion of the network and its weights are copied as a fixed model, while the remaining portion serves as a non-fixed model. Next, a new diffusion model (referencing the image editing model) can be trained, fine-tuning its corresponding parameters layer by layer, thereby adjusting the parameters of the original diffusion model. That is, the parameters of other network layers in the new diffusion model are first fixed, while the parameters of the current layer are adjusted. The initial parameters of the new diffusion model are the same as those of the original diffusion model.

[0071] In this embodiment, the parameters of the new diffusion model can be controllably adjusted layer by layer according to the control conditions represented by the sample guidance image and sample description text. Since the model parameters can be adjusted based on control conditions such as the sample guidance image and sample description text, controllable model training is achieved. Furthermore, because the fused image possesses more complete edge and detail texture information, and the network weight settings of the supernetwork remain consistent with the original diffusion model, the original diffusion model can be easily fine-tuned, allowing the output of the diffusion model to be guided by the supernetwork and generate an image that meets the requirements.

[0072] When training the reference image editing model, random noise can be added to the sample images to obtain pure noise image latent vectors. The sample guidance image and sample description text are used as control conditions. Guidance features and text features are used as input control conditions and fused with time-step features. Denoising is iteratively performed starting from the pure noise image latent vector, and text and guidance features are incorporated into the image latent vector to obtain a denoised image latent vector with semantic and guidance information. A reference image is then generated based on the denoised image latent vector. The reference image can be an image generated based on the sample image and associated with the sample description text and sample guidance image.

[0073] After obtaining the reference image, to assess accuracy, the relevance between the reference image and the sample guide image and sample descriptive text can be determined. Specifically, the similarity between the reference image and the sample guide image can be calculated to obtain their relevance; alternatively, the similarity between the reference image and the sample descriptive text can be calculated to obtain their relevance. The relevance is considered satisfied when the relevance with both the sample guide image and the sample descriptive text is greater than the corresponding relevance threshold.

[0074] Next, the loss function can be determined based on the noise distribution, predicted noise distribution, and Gaussian distribution corresponding to the current time step. The noise distribution, predicted noise distribution, and Gaussian distribution can be logically combined to obtain the loss function, as shown in formula (1).

[0075]

[0076] Where ε is the noise distribution at the current time step t, ε θ This represents the distribution of prediction noise generated by the network at the current time step t. Let be the variance of the Gaussian distribution at the current time step t.

[0077] After obtaining the loss function, the reference image editing model can be trained layer by layer based on the loss function until the relevance meets the relevance condition, thus obtaining the trained reference image editing model.

[0078] It should be noted that an image editing model can also be trained directly from the fused sample image without using a Matteformer model for fine-tuning. In this case, after obtaining the trained image editing model, a reference image can be generated using the trained model. The Matteformer model can then be used to fine-tune the reference image and its corresponding mask image, thus achieving model training.

[0079] In step S720, the parameters of the image editing model are adjusted according to the trained reference image editing model to obtain the trained image editing model.

[0080] In this embodiment of the present disclosure, after the reference image editing model is trained, the parameters of the non-fixed model in the image editing model can be adjusted according to the parameters of the trained reference image editing model, while keeping the parameters of the fixed model in the image editing model unchanged, thereby obtaining the trained image editing model.

[0081] In this embodiment, a sample guidance image is obtained by fusing sample edge images and sample mask images. The image editing model is then trained based on this guide image and sample description text. Since the guide image contains more detailed information, it allows for accurate training of the image editing model, improving its accuracy. Furthermore, because the guide image possesses more complete edge and detail information, and the training of the image editing model is guided by a reference image editing model, the controllability of the output is ensured. In addition, because the guide image has more complete edge and detail texture information, and the network weights of the supernetwork are consistent with the original image editing model, fine-tuning of the original model is convenient. This allows the output of the generator network to be guided by the supernetwork, improving the convenience and efficiency of model training, as well as its versatility and flexibility. By introducing parameter fine-tuning into the toolchain, images that meet specific requirements can be generated, improving the effectiveness of the data generation task.

[0082] This disclosure also provides an image processing method, referring to... Figure 8 As shown, this image processing method can be applied to the inference stage, specifically including the following steps:

[0083] In step S810, the image to be processed and the mask image corresponding to the image to be processed are obtained;

[0084] In step S820, the edge image corresponding to the image to be processed is obtained, and the edge image is fused with the mask image to obtain the guide image;

[0085] In step S830, the image to be processed, the guide image, and the descriptive text corresponding to the image to be processed are input into the trained image editing model to obtain the target image.

[0086] In the technical solution provided by this disclosure, on the one hand, since the guide image is obtained by fusing the edge image and the mask image, and the guide image has more complete edge and detail textures, the target image can be generated based on accurate input, thus improving the accuracy and quality of the target image. On the other hand, the guide image and descriptive text can be processed by the trained image editing model in the data factory link, realizing a data production process that generates a large number of target images that conform to the guide image and descriptive text from a small number of images to be processed, thereby improving data production efficiency.

[0087] Next, refer to Figure 8 Each step of the image processing method in the embodiments of this disclosure will be described in detail.

[0088] In step S810, the image to be processed and the mask image corresponding to the image to be processed are obtained.

[0089] In this embodiment, the image to be processed can be any type of image, and the objects contained therein can be people, animals, still life, or any type of object. The mask image is a binary image composed of 0s and 1s. Specific images can be extracted from an image using the mask image. The mask image corresponding to the image to be processed can be obtained by segmenting the image to be processed using an image segmentation algorithm. The image segmentation algorithm can be a semantic segmentation algorithm or an instance segmentation algorithm; no specific limitation is made here.

[0090] Based on this, the image to be processed and the mask image can be combined into an image pair, which can then be used as the input to the data factory link to generate data.

[0091] In step S820, the edge image corresponding to the image to be processed is obtained, and the edge image is fused with the mask image to obtain the guide image.

[0092] In this embodiment, the guide image represents the control conditions for image processing of the image to be processed. The guide image can represent the area range in the image to be processed where image editing is performed. Using this guide image, guided image data of the same type can be generated in batches based on a small number of images, enabling large-scale data production. The guide image can include any one of the following: edge image, mask image, fused image corresponding to the edge image and mask image, or other types of images, specifically determined according to the accuracy and requirements of the downstream data production task.

[0093] In this embodiment of the disclosure, in order to enable the network to learn more texture information from the image used as the guiding input and to preserve the edge features of the image, edge detection can be performed on the image to be processed to obtain an edge image. The edge image and the mask image are then combined to determine the guiding image. Specifically, the edge image and the mask image can be fused to obtain a fused image, and the guiding image containing texture details can be determined based on the fused image. After obtaining the fused image, it can be directly used as the guiding image. Alternatively, to obtain a more refined edge segmentation result, the segmentation boundaries in the fused image can be finely adjusted to obtain an adjusted fused image, and the guiding image can be determined based on the adjusted fused image.

[0094] Alternatively, edge extraction can be performed on the image to be processed based on the HED model to obtain an edge image. When performing fine adjustments, the Matteformer model can still be used to finely adjust the segmentation boundaries of the fused image. The specific processing procedure is the same as the execution process in steps S310 to S330, and will not be repeated here.

[0095] Furthermore, the fused image can be used directly as the guide image, or the adjusted fused image can be used as the guide image, depending on the actual needs.

[0096] In step S830, the image to be processed, the guide image, and the descriptive text corresponding to the image to be processed are input into the trained image editing model to obtain the target image.

[0097] In this embodiment, the data factory chain can include multiple models, such as, but not limited to, a BLIP (Bootstrapping Language-Image Pre-training, a stable large graph-text model) model, a trained image editing model, etc. Based on this, descriptive text can first be obtained from the BLIP model in the data factory chain. The descriptive text can be used to represent information about the image to be generated from the image to be processed, and can be specifically determined based on the information contained in the image to be processed. For example, it can include, but is not limited to, objects, environment, actions, quantity, size, style, etc., contained in the image, depending on actual needs. For example, the image to be processed is input into the BLIP model, and after processing, a text description prompt corresponding to the image scene is output.

[0098] Furthermore, the image to be processed, the descriptive text, and the guiding image can be used as inputs to a trained image editing model in the data factory chain to generate images, thereby realizing data production based on the data factory chain. The trained image editing model can be a trained diffusion model capable of converting text to images. For example, the image to be processed, the guiding image, and the descriptive text can be input into the trained image editing model to generate multiple generated images; further, the differences between each generated image and the image to be processed can be determined, and the target image can be determined from the multiple generated images based on these differences. (Reference) Figure 9 As shown, the image to be processed, the guide image, and the descriptive text can be used as inputs to the trained image editing model to obtain multiple generated images as outputs. These generated images are then filtered to obtain the target image. The number of target images can be one or more, without specific limitation. It should be noted that the mask image of the target image is the same as the mask image of the image to be processed; that is, the target image and the image to be processed are of the same type of image and share the same mask image.

[0099] refer to Figure 10 As shown, the trained image editing model can include a text encoder, an image information creator, and a decoder. During image generation, the text encoder first extracts features from the descriptive text to obtain text features. Next, since the guiding image and descriptive text are used as control conditions, the guiding features corresponding to the guiding image and the text features are fused to obtain fused features. These fused features are then used to fit the image features of the image to be processed, which has random features, and the time step features to obtain intermediate features as the output of the image information creator. For example, using text features and guiding features as control conditions, iterative denoising is performed starting from the image latent vector of pure noise (the image features of the image to be processed with random noise), and text features and guiding features are incorporated into the image features to obtain denoised image features with semantic and guiding information. During denoising, text features and guiding features can be used as input control conditions, along with the time step, to guide the denoising of low-dimensional features through simple connections or cross-attention, thereby obtaining intermediate features. Furthermore, the intermediate features can be decoded based on the decoder to obtain multiple generated images.

[0100] After obtaining multiple generated images, the differences between the generated images and the image to be processed can be determined, and the target image can be determined from the multiple generated images based on the differences. The differences are used to describe the magnitude of the difference between the generated images and the image to be processed. The differences can be determined based on one or more of a first metric and a second metric. The first metric can be PSNR (Peak Signal-to-Noise Ratio), and the second metric can be SSIM (Structural Similarity). The first metric can be calculated using formula (2):

[0101]

[0102] Among them, This represents the maximum possible pixel value of the image. If each pixel is represented by 8 bits, then... It can be 255. MSE is the mean square error for each pixel, which can be calculated using formula (3):

[0103]

[0104] The second indicator can be determined using formula (4):

[0105]

[0106] Where x and y represent the original image (image to be processed) and the generated image, respectively, μ x Let μ be the mean of x. y Let σ be the mean of y. x 2 Let σ be the variance of x. y 2 Let σ be the variance of y. xy Let c1 be the covariance of x and y, and c1 = (k1L) 2 c2 = (k2L) 2 Let L be two constants, where L is the range of pixel values. B -1, k1 = 0.01, k2 = 0.03.

[0107] Based on this, the difference can be determined solely by the first or second indicator, or a weighted sum of the first and second indicators can be used to determine the difference; no specific limitation is made here. After obtaining the differences, each difference can be arranged in ascending order, and the generated images corresponding to the top N differences can be used as the target images. The number N of target images can be determined according to actual business needs. If only one target image is needed, the generated image with the smallest difference can be directly used as the target image; alternatively, the top N differences can be displayed on the interactive interface, and the target image can be determined based on the user's selection of the generated image on the interactive interface.

[0108] If a large number of target images are needed, and the number of generated images with differences less than the difference threshold meets the quantity requirement, all generated images with differences less than the difference threshold can be directly used as target images to facilitate downstream tasks. If a large number of target images are needed, but the number of generated images with differences less than the difference threshold does not meet the quantity requirement, the image editing model can be retrained to regenerate data based on the retrained model, thereby enabling downstream tasks to be performed based on the target images.

[0109] Based on this, a data factory chain consisting of an image generation model and a trained image editing model can automatically generate similar datasets in batches, thus obtaining a dataset with a specified pose and sharing the same mask image as the input image to be processed. Based on the original image to be processed, a guide image obtained by fusing the edge image and the mask image can be determined. Further data generation based on the image to be processed and the fused image can then yield the target image. Furthermore, fine-tuning of the target image and the mask image can be performed to improve accuracy.

[0110] After obtaining the dataset of target images corresponding to the image to be processed, the data factory link outputs an image pair consisting of the target image and its corresponding mask image, such as <target image, mask image>. Further, downstream tasks can be executed based on the image pairs output by the data factory link. Downstream tasks can be, for example, image segmentation or other types of operations, depending on the actual business requirements. Here, we will use image segmentation as an example of a downstream task for illustration.

[0111] For example, image pairs consisting of all target images and their corresponding mask images can be used as input to train the image segmentation model without annotation, thus obtaining a trained image segmentation model. Since the output consists of the target images and their corresponding mask images, there is no need to re-annotate the target images. Therefore, the image segmentation model can be accurately trained based on the generated target images with detailed information to complete downstream tasks.

[0112] In this embodiment, by combining the BLIP model and the diffusion model, an automated descriptive text prediction process is achieved. This enables the network to perform data generation tasks without human intervention, and the generated data is more consistent with the semantics and scene features of the input image. Furthermore, through the data factory link, a dataset with specified poses and sharing the same mask image can be generated using a small number of images to be processed, realizing the function of automatically generating similar data in batches.

[0113] Figure 11 The diagram illustrates the overall process of data production. (See reference) Figure 11 As shown, it mainly includes the model training phase and the model application phase, including the following steps:

[0114] In step S1110, a sample image and a sample mask image of the sample image are acquired;

[0115] In step S1120, the sample edge image corresponding to the sample image is obtained;

[0116] In step S1130, the sample edge image and the sample mask image are fused to obtain the sample guide image;

[0117] In step S1140, the sample description text corresponding to the sample image is obtained;

[0118] In step S1150, a loss function is determined, and the image editing model is trained based on the loss function.

[0119] In step S1160, it is determined whether the model training has ended; if yes, proceed to step S1170; if no, proceed to step S1110.

[0120] In step S1170, the trained image editing model is obtained.

[0121] In step S1180, the image to be processed, the corresponding guide image, and the descriptive text are input into the trained image editing model.

[0122] In step S1190, the target image is output.

[0123] Figure 12 The diagram illustrates a flowchart of data production via a data factory link. (See reference) Figure 12 As shown, the main steps include:

[0124] In step S1202, the image to be processed and the mask image are obtained;

[0125] In step S1204, the edge image and mask image corresponding to the image to be processed are fused to obtain the guide image;

[0126] In step S1206, the image to be processed is input into the BLIP model to obtain a text description;

[0127] In step S1208, the image to be processed, the text description, and the guide image are input into the trained image editing model;

[0128] In step S1210, the target image is generated by sharing the mask image.

[0129] The technical solution in this embodiment fuses sample edge images and sample mask images to obtain sample guidance images. Image editing models are then trained based on these guidance images and sample description text. Since the sample guidance images contain more detailed information, the image editing models can be trained accurately, improving model accuracy. Furthermore, because the sample guidance images possess more complete edge and detail information, and the training of the image editing model is guided by a reference image editing model, the controllability of the output is ensured. In addition, because the sample guidance images possess more complete edge and detail texture information, and the network weight settings of the supernetwork are consistent with the original image editing model, the original model can be easily fine-tuned. This allows the output of the generator network to be guided by the supernetwork, improving the convenience and efficiency of model training, as well as enhancing the model's versatility and flexibility. By introducing parameter fine-tuning into the toolchain, images that meet specific requirements can be generated, improving the effectiveness of data generation tasks. The generated image obtained from the data factory link, which is of the same type as the image to be processed and shares a mask image, can be used as a reference. Figure 13 As shown in the image.

[0130] This disclosure provides a model training apparatus, with reference to... Figure 14 As shown, the model training device 1400 may include: an image acquisition module 1401, a sample guidance image generation module 1402, and a guided training module 1403; wherein:

[0131] Image acquisition module 1401 can be used to acquire sample images and sample mask images corresponding to the sample images;

[0132] The sample guidance image generation module 1402 can be used to obtain the sample edge image corresponding to the sample image, and fuse the sample edge image with the sample mask image to obtain the sample guidance image;

[0133] The guided training module 1403 can be used to train the image editing model using the sample image, the sample guidance image, and the sample description text corresponding to the sample image, so as to obtain the trained image editing model.

[0134] In one exemplary embodiment of this disclosure, the sample guidance image generation module includes: a generation control module, configured to fuse the sample edge image and the sample mask image pixel by pixel to obtain a sample fusion image, and determine the sample guidance image based on the sample fusion image.

[0135] In one exemplary embodiment of this disclosure, the generation control module includes: a three-part image acquisition module, used to acquire three-part images corresponding to the sample image; an adjustment module, used to determine the grayscale values ​​of pixels in the unknown regions based on prior markers contained in the unknown regions of the three-part images, so as to obtain an adjusted sample fusion image; and a generation module, used to use the adjusted sample fusion image as the sample guidance image.

[0136] In one exemplary embodiment of this disclosure, the adjustment module includes: a self-attention operation module, configured to perform multiple self-attention operations based on the prior label to obtain multiple implicit features; and a grayscale value determination module, configured to decode the multiple implicit features to determine the grayscale value of the pixels in the unknown region, so as to obtain the adjusted sample fusion image.

[0137] In one exemplary embodiment of this disclosure, the guided training module includes: a reference training module, used to train a reference image editing model based on the sample image, the sample guidance image, and the sample description text to obtain a trained reference image editing model; and a fine-tuning module, used to adjust the parameters of the image editing model according to the trained reference image editing model to obtain the trained image editing model.

[0138] In one exemplary embodiment of this disclosure, the reference training module includes: a reference image generation module, configured to generate a reference image from the sample image based on the reference image editing model, using the sample guidance image and the sample description text as control conditions; a relevance determination module, configured to determine the relevance between the sample guidance image and the sample description text and the reference image; a loss function determination module, configured to determine a loss function based on the noise distribution, predicted noise distribution, and Gaussian distribution corresponding to the current time step; and a training control module, configured to train the reference image editing model layer by layer based on the loss function until the relevance meets the relevance condition, thereby obtaining the trained reference image editing model.

[0139] In one exemplary embodiment of this disclosure, the image editing model includes a fixed model and a non-fixed model; the fine-tuning module is configured to: adjust the parameters of the non-fixed model in the image editing model according to the parameters of the trained reference image editing model, while keeping the parameters of the fixed model in the image editing model unchanged, so as to obtain the trained image editing model.

[0140] In this embodiment of the disclosure, an image processing apparatus is also provided, with reference to Figure 15 As shown, the image processing device 1500 may include: an image acquisition module 1501, a guided image determination module 1502, and an image generation module 1503; wherein:

[0141] Image acquisition module 1501 is used to acquire the image to be processed and the mask image corresponding to the image to be processed;

[0142] The guide image determination module 1502 is used to obtain the edge image corresponding to the image to be processed, and fuse the edge image with the mask image to obtain the guide image;

[0143] The image generation module 1503 is used to input the image to be processed, the guide image, and the description text corresponding to the image to be processed into the trained image editing model to obtain the target image; wherein, the trained image editing model is trained according to any of the model training methods described above.

[0144] In one exemplary embodiment of this disclosure, the image generation module includes: an image generation determination module, configured to input the image to be processed, the guiding image, and the descriptive text into the trained image editing model to generate an image, thereby obtaining a plurality of generated images; and a target image determination module, configured to determine the differences between the generated images and the image to be processed, and to determine the target image from the plurality of generated images based on the differences.

[0145] In an exemplary embodiment of this disclosure, the generated image determination module includes: a text feature extraction module, used to extract features from the descriptive text to obtain text features; an intermediate feature fitting module, used to fuse the guidance features of the guidance image and the text features to obtain fused features, and to fit the image features and time step features corresponding to the image to be processed with random noise in combination with the fused features to obtain intermediate features; and a decoding module, used to decode the intermediate features to obtain multiple generated images.

[0146] In an exemplary embodiment of this disclosure, the target image determination module includes: a first index determination module, configured to determine a first index based on the average pixel value of the generated image and the pixel value of any pixel in the image to be processed; a second index determination module, configured to determine a second index based on parameters corresponding to the pixel values ​​of the generated image and the image to be processed; and a difference determination module, configured to determine the difference based on one or more of the first index and the second index.

[0147] In one exemplary embodiment of this disclosure, after obtaining the target image, the apparatus further includes: an image segmentation model training module, configured to train an image segmentation model based on the target image and a mask image corresponding to the target image, to obtain a trained image segmentation model.

[0148] It should be noted that the specific details of each part of the above-mentioned model training device and image processing device have been described in detail in some embodiments of the corresponding methods. For details that are not disclosed, please refer to the implementation content of the method section, and therefore will not be repeated here.

[0149] Exemplary embodiments of this disclosure also provide an electronic device. This electronic device may be the aforementioned mobile terminal device. Generally, the electronic device may include a processor and a memory, the memory storing executable instructions of the processor, and the processor configured to execute the aforementioned image processing method and video text retrieval method by executing the executable instructions.

[0150] The following is based on Figure 16 Taking the mobile terminal 1600 as an example, the construction of this electronic device will be described by way of example. Those skilled in the art will understand that, apart from components specifically designed for mobile purposes, Figure 16 The structure can also be applied to fixed types of equipment.

[0151] like Figure 16 As shown, the mobile terminal 1600 may specifically include: a processor 1601, a memory 1602, a bus 1603, a mobile communication module 1604, an antenna 1, a wireless communication module 1605, an antenna 2, a display screen 1606, a camera module 1607, an audio module 1608, a power module 1609, and a sensor module 1610.

[0152] Processor 1601 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, an encoder, a decoder, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The image processing method and video text retrieval method in this exemplary embodiment can be executed by an AP, GPU, or DSP. When the method involves neural network-related processing, it can be executed by an NPU. For example, the NPU can load neural network parameters and execute neural network-related algorithm instructions.

[0153] An encoder encodes (compresses) images or videos to reduce data size for easier storage or transmission. A decoder decodes (decompresses) the encoded data to restore the original image or video data. The mobile terminal 1600 can support one or more encoders and decoders, such as image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), and BMP (Bitmap), and video formats like MPEG (Moving Picture Experts Group) 1, MPEG10, H.1063, H.1064, and HEVC (High Efficiency Video Coding).

[0154] The processor 1601 can be connected to the memory 1602 or other components via the bus 1603.

[0155] The memory 1602 can be used to store executable program code, which includes instructions. The processor 1601 executes various functional applications and data processing of the mobile terminal 1600 by running the instructions stored in the memory 1602. The memory 1602 can also store application data, such as images, videos, and other files.

[0156] The communication functions of mobile terminal 1600 can be implemented through mobile communication module 1604, antenna 1, wireless communication module 1605, antenna 2, modem processor, and baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Mobile communication module 1604 can provide 3G, 4G, and 5G mobile communication solutions for mobile terminal 1600. Wireless communication module 1605 can provide wireless communication solutions such as wireless LAN, Bluetooth, and near-field communication for mobile terminal 1600.

[0157] The display screen 1606 is used to implement display functions, such as displaying user interfaces, images, and videos. The camera module 1607 is used to implement shooting functions, such as capturing images and videos, and may include a color temperature sensor array. The audio module 1608 is used to implement audio functions, such as playing audio and capturing voice. The power module 1609 is used to implement power management functions, such as charging the battery, supplying power to the device, and monitoring battery status. The sensor module 1610 may include one or more sensors to implement corresponding sensing and detection functions. For example, the sensor module 1610 may include an inertial sensor, which is used to detect the motion posture of the mobile terminal 1600 and output inertial sensing data.

[0158] It should be noted that the present disclosure also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or it may exist alone and not assembled into the electronic device.

[0159] Computer-readable storage media can be, for example—but not limited to—electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0160] A computer-readable storage medium can be sent, propagated, or transmitted for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wireline, optical fiber, RF, etc., or any suitable combination thereof.

[0161] A computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to perform the methods described in the following embodiments.

[0162] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.

[0163] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0164] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0165] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0166] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A model training method, characterized in that, include: Obtain the sample image and the corresponding sample mask image; Obtain the sample edge image corresponding to the sample image, and fuse the sample edge image with the sample mask image to obtain the sample guide image; The image editing model is trained using the sample image, the sample guidance image, and the sample description text corresponding to the sample image to obtain the trained image editing model; The image editing model includes a fixed model and a non-fixed model; the image editing model is trained using the sample image, the sample guidance image, and the sample description text corresponding to the sample image to obtain the trained image editing model, including: The reference image editing model is trained based on the sample image, the sample guidance image, and the sample description text to obtain the trained reference image editing model. Based on the parameters of the trained reference image editing model, the parameters of the non-fixed model in the image editing model are adjusted, while the parameters of the fixed model in the image editing model remain unchanged, so as to obtain the trained image editing model.

2. The model training method according to claim 1, characterized in that, The step of fusing the sample edge image with the sample mask image to obtain the sample guidance image includes: The sample edge image and the sample mask image are fused pixel by pixel to obtain a sample fused image, and the sample guide image is determined based on the sample fused image.

3. The model training method according to claim 2, characterized in that, Determining the sample guidance image based on the sample fusion image includes: Obtain the trisection image corresponding to the sample image; Based on the prior labels contained in the unknown regions of the three-part image, the gray values ​​of the pixels in the unknown regions are determined to obtain the adjusted sample fusion image; The adjusted sample fusion image is used as the sample guide image.

4. The model training method according to claim 3, characterized in that, The step of determining the grayscale values ​​of pixels in the unknown regions based on prior labels contained in the unknown regions of the three-part image to obtain the adjusted sample fusion image includes: Multiple self-attention operations are performed based on the prior labels to obtain multiple implicit features; Decode multiple implicit features to determine the grayscale values ​​of pixels in the unknown region, so as to obtain an adjusted sample fusion image.

5. The model training method according to claim 1, characterized in that, The process of training the reference image editing model based on the sample image, the sample guidance image, and the sample description text to obtain the trained reference image editing model includes: Using the sample guide image and the sample description text as control conditions, a reference image is generated from the sample image based on the reference image editing model. Determine the correlation between the sample guide image and the sample description text and the reference image; The loss function is determined based on the noise distribution, predicted noise distribution, and Gaussian distribution corresponding to the current time step. The reference image editing model is trained layer by layer based on the loss function until the relevance meets the relevance condition, thus obtaining the trained reference image editing model.

6. An image processing method, characterized in that, include: Obtain the image to be processed and the mask image corresponding to the image to be processed; Obtain the edge image corresponding to the image to be processed, and fuse the edge image with the mask image to obtain the guide image; The image to be processed, the guide image, and the descriptive text corresponding to the image to be processed are input into the trained image editing model to obtain the target image; wherein, the trained image editing model is trained by the model training method according to any one of claims 1-5.

7. The image processing method according to claim 6, characterized in that, The step of inputting the image to be processed, the guide image, and the descriptive text corresponding to the image to be processed into the trained image editing model to obtain the target image includes: The image to be processed, the guide image, and the descriptive text are input into the trained image editing model to generate multiple generated images. The differences between the generated image and the image to be processed are determined, and the target image is determined from the plurality of generated images based on the differences.

8. The image processing method according to claim 7, characterized in that, The process involves inputting the image to be processed, the guiding image, and the descriptive text into the trained image editing model to generate multiple generated images, including: Feature extraction is performed on the descriptive text to obtain text features; The guidance features of the guidance image and the text features are fused to obtain fused features. The fused features are then combined with the image features and time step features corresponding to the image to be processed with random noise to obtain intermediate features. The intermediate features are decoded to obtain multiple generated images.

9. The image processing method according to claim 7, characterized in that, Determining the difference between the generated image and the image to be processed includes: The first index is determined based on the average pixel value of the generated image and the pixel value of any pixel in the image to be processed; The second index is determined based on the parameters corresponding to the pixel values ​​of the generated image and the image to be processed; The difference is determined based on one or more of the first indicator and the second indicator.

10. The image processing method according to claim 6, characterized in that, After obtaining the target image, the method further includes: The image segmentation model is trained based on the target image and the corresponding mask image to obtain the trained image segmentation model.

11. A model training device, characterized in that, include: The image acquisition module is used to acquire a sample image and a sample mask image corresponding to the sample image; The sample guidance image generation module is used to obtain the sample edge image corresponding to the sample image, and fuse the sample edge image with the sample mask image to obtain the sample guidance image; The guided training module is used to train the image editing model using the sample image, the sample guidance image, and the sample description text corresponding to the sample image, so as to obtain the trained image editing model. The image editing model includes a fixed model and a non-fixed model; the image editing model is trained using the sample image, the sample guidance image, and the sample description text corresponding to the sample image to obtain the trained image editing model, including: The reference image editing model is trained based on the sample image, the sample guidance image, and the sample description text to obtain the trained reference image editing model. Based on the parameters of the trained reference image editing model, the parameters of the non-fixed model in the image editing model are adjusted, while the parameters of the fixed model in the image editing model remain unchanged, so as to obtain the trained image editing model.

12. An image processing apparatus, characterized in that, include: The image acquisition module is used to acquire the image to be processed and the mask image corresponding to the image to be processed; A guide image determination module is used to obtain the edge image corresponding to the image to be processed, and fuse the edge image with the mask image to obtain a guide image; An image generation module is used to input the image to be processed, the guide image, and the descriptive text corresponding to the image to be processed into a trained image editing model to obtain a target image; wherein the trained image editing model is trained by the model training method according to any one of claims 1-5.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the model training method according to any one of claims 1-5 or the image processing method according to any one of claims 6-10.

14. An electronic device, characterized in that, include: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the model training method of any one of claims 1-5 or the image processing method of any one of claims 6-10 by executing the executable instructions.

Citation Information

Patent Citations

  • Training method of image editing model and image editing method and device

    CN116363261A