Image enhancement model training method and device, electronic equipment and storage medium

CN118608404BActive Publication Date: 2026-09-29BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410665004.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2026-09-29
Estimated Expiration
2044-05-27

AI Technical Summary

Technical Problem

[0004]为了解决上述技术问题或者至少部分地解决上述技术问题,本申请实施例提供一种图像增强模型训练方法、装置、电子设备及存储介质,用以解决目前图像处理算法无法有效保留商品细节和清晰度,同时无法针对某一区域进行注意力增强的问题

Benefits of technology

[0073]本申请实施例提供一种图像增强模型训练方法、装置、电子设备及存储介质,获取训练数据集,训练数据集中包括多个待训练图像样本,以及每个待训练图像样本对应的预标注注意区域;对每个待训练图像样本进行特征提取,得到文本掩膜和商品对象掩膜,文本掩膜用于指示每个待训练图像样本中的文本区域,商品对象掩膜用于指示每个待训练图像样本中的商品区域;通过文本掩膜和训练数据集对初始卷积神经网络中的第一多提示网络进行训练,以及通过商品对象掩膜和训练数据集对初始卷积神经网络中的第二多提示网络进行训练,得到目标图像增强模型;其中,第一多提示网络用于对图像进行质量增强,第二多提示网络用于对图像进行注意力重定向。通过该方案,训练得到了一个可以同时对图像进行质量增强以及注意力重定向的模型,并且在模型中加入了多提示网络,可以同时关注到注意力区域、任务细节和图像内容,显著提高了模型的迁移学习效率和泛化能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608404B_ABST
    Figure CN118608404B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose an image enhancement model training method and device, electronic equipment and storage medium, applied to the technical field of image enhancement, which can solve the problem that current image processing algorithms cannot effectively retain the details and clarity of goods and cannot enhance the attention of a certain area. The method comprises: obtaining a training data set; performing feature extraction on each to-be-trained image sample to obtain a text mask and a commodity object mask; training a first multi-prompt network in an initial convolutional neural network through the text mask and the training data set, and training a second multi-prompt network in the initial convolutional neural network through the commodity object mask and the training data set to obtain a target image enhancement model; wherein the first multi-prompt network is used for quality enhancement of an image, and the second multi-prompt network is used for attention redirection of the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image enhancement technology, and in particular to an image enhancement model training method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the widespread adoption of online shopping, e-commerce images play a crucial role in product display and attracting consumers. However, due to limitations in online storage and bandwidth, e-commerce images often require compression, leading to decreased image quality and unclear product details and advertising text. Furthermore, key products in some images are not prominently displayed, failing to effectively capture consumer attention. These issues limit the practicality and advertising value of e-commerce images.

[0003] While existing algorithms for improving image quality and attention redirection have achieved some success in processing natural images, their application in e-commerce image processing has limitations. E-commerce images are unique in that they typically include product-oriented layouts and product-related text, elements crucial for conveying product information. However, current image processing techniques are not optimized for these specific elements, resulting in an inability to effectively preserve product details and clarity while ensuring the readability of text information when processing e-commerce images. Summary of the Invention

[0004] To address, or at least partially address, the aforementioned technical problems, this application provides an image enhancement model training method, apparatus, electronic device, and storage medium to solve the problem that current image processing algorithms cannot effectively preserve product details and clarity, and cannot perform attention enhancement on a specific area.

[0005] To achieve the above objectives, the technical solutions provided in this application are as follows:

[0006] In a first aspect, embodiments of this application provide an image enhancement model training method, the image enhancement model training method comprising: acquiring a training dataset, the training dataset including multiple image samples to be trained, and a pre-annotated attention region corresponding to each image sample to be trained;

[0007] Feature extraction is performed on each training image sample to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample.

[0008] The first multi-cue network in the initial convolutional neural network is trained using the text mask and the training dataset, and the second multi-cue network in the initial convolutional neural network is trained using the product object mask and the training dataset to obtain the target image enhancement model.

[0009] The first multi-cue network is used to enhance the quality of the image, and the second multi-cue network is used to redirect the attention of the image.

[0010] As an optional implementation, in a first aspect of the embodiments of this application, the step of extracting features from each image sample to be trained to obtain a text mask and a product object mask includes:

[0011] For each training image sample, text extraction and image extraction are performed to obtain the text mask and image mask;

[0012] Obtain task information input by the user, and determine text features based on the task information;

[0013] The product object mask is determined based on the cosine similarity between the image mask and the text features.

[0014] As an optional implementation, in a first aspect of this application, determining the product object mask based on the cosine similarity between the image mask and the text features includes:

[0015] Calculate the cosine similarity between each image mask and the text features;

[0016] The image mask with the highest cosine similarity is determined as the product object mask.

[0017] As an optional implementation, in a first aspect of this application, training a first multi-cue network in an initial convolutional neural network using the text mask and the training dataset, and training a second multi-cue network in the initial convolutional neural network using the product object mask and the training dataset to obtain a target image enhancement model, includes:

[0018] The text mask and the plurality of training image samples are input into a first multi-cue network to obtain a first visual cue, and the product object mask and the plurality of training image samples are input into a second multi-cue network to obtain a second visual cue;

[0019] The first visual cue and the plurality of image samples to be trained are input into the backbone network and the first decoder network to obtain image samples with image enhancement.

[0020] The second visual cue and the image-enhanced training image sample are input into the backbone network and the second decoder network to obtain the attention-redirected image sample. The second decoder network includes multiple fully connected layers, each corresponding to multiple parameters of the image editing operation.

[0021] The first multi-cue network and the second multi-cue network are trained based on the image samples after image enhancement and the image samples after attention redirection to obtain the target image enhancement model.

[0022] As an optional implementation, in a first aspect of the embodiments of this application, the step of inputting the text mask and the plurality of image samples to be trained into a first multi-cue network to obtain a first visual cue includes:

[0023] The text mask is input into the mask encoder in the first multi-cue network to obtain the first attention cue;

[0024] The multiple image samples to be trained are input into the content encoder in the first multi-cue network to obtain the first content cue.

[0025] Based on the task information entered by the user, the first task prompt is obtained;

[0026] The first visual cue is determined based on the first attention cue, the first content cue, and the first task cue;

[0027] And, the step of inputting the product object mask and the plurality of training image samples into the second multi-cue network to obtain the second visual cue includes:

[0028] The product object mask is input into the mask encoder in the second multi-cue network to obtain the second attention cue;

[0029] The multiple image samples to be trained are input into the content encoder in the second multi-cue network to obtain the second content cue;

[0030] Based on the task information input by the user, a second task prompt is obtained;

[0031] The second visual cue is determined based on the second attention cue, the second content cue, and the second task cue.

[0032] As an optional implementation, in a first aspect of this application, training the first multi-cue network and the second multi-cue network respectively based on the image-enhanced image samples and the attention-redirected image samples to obtain the target image enhancement model includes:

[0033] A first loss function is determined based on the image samples after image enhancement and the pre-labeled attention regions corresponding to each image sample to be trained; and a second loss function is determined based on the image samples after attention redirection and the pre-labeled attention regions corresponding to each image sample to be trained.

[0034] The first multi-cue network is optimized based on the first loss function, and the second multi-cue network is optimized based on the second loss function to obtain the target image enhancement model.

[0035] As an optional implementation, in a first aspect of this application, determining the second loss function based on the attention-redirected image samples and the pre-labeled attention regions corresponding to each training image sample includes:

[0036] The saliency loss is determined based on the target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained;

[0037] The target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained are input into the authenticity evaluation network to obtain the authenticity loss;

[0038] The second loss function is determined based on the significance loss and the truth loss.

[0039] Secondly, embodiments of this application provide an image enhancement model training device, the image enhancement model training device comprising: an acquisition module, used to acquire a training dataset, the training dataset including multiple image samples to be trained, and a pre-annotated attention region corresponding to each image sample to be trained;

[0040] The processing module is used to extract features from each training image sample to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample.

[0041] The processing module is further configured to train the first multi-cue network in the initial convolutional neural network using the text mask and the training dataset, and to train the second multi-cue network in the initial convolutional neural network using the product object mask and the training dataset, to obtain a target image enhancement model;

[0042] The first multi-cue network is used to enhance the quality of the image, and the second multi-cue network is used to redirect the attention of the image.

[0043] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to perform text extraction and image extraction on each image sample to be trained, to obtain the text mask and the image mask;

[0044] The processing module is specifically used to acquire task information input by the user and determine text features based on the task information;

[0045] The processing module is specifically used to determine the product object mask based on the cosine similarity between the image mask and the text features.

[0046] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to calculate the cosine similarity between each image mask and the text feature;

[0047] The processing module is specifically used to determine the product object mask based on the image mask with the highest cosine similarity.

[0048] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to input the text mask and the plurality of image samples to be trained into a first multi-cue network to obtain a first visual cue, and to input the product object mask and the plurality of image samples to be trained into a second multi-cue network to obtain a second visual cue;

[0049] The processing module is specifically used to input the first visual cue and the plurality of image samples to be trained into the backbone network and the first decoder network to obtain image samples after image enhancement.

[0050] The processing module is specifically used to input the second visual cue and the image-enhanced training image sample into the backbone network and the second decoder network to obtain the attention-redirected image sample. The second decoder network includes multiple fully connected layers, which correspond to multiple parameters of the image editing operation.

[0051] The processing module is specifically used to train the first multi-cue network and the second multi-cue network respectively based on the image samples after image enhancement and the image samples after attention redirection, so as to obtain the target image enhancement model.

[0052] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to input the text mask into the mask encoder in the first multi-cue network to obtain a first attention cue;

[0053] The processing module is specifically used to input the plurality of image samples to be trained into the content encoder in the first multi-cue network to obtain the first content cue;

[0054] The processing module is specifically used to obtain a first task prompt based on the task information input by the user;

[0055] The processing module is specifically used to determine the first visual cue based on the first attention cue, the first content cue, and the first task cue;

[0056] The processing module is specifically used to input the product object mask into the mask encoder in the second multi-cue network to obtain the second attention cue;

[0057] The processing module is specifically used to input the plurality of image samples to be trained into the content encoder in the second multi-cue network to obtain the second content cue;

[0058] The processing module is specifically used to obtain a second task prompt based on the task information input by the user;

[0059] The processing module is specifically used to determine the second visual cue based on the second attention cue, the second content cue, and the second task cue.

[0060] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically configured to determine a first loss function based on the image samples after image enhancement and the pre-labeled attention regions corresponding to each image sample to be trained, and to determine a second loss function based on the image samples after attention redirection and the pre-labeled attention regions corresponding to each image sample to be trained;

[0061] The processing module is specifically used to optimize the first multi-cue network according to the first loss function and optimize the second multi-cue network according to the second loss function to obtain the target image enhancement model.

[0062] As an optional implementation, in a second aspect of the embodiments of this application, the processing module is specifically used to determine the saliency loss based on the target attention region in the attention-redirected image sample and the pre-labeled attention region corresponding to each training image sample;

[0063] The processing module is specifically used to input the target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained into the authenticity evaluation network to obtain the authenticity loss.

[0064] The processing module is specifically used to determine the second loss function based on the saliency loss and the true loss.

[0065] Thirdly, embodiments of this application provide an electronic device, the electronic device comprising:

[0066] Memory containing executable program code;

[0067] A processor coupled to the memory;

[0068] The processor calls the executable program code stored in the memory to execute the image enhancement model training method in the first aspect of the embodiments of this application.

[0069] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that causes a computer to execute the image enhancement model training method of the first aspect of embodiments of this application. The computer-readable storage medium includes ROM / RAM, a magnetic disk, or an optical disk, etc.

[0070] Fifthly, embodiments of this application provide a computer program product that, when run on a computer, causes the computer to perform some or all of the steps of any of the methods of the first aspect.

[0071] Sixthly, embodiments of this application provide an application publishing platform for publishing computer program products, wherein when the computer program product is run on a computer, the computer performs some or all of the steps of any of the methods of the first aspect.

[0072] Compared with the prior art, the embodiments of this application have the following beneficial effects:

[0073] This application provides a method, apparatus, electronic device, and storage medium for training an image enhancement model. The method involves acquiring a training dataset, which includes multiple image samples to be trained and pre-labeled attention regions corresponding to each image sample. Feature extraction is performed on each image sample to obtain a text mask and a product object mask. The text mask indicates the text region in each image sample, and the product object mask indicates the product region. A first multi-cue network in an initial convolutional neural network is trained using the text mask and the training dataset, and a second multi-cue network in the same network is trained using the product object mask and the training dataset to obtain a target image enhancement model. The first multi-cue network is used for image quality enhancement, and the second multi-cue network is used for image attention redirection. This approach trains a model capable of simultaneously enhancing image quality and redirecting attention. The addition of a multi-cue network allows the model to simultaneously focus on attention regions, task details, and image content, significantly improving the model's transfer learning efficiency and generalization ability. Attached Figure Description

[0074] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0075] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0076] Figure 1 This is a flowchart illustrating an image enhancement model training method provided in an embodiment of this application. Figure 1 ;

[0077] Figure 2 This is a flowchart illustrating an image enhancement model training method provided in an embodiment of this application. Figure 2 ;

[0078] Figure 3 This is a schematic diagram of an image enhancement model training method provided in an embodiment of this application. Figure 1 ;

[0079] Figure 4 This is a flowchart illustrating an image enhancement model training method provided in an embodiment of this application. Figure 3 ;

[0080] Figure 5This is a schematic diagram of an image enhancement model training method provided in an embodiment of this application. Figure 2 ;

[0081] Figure 6 This is a schematic diagram of an image enhancement model training method provided in an embodiment of this application. Figure 3 ;

[0082] Figure 7 This is a schematic diagram showing the result of an image enhancement model training method provided in an embodiment of this application;

[0083] Figure 8 This is a schematic diagram of the structure of an image enhancement model training device provided in an embodiment of this application;

[0084] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0085] To better understand the above-mentioned objectives, features, and advantages of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features of this application can be combined with each other. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0086] The terms "first" and "second," etc., used in the specification and claims of this application are used to distinguish different objects, rather than to describe a specific order of objects.

[0087] The terms “comprising” and “having”, and any variations thereof, in this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to such process, method, product, or device.

[0088] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0089] With the widespread adoption of online shopping, e-commerce images play a crucial role in product display and consumer attraction. However, due to limitations in online storage and bandwidth, e-commerce images often require compression, leading to decreased image quality and unclear product details and advertising text. Furthermore, key products in some images are not prominently displayed, failing to effectively capture consumer attention. These issues limit the practicality and advertising value of e-commerce images.

[0090] In the field of image processing, researchers have proposed various techniques to improve image quality. For example, the SwinIR algorithm, which employs the Swin Transformer structure, effectively restores images by addressing various degradation problems such as noise and blur. SwinIR improves the quality of the restored image by utilizing both local and global features during the restoration process, particularly excelling in preserving detail and texture.

[0091] Another related study is the Realistic Saliency Guided Image Enhancement method, which focuses on attention redirection in images. By adjusting the distribution of saliency in an image, this method guides the viewer's visual attention, making key areas of the image more prominent. This approach improves the visual appeal of an image by optimizing the saliency map, making important information more noticeable.

[0092] While existing methods have achieved some success in processing natural images, their application in e-commerce image processing has limitations. E-commerce images are unique in that they typically include product-oriented layouts and product-related text, elements crucial for conveying product information. However, current image processing techniques are not optimized for these specific elements, resulting in an inability to effectively preserve product details and clarity while ensuring the readability of text information when processing e-commerce images.

[0093] Furthermore, key products in e-commerce images typically need to be highlighted to attract consumers' attention, but existing attention retargeting techniques primarily target natural images and lack specific processing for product salience in e-commerce images. This results in important product features potentially failing to be properly emphasized in e-commerce images, thereby reducing the advertising effectiveness of the images and consumers' willingness to purchase.

[0094] High-quality e-commerce image datasets are relatively small in scale, and it's difficult to obtain a large number of images with both realism enhancement and attention redirection annotations. This data scarcity limits the application of deep learning methods in e-commerce image enhancement, as these methods typically require a large amount of labeled data to train effective models. The lack of sufficient training data not only affects model performance but also limits the algorithm's ability to generalize to different types of e-commerce images.

[0095] Therefore, the shortcomings of existing technologies mainly stem from their inability to fully adapt to the specific needs of e-commerce images, and the lack of sufficient high-quality e-commerce image data to support and optimize image processing algorithms. These limiting factors collectively result in unsatisfactory processing effects for e-commerce images in terms of quality improvement and attention guidance.

[0096] To address some or all of the aforementioned technical problems, this application provides an image enhancement model training method, apparatus, electronic device, and storage medium. The method involves acquiring a training dataset, which includes multiple image samples to be trained and pre-labeled attention regions corresponding to each image sample. Feature extraction is performed on each image sample to obtain a text mask and a product object mask. The text mask indicates the text region in each image sample, and the product object mask indicates the product region. A first multi-cue network in an initial convolutional neural network is trained using the text mask and the training dataset, and a second multi-cue network in the same network is trained using the product object mask and the training dataset to obtain a target image enhancement model. The first multi-cue network is used for image quality enhancement, and the second multi-cue network is used for image attention redirection. This approach trains a model capable of simultaneously enhancing image quality and redirecting attention. The addition of a multi-cue network allows the model to simultaneously focus on attention regions, task details, and image content, significantly improving the model's transfer learning efficiency and generalization ability.

[0097] like Figure 1 As shown, Figure 1 A flowchart of an image enhancement model training method provided in this application embodiment, the method may include the following steps:

[0098] 101. Obtain the training dataset.

[0099] In this embodiment, the training dataset may include: multiple training image samples and pre-annotated attention regions corresponding to each training image sample. The training image samples may be e-commerce images, including product images, product names, product descriptions, and other information. The pre-annotated attention regions may be pre-annotated by staff on the training image samples. Based on the characteristics of e-commerce images, in order to showcase products and attract consumers, it is necessary to enhance or enlarge certain areas of the image. Therefore, to train the model, staff can pre-annotate the samples, that is, mark the attention regions that need to be highlighted in the training image samples.

[0100] 102. Extract features from each image sample to be trained to obtain text mask and product object mask.

[0101] In this embodiment of the application, feature extraction can be performed on each training image sample. Since the training image sample includes text content and image content, they are extracted separately to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample.

[0102] It should be noted that image masking involves using selected images, graphics, or objects to occlude (fully or partially) an image to be processed, thereby controlling the area or process of image processing. In image processing, a mask is typically a binary or Boolean image of the same size as the original image, where the selected area is marked as 1 (or True), while the remaining area is marked as 0 (or False).

[0103] In digital image processing, a mask is a two-dimensional matrix array, sometimes also using multi-valued images. Its main applications include: extracting regions of interest (ROIs): multiplying a pre-made ROI mask with the image to be processed to obtain the ROI image, where image values ​​within the ROI remain unchanged, while image values ​​outside the ROI are all 0; masking: using a mask to shield certain regions of an image, preventing them from participating in processing or parameter calculations, or processing or statistically analyzing only the shielded areas; structural feature extraction: detecting and extracting structural features in the image similar to the mask using similarity variables or image matching methods; and creating images of specific shapes: defining regions of specific shapes using masks and processing them to generate images of those specific shapes.

[0104] Additionally, when using masks for image processing, it's important to note that the mask must have the same size and resolution as the original image, and that regions with a value of 1 represent areas that need processing, while regions with a value of 0 represent areas that do not need processing. Furthermore, when using masks for image processing, it's necessary to select the appropriate operation (such as filtering, edge detection, region extraction, etc.) based on specific requirements.

[0105] In some embodiments, feature extraction is performed on each training image sample to obtain a text mask and a product object mask, specifically through... Figure 2 This is achieved through steps 1021 to 1023 shown.

[0106] 1021. Perform text extraction and image extraction on each training image sample to obtain a text mask and an image mask.

[0107] 1022. Obtain the task information input by the user and determine the text features based on the task information.

[0108] 1023. Determine the product object mask based on the cosine similarity between the image mask and text features.

[0109] In this embodiment of the application, the above steps can be implemented through an interactive condition unit, which enables the user to specify the desired object and then generate the corresponding attention area, which can associate the user's specific product name with the image area of ​​different objects.

[0110] In some embodiments, such as Figure 3 As shown, in the interactive conditional unit, each image sample to be trained is input into the attention generator to produce a text mask M. t As text attention (i.e., text mask), and N dominant object masks (i.e., the image mask). The image mask can then be input into the image encoder E. img In order to extract visual features Specifically, it can be expressed as:

[0111]

[0112] Here, I represents the input image (the image sample to be trained), and ⊙ denotes element-wise multiplication. Element-wise multiplication (also known as element-level multiplication or Hadamard product) is an operation between two matrices or arrays of the same shape, where each element of the output array is the product of the corresponding elements of the input array. This is very useful in many mathematical and computational applications, especially in deep learning, image processing, and other types of matrix computations.

[0113] Meanwhile, for the text side, the user-specific product name is used to construct a pre-designed prompt template, such as "a product photo of [name]", which is then input into the text encoder E. text In order to extract text features f text :

[0114] f text =E text ("aproductphotoof[name]")

[0115] In some embodiments, the above-mentioned prompt templates may include multiple formats. To compare the performance of different prompt templates, mean AR (mAR), recall (R), and precision (P) can be used to measure the difference between the generated object attention region and the annotated product. For comparison, a simple convolutional binary classifier is introduced, which is trained to determine whether each input object mask is a product or something else. As shown in Table 1, the interactive graph proposed by this method performs better than the baseline model. Furthermore, the template "a product photo of [name]" enables the interactive conditional module of this method to achieve optimal performance, and therefore this template was ultimately selected in this method.

[0116] Table 1. Performance Comparison Results of Different Prompt Templates

[0117]

[0118] In obtaining text features f of user-specific objects text Visual features of all object candidates Next, the cosine similarity score is calculated to match the object mask to the product name. Here, since the user input can be considered task information—that is, the content the user needs to highlight—text features are determined based on the task information. Then, visual features and text features are compared sequentially to determine the closest visual feature, thus obtaining the product object mask.

[0119] In some embodiments, specifically, the cosine similarity between each image mask and text feature is calculated; the image mask with the highest cosine similarity is determined as the product object mask.

[0120] It should be noted that since the image mask has been transformed into a visual feature after being processed by the image encoder, the cosine similarity between each visual feature and the text feature can be calculated. Then, the visual feature with the highest cosine similarity can be selected, and the image mask corresponding to that visual feature can be determined as the product object mask.

[0121] Specifically, the cosine similarity can be expressed as:

[0122]

[0123] Where ⊙ represents element-wise multiplication, and ‖·‖ represents the Euclidean norm. The Euclidean norm, also known as the L2 norm or L2 standard, is a widely used distance metric in mathematics and computer science. It measures the similarity between vectors and is specifically defined as the length of a vector, which is the square root of the sum of the squares of its elements. For an n-dimensional vector x = (x1, x2, ..., xn), its Euclidean norm is calculated as follows:

[0124] ||x||2=sqrt(x1 2 +x2 2 +…+x n 2 )

[0125] Where ||x||2 represents the Euclidean norm of vector x.

[0126] The Euclidean norm has the following properties: Nonnegativity: For any vector x, the Euclidean norm is always nonnegative. Ternary property: For any vector x and a real number c, the Euclidean norm of c·x is equal to the Euclidean norm of |c|·x. Homogeneity: Similar to the ternary property, for any vector x and a real number c, the Euclidean norm of x is equal to the Euclidean norm of |c|·x. Inverse symmetry: For any vectors x and y, the Euclidean norm of x is equal to the Euclidean norm of y if and only if x and y are opposite vectors. It should be noted that the Euclidean norm can also be 0 when a vector contains zero elements.

[0127] 103. Train the first multi-cue network in the initial convolutional neural network using a text mask and a training dataset, and train the second multi-cue network in the initial convolutional neural network using a commodity object mask and a training dataset to obtain the target image enhancement model.

[0128] In this application, unlike the single visual cue used in most existing methods, a multi-cue network is proposed to learn enhanced visual cues, including three types: attention cues, content cues, and task cues. In this way, attention regions, task details, and image content can be embedded to cue the downstream tasks of the quality enhancement unit and the attention retargeting unit. This significantly improves the transfer learning efficiency and generalization ability of the proposed method.

[0129] In some embodiments, the initial convolutional neural network includes a quality enhancement unit and an attention redirection unit. The quality enhancement unit includes a first multi-cue network, and the attention redirection unit includes a second multi-cue network, which are used to enhance the image quality and redirect the attention, respectively.

[0130] In some embodiments, a first multi-cue network in the initial convolutional neural network is trained using a text mask and a training dataset, and a second multi-cue network in the initial convolutional neural network is trained using a product object mask and a training dataset to obtain a target image enhancement model. Specifically, this can be achieved through... Figure 4 This is achieved through steps 1031 to 1034 shown.

[0131] 1031. Input the text mask and multiple training image samples into the first multi-cue network to obtain the first visual cue, and input the product object mask and multiple training image samples into the second multi-cue network to obtain the second visual cue.

[0132] In the embodiments of this application, the first multi-cue network and the second multi-cue network are essentially the same, both of which integrate multiple cues into a visual cue, but the specific input content is different.

[0133] Combined with the attention mask M in the interactive conditional unit v (M p Or M t It generates attention cues to indicate the attention areas of subsequent quality enhancement units and attention redirection units.

[0134] In some embodiments, the generation processes of the first visual cue and the second visual cue are similar, wherein: a text mask is input into the mask encoder in a first multi-cue network to obtain a first attention cue; multiple training image samples are input into the content encoder in the first multi-cue network to obtain a first content cue; a first task cue is obtained based on task information input by the user; a first visual cue is determined based on the first attention cue, the first content cue, and the first task cue; and a product object mask is input into the mask encoder in a second multi-cue network to obtain a second attention cue; multiple training image samples are input into the content encoder in the second multi-cue network to obtain a second content cue; a second task cue is obtained based on task information input by the user; and a second visual cue is determined based on the second attention cue, the second content cue, and the second task cue.

[0135] In some embodiments, such as Figure 5 As shown, a design with two residual blocks E was created. Res Using residual structures to generate attention cues P att :

[0136] Patt =E UP (E Res (M v ))+M v

[0137] In the formula, E UP It is an upsampling layer with pixel shuffling operations to restore the size of the input image. Where M... v Specifically, this can include text masks and product object masks. The input in the first multi-hint network is a text mask, and the input in the second multi-hint network is a product object mask.

[0138] Depend on Figure 5 As can be seen, the mask input into the multi-cue network is actually processed by a mask encoder, which includes two sets of convolutional layers (residual structures) and an upsampling layer, ultimately yielding attention cues.

[0139] In some embodiments, content prompt P cont ∈R H×W×3 It is generated to improve the generalization ability of the backbone network. Specifically, content cues are generated through a series of convolutional layers E VGG Extracted, specific, content hint P cont It can be represented as:

[0140] P cont =E UP (E VGG (I)),

[0141] Among them, E UP Same as the upsampling layer in the attention cue. Figure 5 As can be seen, the input of multiple training image samples into the multi-cue network is actually processed by the content encoder, which includes multiple sets of convolutional layers, upsampling layers and activation functions, and finally obtains the content cues.

[0142] In some embodiments, define task prompt P task ∈R H×W×3 As a learnable parameter, its size is the same as the input image. In this way, task information can be encoded to prompt the frozen backbone network to perform different tasks, namely quality enhancement units and attention redirection units.

[0143] Depend on Figure 5 As can be seen, the task prompt P task It is generated directly; the task prompt is actually obtained based on the task information entered by the user.

[0144] Finally, the three types of cues mentioned above (content cues, task cues, and attention cues) are combined into the final visual cue P.v Specifically, it can be expressed as:

[0145] P v =P task +λ·P cont +γ·P att

[0146] In the formula, + indicates element-wise addition, and λ and γ are the corresponding weights.

[0147] In some embodiments, the effectiveness of the multi-cue network for quality enhancement was further investigated through three ablation experiments: (1) no content cues; (2) no content and attention cues; and (3) no multi-cue network. The experimental results are shown in Table 2. It can be seen that removing the entire multi-cue network resulted in a PSNR decrease of 0.15 dB. Meanwhile, removing content cues and attention cues resulted in PSNR decreases of 0.03 dB and 0.06 dB, respectively. Similar results were observed for SSIM and LPIPS. This indicates that the developed cue-based framework contributes to the final performance of our method.

[0148] Table 2. Performance Comparison Results of Various Ablation Methods

[0149] Full model 32.53 0.925 0.178 w / o content prompt 32.50 0.925 0.181 w / o conten&attention prompt 32.44 0.924 0.180 w / o multi-prompt net 32.38 0.924 0.180

[0150] 1032. Input the first visual cue and multiple training image samples into the backbone network and the first decoder network to obtain image samples after image enhancement.

[0151] In this embodiment of the application, the aforementioned backbone network E back First decoder network D QE Together with the first multi-cue network, they form a quality enhancement unit, which aims to improve the objective quality of compressed e-commerce images. The backbone network E... back It is pre-trained and frozen, consisting of a set of Swin Transformer blocks and convolutional layers. Additionally, the first decoder network D... QE It is lightweight, has two convolutional layers, and is trainable.

[0152] Specifically, the structure of the mass enhancement unit is as follows: Figure 6 As shown, image sample I and visual cue P QE The features (obtained from the multi-cue network) are combined and then fed into a pre-trained backbone network for feature extraction. Subsequently, the extracted features are input into the decoder to reconstruct a high-quality e-commerce image. Specifically, it can be expressed as:

[0153]

[0154] 1033. Input the second visual cue and the image sample to be trained after image enhancement into the backbone network and the second decoder network to obtain the image sample after attention redirection. The second decoder network includes multiple fully connected layers, which correspond to multiple parameters of the image editing operation.

[0155] In this embodiment of the application, the aforementioned backbone network E back First decoder network D AR Together with the second multi-cue network, they form the attention redirection unit, which improves the salience of the desired object by performing pixel-level operations on the object region. It can be seen that the attention redirection unit and the quality enhancement unit use the same backbone network E. back Unlike the quality enhancement unit, the attention retargeting unit's decoder includes global average pooling and five fully connected layers. Here, each fully connected layer learns parameters for image editing operations, including adjustments to color, saturation, contrast, white balance, and sharpness. These parameters are then used to manipulate the object region using corresponding conventional image processing methods.

[0156] Specifically, the structure of the attention redirection unit is as follows: Figure 6 As shown, based on the output image of the quality enhancement unit and learned visual cues P AR The final output image is obtained. It can be represented as:

[0157]

[0158] 1034. Based on the image samples after image enhancement and the image samples after attention redirection, train the first multi-cue network and the second multi-cue network respectively to obtain the target image enhancement model.

[0159] In this embodiment of the application, in order to optimize the model, the first multi-cue network and the second multi-cue network in the quality enhancement unit and the attention redirection unit can be trained and optimized respectively to obtain the final target image enhancement model.

[0160] In some embodiments, a first loss function is determined based on the image samples after image enhancement and the pre-labeled attention regions corresponding to each image sample to be trained, and a second loss function is determined based on the image samples after attention redirection and the pre-labeled attention regions corresponding to each image sample to be trained; a first multi-cue network is optimized based on the first loss function, and a second multi-cue network is optimized based on the second loss function to obtain the target image enhancement model.

[0161] It should be noted that, in order to train the quality enhancement module, the original image I can be measured. rawand image enhancement The Charbonnier distance between them. Specifically, the quality enhancement loss. The (first loss function) can be expressed as:

[0162]

[0163] Where ∈ is a constant factor with a value of 1 × 10⁻⁶. -3 After obtaining the loss function, the parameters of the first multi-cue network can be optimized, and training can continue until convergence.

[0164] It should be noted that an attention redirection loss was designed to obtain true results with the desired significance distribution. To supervise the AR module. However, unlike the optimization method for the first multi-cue network in the quality enhancement unit, the optimization of the second multi-cue network in the attention redirection unit mainly considers two factors: the saliency change ΔS of the input and edited images and the realism change ΔR.

[0165] In some embodiments, a saliency loss is determined based on the target attention region in the attention-redirected image samples and the pre-labeled attention region corresponding to each training image sample; the target attention region in the attention-redirected image samples and the pre-labeled attention region corresponding to each training image sample are input into a realism evaluation network to obtain a realism loss; a second loss function is determined based on the saliency loss and the realism loss.

[0166] Specifically, the saliency change ΔS represents the average difference between the saliency maps of the input and edited images in the attention region:

[0167]

[0168] Where E[·] is the expectation function, E sal (·) is a deep significance prediction model.

[0169] Furthermore, the actual change ΔR is expressed as:

[0170]

[0171] Among them, E rel (·) is a realism evaluation network containing multiple convolutional and fully connected layers, where s is a preset offset. To simultaneously optimize saliency and realism, the aforementioned saliency change ΔS and realism change ΔR can be mixed to obtain a second loss function, which is then used to optimize the second multi-cue network.

[0172] Specifically, attention redirection loss The second loss function can be expressed as:

[0173]

[0174] Among them, w s It is a hyperparameter used to adjust the degree of change in significance and truth.

[0175] This application provides an image enhancement model training method. The method involves acquiring a training dataset, which includes multiple training image samples and pre-labeled attention regions corresponding to each sample. Feature extraction is performed on each training image sample to obtain a text mask and a product object mask. The text mask indicates the text region in each training image sample, and the product object mask indicates the product region. A first multi-cue network in an initial convolutional neural network is trained using the text mask and the training dataset, and a second multi-cue network in the same network is trained using the product object mask and the training dataset to obtain a target image enhancement model. The first multi-cue network is used for image quality enhancement, and the second multi-cue network is used for image attention redirection. This approach trains a model capable of simultaneously enhancing image quality and redirecting attention. The addition of a multi-cue network allows the model to simultaneously focus on attention regions, task details, and image content, significantly improving the model's transfer learning efficiency and generalization ability.

[0176] In some embodiments, during the training process of the target image enhancement model described above, the Adam optimizer and an initial learning rate of 2×10⁻⁶ can be used. -3 The quality enhancement unit and attention redirection unit were optimized. Furthermore, multi-step scheduling was applied to automatically adjust the learning rate. For the multi-cue network, both λ and γ were set to 1. For the quality enhancement unit, the patch size was set to 126×126. For the attention redirection unit, the input image was resized to 384×384 to generate editing parameters, which were then applied to the original image. The loss function for the attention redirection unit includes s and w. s The values ​​were set to 0.25 and 0.3, respectively. All experiments were conducted on a server equipped with two NVIDIA 4090 GPUs.

[0177] In addition, for model training, this method randomly selected 871 and 350 images from the database as the training and validation sets, respectively. For the inference phase, this method uses the remaining 311 images from the database as the test set. Simultaneously, to evaluate the generalization ability of this method, 324 and 321 test images were randomly selected from the M5Product and K3M databases, respectively. Similar to the databases proposed in this method, the e-commerce images in M5Product and K3M are also compressed using a JPEG codec with QF values ​​of 10, 20, 30, and 40. To evaluate attention retargeting, this method further manually annotated the attention regions in M5Product and K3M, identical to the annotations in the database used in this method.

[0178] Furthermore, our proposed method was compared with several other learning-based quality enhancement methods, including SwinIR, MIRNet-v2, DnCNN, and DCAD. All compared methods were retrained and evaluated on the same data as our proposed method. Subsequently, the quality enhancement performance of our proposed and compared methods was evaluated using three metrics: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measurement (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). Note that higher PSNR and SSIM, and lower LPIPS, indicate better quality. The results for the three databases at four QF values ​​are shown in Table 3. It can be observed that our proposed method achieved the best performance in terms of PSNR, SSIM, and LPIPS across all three databases. For example, in the database proposed by our method, at QFs of 10, 20, and 40, our method improved the PSNR by 0.03 dB, 0.1 dB, and 0.16 dB, respectively, compared to the second-best method. This validates the effectiveness of our proposed method in quality enhancement. More importantly, this method also achieved the highest image quality on the M5Product and K3M databases, indicating that the method has good generalization ability on multiple databases.

[0179] Table 3 Performance comparison of different methods on multiple databases

[0180]

[0181] In addition, two best-performing quality enhancement methods and two state-of-the-art attention redirection methods were selected, and their combination was used as the baseline method. Specifically, based on the results in Table 1, SwinIR and MIRNet-v2 were selected for quality enhancement. For attention redirection, two learning-based attention redirection methods, RelEn and GazeShift, were employed. To make quantitative comparisons, a user study was conducted to evaluate the performance of the proposed method and the baseline methods. Specifically, 105 images were randomly selected from three databases and enhanced using the proposed method and the baseline methods, respectively. Next, the images enhanced by different methods were randomly presented along with the original images, and 20 volunteers were recruited to select the best enhanced image overall. The following aspects were considered in the selection process: (1) the overall image quality was enhanced; (2) the product area was more prominent than the original image; and (3) the enhanced image looked realistic. Then, the proportion of each method selected was calculated. As shown in Table 4 below, the proposed method had a selection rate of 62.14% in the database, which was significantly better than the second-best baselines SwinIR and RelEn's 13.14%. Meanwhile, in two other databases, this method also achieved excellent results of 56.85% and 52.71%, respectively. This demonstrates the effectiveness of this method in quality enhancement and attention retargeting of e-commerce images. Additionally, as... Figure 7 As shown, the masking and processing results of this method and various baseline methods for input images from different databases are presented.

[0182] Table 4: Subjective selection rates of multiple methods and multiple databases among multiple users.

[0183] SwinIR+RelEn 13.14% 14.42% 16.14% SwinIR+GazeShift 5.28% 7.28% 7.00% MIRNet-v2+RelEn 13.00% 16.42% 16.85% MIRNet-v2+GazeShift 6.42% 5.00% 7.71% This method 62.14% 56.85% 52.71%

[0184] like Figure 8 As shown, this application embodiment provides an image enhancement model training device, which may include:

[0185] The acquisition module 801 is used to acquire the training dataset, which includes multiple image samples to be trained and the pre-annotated attention region corresponding to each image sample to be trained.

[0186] The processing module 802 is used to extract features from each training image sample to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample.

[0187] The processing module 802 is also used to train the first multi-cue network in the initial convolutional neural network using a text mask and a training dataset, and to train the second multi-cue network in the initial convolutional neural network using a commodity object mask and a training dataset, to obtain a target image enhancement model.

[0188] The first multi-cue network is used for image quality enhancement, and the second multi-cue network is used for image attention redirection.

[0189] In some embodiments, the processing module 802 is specifically used to perform text extraction and image extraction on each training image sample to obtain a text mask and an image mask;

[0190] The processing module 802 is specifically used to obtain task information input by the user and determine text features based on the task information;

[0191] The processing module 802 is specifically used to determine the product object mask based on the cosine similarity between the image mask and the text features.

[0192] In some embodiments, the processing module 802 is specifically used to calculate the cosine similarity between each image mask and text feature;

[0193] The processing module 802 is specifically used to determine the product object mask based on the image mask with the highest cosine similarity.

[0194] In some embodiments, the processing module 802 is specifically used to input a text mask and multiple image samples to be trained into a first multi-cue network to obtain a first visual cue, and to input a product object mask and multiple image samples to be trained into a second multi-cue network to obtain a second visual cue;

[0195] Processing module 802 is specifically used to input the first visual cue and multiple image samples to be trained into the backbone network and the first decoder network to obtain image samples after image enhancement;

[0196] The processing module 802 is specifically used to input the second visual cue and the image enhancement training image sample into the backbone network and the second decoder network to obtain the attention-redirected image sample. The second decoder network includes multiple fully connected layers, which correspond to multiple parameters of the image editing operation.

[0197] The processing module 802 is specifically used to train the first multi-cue network and the second multi-cue network respectively based on the image samples after image enhancement and the image samples after attention redirection, so as to obtain the target image enhancement model.

[0198] In some embodiments, the processing module 802 is specifically used to input the text mask into the mask encoder in the first multi-cue network to obtain a first attention cue;

[0199] Processing module 802 is specifically used to input multiple image samples to be trained into the content encoder in the first multi-cue network to obtain the first content cue;

[0200] The processing module 802 is specifically used to obtain the first task prompt based on the task information input by the user;

[0201] The processing module 802 is specifically used to determine the first visual cue based on the first attention cue, the first content cue, and the first task cue;

[0202] Processing module 802 is specifically used to input the product object mask into the mask encoder in the second multi-cue network to obtain the second attention cue;

[0203] Processing module 802 is specifically used to input multiple image samples to be trained into the content encoder in the second multi-cue network to obtain the second content cue;

[0204] The processing module 802 is specifically used to obtain a second task prompt based on the task information input by the user;

[0205] The processing module 802 is specifically used to determine the second visual cue based on the second attention cue, the second content cue, and the second task cue.

[0206] In some embodiments, the processing module 802 is specifically configured to determine a first loss function based on the image samples after image enhancement and the pre-labeled attention region corresponding to each image sample to be trained, and to determine a second loss function based on the image samples after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained;

[0207] The processing module 802 is specifically used to optimize the first multi-cue network according to the first loss function and optimize the second multi-cue network according to the second loss function to obtain the target image enhancement model.

[0208] In some embodiments, the processing module 802 is specifically used to determine a saliency loss based on the target attention region in the attention-redirected image sample and the pre-labeled attention region corresponding to each training image sample;

[0209] The processing module 802 is specifically used to input the target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained into the authenticity evaluation network to obtain the authenticity loss.

[0210] The processing module 802 is specifically used to determine the second loss function based on the significance loss and the true loss.

[0211] In this embodiment, each module can implement the image enhancement model training method provided in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0212] like Figure 9 As shown in the embodiments of this application, an electronic device is also provided, which may include:

[0213] Memory 901 storing executable program code;

[0214] Processor 902 coupled to memory 901;

[0215] Specifically, the processor 902 calls the executable program code stored in the memory 901 to execute the image enhancement model training method executed by the electronic device in the above method embodiments.

[0216] This application provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the image enhancement model training method described in the above method embodiments and achieves the same technical effect. To avoid repetition, further details are omitted here.

[0217] This application also provides a computer program product, which stores a computer program. When the computer program is executed by a processor, it implements each process of the image enhancement model training method in the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0218] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.

[0219] It should be understood, in the several embodiments provided in this application, that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0220] In this application, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0221] In this application, memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0222] In this application, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, including permanent and non-permanent, removable and non-removable storage media. The storage medium can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), other types of random access memory (RAM), read-only memory (ROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information that can be accessed by a computing device. As defined in this document, computer-readable media do not include transient media, such as modulated data signals and carrier waves.

[0223] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the term "comprising" or any other variations thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0224] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. Those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application. The above-mentioned multiple embodiments are not necessarily multiple independent embodiments; they are divided into multiple embodiments only to highlight different technical features in different embodiments. Those skilled in the art should understand that the above-mentioned multiple embodiments can also be combined arbitrarily.

[0225] In the various embodiments of this application, it should be understood that the sequence number of each process does not necessarily imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0226] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they can be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0227] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0228] If the aforementioned integrated units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-accessible memory. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several requests to cause a computer device (which can be a personal computer, server, or network device, specifically a processor in the computer device) to execute some or all of the steps of the methods described in the various embodiments of this application.

[0229] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for training an image enhancement model, characterized in that, The method includes: Obtain a training dataset, which includes multiple image samples to be trained and a pre-annotated attention region corresponding to each image sample to be trained; Feature extraction is performed on each training image sample to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample. The first multi-cue network in the initial convolutional neural network is trained using the text mask and the training dataset, and the second multi-cue network in the initial convolutional neural network is trained using the product object mask and the training dataset to obtain the target image enhancement model. The first multi-cue network is used to enhance the quality of the image, and the second multi-cue network is used to redirect the attention of the image.

2. The method according to claim 1, characterized in that, The step of extracting features from each training image sample to obtain a text mask and a product object mask includes: For each training image sample, text extraction and image extraction are performed to obtain the text mask and image mask; Obtain task information input by the user, and determine text features based on the task information; The product object mask is determined based on the cosine similarity between the image mask and the text features.

3. The method according to claim 2, characterized in that, Determining the product object mask based on the cosine similarity between the image mask and the text features includes: Calculate the cosine similarity between each image mask and the text features; The image mask with the highest cosine similarity is determined as the product object mask.

4. The method according to claim 1, characterized in that, The step of training a first multi-cue network in the initial convolutional neural network using the text mask and the training dataset, and training a second multi-cue network in the initial convolutional neural network using the product object mask and the training dataset to obtain a target image enhancement model includes: The text mask and the plurality of training image samples are input into a first multi-cue network to obtain a first visual cue, and the product object mask and the plurality of training image samples are input into a second multi-cue network to obtain a second visual cue; The first visual cue and the plurality of image samples to be trained are input into the backbone network and the first decoder network to obtain image samples with image enhancement. The second visual cue and the image-enhanced training image sample are input into the backbone network and the second decoder network to obtain the attention-redirected image sample. The second decoder network includes multiple fully connected layers, each corresponding to multiple parameters of the image editing operation. The first multi-cue network and the second multi-cue network are trained based on the image samples after image enhancement and the image samples after attention redirection to obtain the target image enhancement model.

5. The method according to claim 4, characterized in that, The step of inputting the text mask and the plurality of image samples to be trained into a first multi-cue network to obtain a first visual cue includes: The text mask is input into the mask encoder in the first multi-cue network to obtain the first attention cue; The multiple image samples to be trained are input into the content encoder in the first multi-cue network to obtain the first content cue; Based on the task information entered by the user, the first task prompt is obtained; The first visual cue is determined based on the first attention cue, the first content cue, and the first task cue; And, the step of inputting the product object mask and the plurality of training image samples into the second multi-cue network to obtain the second visual cue includes: The product object mask is input into the mask encoder in the second multi-cue network to obtain the second attention cue; The multiple image samples to be trained are input into the content encoder in the second multi-cue network to obtain the second content cue; Based on the task information input by the user, a second task prompt is obtained; The second visual cue is determined based on the second attention cue, the second content cue, and the second task cue.

6. The method according to claim 4, characterized in that, The step of training the first multi-cue network and the second multi-cue network respectively based on the image-enhanced image samples and the attention-redirected image samples to obtain the target image enhancement model includes: A first loss function is determined based on the image samples after image enhancement and the pre-labeled attention regions corresponding to each image sample to be trained; and a second loss function is determined based on the image samples after attention redirection and the pre-labeled attention regions corresponding to each image sample to be trained. The first multi-cue network is optimized based on the first loss function, and the second multi-cue network is optimized based on the second loss function to obtain the target image enhancement model.

7. The method according to claim 6, characterized in that, The step of determining the second loss function based on the attention-redirected image samples and the pre-labeled attention regions corresponding to each training image sample includes: The saliency loss is determined based on the target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained; The target attention region in the image sample after attention redirection and the pre-labeled attention region corresponding to each image sample to be trained are input into the authenticity evaluation network to obtain the authenticity loss; The second loss function is determined based on the significance loss and the truth loss.

8. An image enhancement model training device, characterized in that, include: The acquisition module is used to acquire the training dataset, which includes multiple image samples to be trained and a pre-annotated attention region corresponding to each image sample to be trained. The processing module is used to extract features from each training image sample to obtain a text mask and a product object mask. The text mask is used to indicate the text region in each training image sample, and the product object mask is used to indicate the product region in each training image sample. The processing module is further configured to train the first multi-cue network in the initial convolutional neural network using the text mask and the training dataset, and to train the second multi-cue network in the initial convolutional neural network using the product object mask and the training dataset, to obtain a target image enhancement model; The first multi-cue network is used to enhance the quality of the image, and the second multi-cue network is used to redirect the attention of the image.

9. An electronic device, characterized in that, include: Memory containing executable program code; and the processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the image enhancement model training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, include: The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the image enhancement model training method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Certificate classification method, device and equipment and storage medium

    CN117115833A

  • Blind image quality evaluation method and device based on attention and multiple tasks, equipment and medium

    CN118014962A