Image processing method, image display method, equipment, storage medium and product

By retrieving similar images from the image database and merging the foreground and background, and selecting the target image in combination with click-through rate prediction, the problem of insufficient visual effects in the existing technology is solved and image generation with high click-through rate is achieved.

CN120705348APending Publication Date: 2025-09-26RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510830691.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image restoration technologies fail to fully consider the special needs of product images, resulting in insufficient visual effects in the generated product images and an inability to ensure high click-through rates.

Method used

By retrieving similar images from the image database, merging the foreground image and the background image, and selecting the target image using click-through rate prediction, a variety of combined images are generated to ensure the diversity and relevance of the background image.

Benefits of technology

The diversity and click-through rate of images are improved, ensuring that the final selected images have a high click-through rate, and improving the efficiency and practicality of the image processing method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705348A_ABST
    Figure CN120705348A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image processing method, an image display method, equipment, a storage medium and a product, and relates to the technical field of computers, the method comprises the following steps: retrieving each similar image of a predetermined image from a predetermined image database; respectively combining the foreground image in the predetermined image with the background image in each similar image to obtain each combined image; predicting the click rate of each combined image to obtain a click rate prediction result corresponding to each combined image; the click rate prediction result is used for representing the probability that the combined image is clicked by the user after the combined image is displayed; and based on the click rate prediction result of each combined image, selecting a target image to be displayed from each combined image. According to the technical scheme provided by the embodiment of the invention, the effect of generating the image can be directly evaluated, and the finally selected image is ensured to have a relatively high click rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image processing method, an image display method, a device, a storage medium, and a product. Background Art

[0002] On product service platforms, products can be recommended to users by displaying product images. In related technologies, to increase the click-through rate of product images, platforms use image restoration technology to beautify product images provided by merchants and then display the beautified product images.

[0003] However, existing image restoration technologies fail to fully consider the specific requirements of product images, such as the coordination between the background and the product itself, when generating beautified product images. This results in insufficient visual quality and, consequently, a failure to ensure high click-through rates. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, an image display method, a device, a storage medium, and a product to alleviate or solve one or more technical problems existing in the prior art.

[0005] In a first aspect, an embodiment of the present application provides an image processing method, comprising: retrieving similar images of a predetermined image from a predetermined image database; merging the foreground image in the predetermined image with the background images in the similar images to obtain merged images; predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image; the click-through rate prediction result is used to characterize the probability that the merged image will be clicked by the user after the merged image is displayed; and selecting a target image to be displayed from the merged images based on the click-through rate prediction results of the merged images.

[0006] In the second aspect, an embodiment of the present application provides an image display method, including: receiving a product image of a reserved product uploaded by a merchant client of a reserved product service platform; retrieving each similar image of the product image from a predetermined image database; merging the foreground image of the product image with the background image in each similar image to obtain each merged image; predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image, the click-through rate prediction result being used to characterize the probability that the merged image will be clicked by the user after the merged image is displayed; based on the click-through rate prediction results of each merged image, selecting a main product image to be displayed from each merged image; and displaying the main product image on the target page of the reserved product service platform.

[0007] In a third aspect, an embodiment of the present application provides an image processing method, comprising: sending a predetermined image to a server of a predetermined product service platform; receiving a target image sent by the server, wherein the target image is an image obtained by the server executing the image processing method of the first aspect on the predetermined image.

[0008] In a fourth aspect, an embodiment of the present application provides an image display method, comprising: displaying a target page, wherein the display content of the target page includes a target image, and the target image is an image obtained by the server according to the image processing method of the first aspect.

[0009] In the fifth aspect, an embodiment of the present application provides an image processing device, comprising: a retrieval module for retrieving similar images of a predetermined image from a predetermined image database; a merging module for merging the foreground image in the predetermined image with the background images in the similar images to obtain merged images; a prediction module for predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image; the click-through rate prediction result is used to characterize the probability that the merged image will be clicked by the user after the merged image is displayed; and a selection module for selecting a target image to be displayed from the merged images based on the click-through rate prediction results of the merged images.

[0010] In the sixth aspect, an embodiment of the present application provides an image display device, comprising: a receiving module for receiving a product image of a predetermined product uploaded by a merchant client of a predetermined product service platform; a retrieval module for retrieving similar images of the product image from a predetermined image database; a merging module for merging the foreground image of the product image with the background images in the similar images to obtain merged images; a prediction module for predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image, wherein the click-through rate prediction result is used to characterize the probability that the merged image is clicked by the user after the merged image is displayed; a selection module for selecting a main product image to be displayed from the merged images based on the click-through rate prediction results of the merged images; and a display module for displaying the main product image on the target page of the predetermined product service platform.

[0011] In the seventh aspect, an embodiment of the present application provides an image processing device, including: a sending module for sending a predetermined image to a server of a predetermined product service platform; a receiving module for receiving a target image sent by the server, wherein the target image is an image obtained by the server executing the image processing method of the first aspect on the predetermined image.

[0012] In an eighth aspect, an embodiment of the present application provides an image display device, comprising: a display module for displaying a target page, wherein the display content of the target page includes a target image, and the target image is an image obtained by the server according to the image processing method of the first aspect.

[0013] In a ninth aspect, an embodiment of the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements any method of the embodiment of the present application when executing the computer program.

[0014] In a tenth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method of any one of the embodiments of the present application is implemented.

[0015] In the eleventh aspect, an embodiment of the present application provides a computer program product, including a computer program, which implements any method of the embodiments of the present application when executed by a processor.

[0016] According to the image processing method of the embodiment of the present application, images similar to the predetermined image can be retrieved from the image database, which is beneficial to ensuring the diversity and relevance of the background image, and the foreground image of the predetermined image is merged with the background image of the retrieved similar image to generate a plurality of combined images, and then the target image to be displayed is screened from each combined image through click-through rate prediction. Compared with the image restoration technology in the related art, the method of the embodiment of the present application can generate a plurality of combined images by retrieving similar images and merging the foreground and background, thereby increasing the diversity of the image; and click-through rate prediction is beneficial to directly evaluate the effect of the generated image, ensuring that the final selected image has a high click-through rate, which is beneficial to improving the efficiency and practicality of the image processing method.

[0017] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0019] Figure 1 Schematic diagram showing a comparison of click-through rates of food images with the same foreground image and different background images;

[0020] Figure 2 A flowchart of an image processing method according to an embodiment of the present application is shown;

[0021] Figure 3 A schematic diagram showing a database construction of an exemplary embodiment of the present application;

[0022] Figure 4 A schematic diagram illustrating the processing process of image retrieval and background replacement in an exemplary embodiment of the present application;

[0023] Figure 5 A schematic diagram showing the working principle of a click-through rate prediction model according to an exemplary embodiment of the present application;

[0024] Figure 6 Schematic diagram of the processing flow of the image display method according to an embodiment of the present application;

[0025] Figure 7 A schematic diagram illustrating the architecture of an image processing system according to an exemplary embodiment of the present application;

[0026] Figure 8 A flowchart of an image processing method according to another embodiment of the present application is shown;

[0027] Figure 9 A flowchart showing an image display method according to another embodiment of the present application is shown;

[0028] Figure 10 A schematic structural diagram of an image processing device according to an embodiment of the present application is shown;

[0029] Figure 11 A schematic structural diagram of an image display device according to an embodiment of the present application is shown;

[0030] Figure 12 A schematic structural diagram of an image processing device according to an embodiment of the present application is shown;

[0031] Figure 13 A schematic structural diagram of an image display device according to an embodiment of the present application is shown;

[0032] Figure 14 A block diagram of an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0033] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0034] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.

[0035] On product service platforms, product recommendations can be presented to users by displaying product images. When users browse product service platforms, visual factors are often the key factor in determining whether they click. The visual quality of product images can significantly influence user clicks or visit behavior, necessitating the development of visually compelling image generation strategies. Background generation experiments have shown a positive correlation between appropriately chosen background images in product images and user engagement. For example, in the case of food images, an appropriately chosen background is one that visually harmonizes with the main subject of the food, aligns with the style, and enhances the overall appeal of the image. This choice of background plays a crucial role in increasing user engagement (e.g., click-through rate, dwell time, etc.).

[0036] Figure 1 The following schematically illustrates a comparison of click rates of food images with the same foreground image and different background images. Figure 1 Figure 4 shows four groups of food images, arranged from left to right. Each group consists of two images, one above and one below. Each group contains two images with the same foreground image and different background images, meaning the same food image has different background images. The letter above each food image is the image number, and the number below is the corresponding image click-through rate.

[0037] In the first set of images, numbered a1 and a2, the foreground image is a bowl of food. Image a1's background features a table and a woven basket, and its click-through rate is 0.045. Image a2's background is a solid color, and its click-through rate is 0.025.

[0038] In the second set of images, numbered b1 and b2, both feature meat dishes as foreground images. Image b1, with its background featuring a wooden cutting board and wood-grain wallpaper, had a click-through rate of 0.114. Image b2, with its background featuring a solid color, had a click-through rate of 0.094.

[0039] In the third set of images, numbered c1 and c2, the foreground image is a drink. Image c1's background features a wooden tabletop, fruit, and greenery, and has a click-through rate of 0.080. Image c2's background features a solid-color tabletop and wall, and has a click-through rate of 0.066.

[0040] In the fourth set of images, numbered d1 and d2, both feature desserts as foreground images. Image d1's background features curtains and flowers, and has a CTR of 0.040. Image d2's background is a solid color, and has a CTR of 0.029.

[0041] pass Figure 1 As can be seen, images a1, b1, c1, and d1 have richer background elements in their background images. These images are all more visually appealing than the other images in the same group. Consequently, these images also have higher click-through rates. This shows that improving the visual appeal of images, especially the background image design, is crucial for increasing click-through rates.

[0042] It should be noted that image click-through rates can be determined through statistical experiments, for example. In these experiments, food images with the same foreground image are presented to a user group with different background images. The click-through rate for each image can be calculated based on users' click behavior over a certain period of time. This rate reflects users' actual engagement and interest in the different background images.

[0043] In related technologies, several image restoration methods can be used to generate and improve image backgrounds, thereby beautifying product images. For example, Generative Adversarial Networks (GANs) are a deep learning model that can generate new, more realistic images through adversarial learning between two networks. Another example is a Convolutional Neural Network (CNN), which can learn image features and use these features to fill in missing or damaged parts of the image, achieving the purpose of restoring product images.

[0044] However, while these image restoration methods can improve the visual quality of images, they may not necessarily significantly increase user interest in clicking on them. To address this issue, a Creative Generation for Click-Through Rate (CG4CTR) algorithm based on click-through rate optimization can be used to generate product images. Specifically: First, a model is pre-trained using e-commerce data to infuse the pre-trained model with e-commerce knowledge. Second, a reward model is constructed and fine-tuned to enable the pre-trained model to predict the click-through rate of product images. The predicted click-through rate results are then used to drive image optimization, gradually improving the click-through rate of product images. Finally, through multiple iterations of optimization, product images with a high click-through rate are generated. However, this method fails to specifically optimize the special requirements of some product images, such as the coordination between the background and the product body, resulting in insufficient visual quality for the product images. Consequently, this method lacks the ability to generate a high-quality background in the initial image generation stage, failing to achieve both visual quality and a high click-through rate.

[0045] It should be noted that the above-mentioned application scenarios or application examples provided in the embodiments of the present application are for ease of understanding, and the embodiments of the present application do not specifically limit the application of the technical solution. In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0046] The following describes in detail the technical solution of this application and how it solves the aforementioned technical problems using specific embodiments. The several specific embodiments listed can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The following describes the embodiments of this application in detail with reference to the accompanying drawings.

[0047] Figure 2 1 shows a flow chart of an image processing method according to an embodiment of the present application. In some embodiments, the image processing method can be applied to a server. Figure 2 As shown, the method may include steps S201 to S204.

[0048] Step S201: Retrieve similar images of a predetermined image from a predetermined image database.

[0049] Step S202 : Merge the foreground image in the predetermined image with the background images in each similar image to obtain merged images.

[0050] Step S203 , predicting the click rate of each merged image to obtain a click rate prediction result corresponding to each merged image; the click rate prediction result is used to represent the probability of the merged image being clicked by the user after the merged image is displayed.

[0051] Step S204 : selecting a target image to be displayed from each merged image based on the click rate prediction results of each merged image.

[0052] According to the method of the embodiment of the present application, images similar to the predetermined image are retrieved from the image database, which is conducive to ensuring the diversity and relevance of the background image. The foreground image of the predetermined image is merged with the background image of the retrieved similar image to generate a plurality of combined images. Then, the target image to be displayed is screened from each combined image through click-through rate prediction. Compared with the image restoration technology in the related art, the method of the embodiment of the present application can generate a plurality of combined images by retrieving similar images and merging the foreground and background, thereby increasing the diversity of the image. The click-through rate prediction is also conducive to directly evaluating the effect of the generated image, ensuring that the final selected image has a high click-through rate, which is conducive to improving the efficiency and practicality of the image processing method.

[0053] Compared to the creative generation algorithm based on click-through rate optimization that uses the predicted results of click-through rate to drive image optimization, the solution of the embodiment of the present application can generate multiple merged images by retrieving similar images and merging the foreground and background, and then screen the target images to be displayed through click-through rate prediction. This method focuses more on the diversity of image combinations and the verification of actual effects. In addition, compared to the creative generation algorithm based on click-through rate optimization that generates product images with high click-through rates through multiple iterative optimizations, the method of the embodiment of the present application can quickly screen out target images to be displayed with high click-through rates through each merged image and click-through rate prediction. The method of the embodiment of the present application combines the diversity of image combinations and the actual click-through rate effect verification to screen target images, avoiding a complex model optimization process and having higher efficiency and practicality.

[0054] In step S201, a predetermined image is used to display a predetermined item. The predetermined item is the main object. In an image, the main object refers to the object that occupies the primary visual position and serves as the core display object. It is the core content of the image and is usually highlighted to attract the viewer's attention.

[0055] The image database contains images of various items. These include food, clothing, household items, cosmetics, electronics (cell phones, computers, headphones, etc.), decorative items, automotive accessories, and artwork. Images of these items all require a high level of coordination between the background image and the main object. The background must not only be aesthetically pleasing but also highlight the main object's characteristics and style, enhancing the overall visual effect and appeal.

[0056] In some embodiments, step S201 may specifically include: encoding a predetermined image to obtain a first image feature; retrieving at least one target regenerated background image from the regenerated background images corresponding to each sample image contained in the image database, and the similarity between the image feature of the target regenerated background image and the first image feature satisfies a predetermined similarity condition; any regenerated background image and the corresponding sample image contain the same foreground image and a different background image; and using each target regenerated background image as a similar image of the predetermined image.

[0057] As an example, a predetermined image may be encoded using a predetermined encoder, wherein the predetermined encoder includes at least one of a visual task encoder (Swin Transformer), a multimodal task (Contrastive Language-Image Pre-training, CLIP) encoder, a convolutional neural network (CNN), and a visual transformer (Vision Transformer, VIT).

[0058] Among them, the visual task encoder can be used for image feature extraction and is applicable to various visual tasks such as image classification, object detection, and semantic segmentation. The multimodal task encoder is a multimodal model consisting of an image encoder and a text encoder. The image encoder can be used to map images to a shared feature space, while the text encoder is a transformer-based network used to map text to the same feature space. Through contrastive learning, the multimodal task encoder can learn the semantic association between images and text. Convolutional neural networks are deep learning models for image processing that extract image features through multiple layers of convolutional layers and pooling layers; the visual transformer applies the transformer model to image processing by dividing the image into small patches and using the encoder part of the transformer to extract image features.

[0059] It should be understood that in actual applications, other types of encoders can be selected to encode the predetermined image according to actual needs, and the embodiments of the present application do not make specific limitations.

[0060] As an example, the similarity meeting the predetermined similarity condition means that the similarity between the image feature of the target regenerated background image and the first image feature is greater than or equal to a predetermined similarity threshold. The predetermined similarity threshold can be customized according to actual needs and is not specifically limited in this embodiment of the application.

[0061] In this embodiment, regenerated background images similar to the predetermined image are retrieved from an image database. The similarity between the image features of these regenerated background images and the features of the first image satisfies a predetermined similarity condition. This allows for rapid identification of background images visually similar to the predetermined image, ensuring the harmony between the background and the subject. Because each regenerated background image and its corresponding sample image contain the same foreground image and a different background image, this helps maintain foreground image consistency while also increasing visual diversity through the use of different background images.

[0062] In some embodiments, before retrieving at least one target regenerated background image from the regenerated background images corresponding to each sample image contained in the image database, the above steps also include: encoding the pre-acquired sample image to obtain image features of the sample image; generating at least one regenerated background image corresponding to the sample image based on the sample image and a preset text prompt, the preset text prompt being used to describe the sample image in words; and storing the sample image and the corresponding regenerated background images in a predetermined image database.

[0063] For example, the pre-acquired sample image is encoded. The encoder used can be found in the description of the above embodiment and will not be repeated here. The preset text is used to describe the sample image in words, which can provide specific guidance for the generation of the background image and ensure that the generated background image is visually consistent with the sample image.

[0064] Exemplarily, storing the sample images and their corresponding regenerated background images in a predetermined image database can be implemented in various ways. For example, one-to-many key-value pairs can be stored, using the image features of the sample images as keys and the corresponding regenerated background images as values. Alternatively, one-to-many key-value pairs can be stored, using the sample images as keys and the corresponding regenerated background images as values.

[0065] In this image database, key-value storage supports a one-to-many relationship, meaning that one sample image can correspond to multiple regenerated background images. This design increases the diversity of background images, meeting diverse creative needs for pre-defined images. Furthermore, key-value storage is easily scalable, allowing for the easy addition of new sample images and their corresponding regenerated background images to accommodate growing data demands.

[0066] In this embodiment, the technical solution is beneficial to improving the visual effect and diversity of the sample images and improving the data management efficiency through image encoding of the sample images, generation of the regenerated background image and data storage.

[0067] In some embodiments, the above-mentioned step of generating at least one regenerated background image corresponding to the sample image based on the sample image and the preset text prompt may specifically include: removing the background of the sample image to obtain a first foreground image; using a preset image generation model to generate at least one new background for the first foreground image based on the first foreground image and the preset text prompt; merging the first foreground image with each new background respectively to obtain each regenerated background image corresponding to the sample image.

[0068] As an example, background removal of a sample image can be achieved in a variety of ways. For example, the pixel intensity distribution of the image can be used to select one or more thresholds to separate the foreground and background; alternatively, the image can be segmented into foreground and background by analyzing regional features in the image, such as color, texture, etc. Alternatively, a generative adversarial network (GAN) method can be used to generate a segmentation mask for the foreground image and background image, for example, where foreground pixels are marked as 1 (or 255) and background pixels are marked as 0. The generated segmentation mask is applied to the sample image, and the foreground is extracted by performing a pixel-by-pixel logical operation on the mask and the sample image. The logical operation, for example, includes multiplying each pixel value of the sample image by its corresponding mask pixel value, thereby achieving mask-based image segmentation and obtaining the foreground image in the sample image. During this logical operation, if the pixel value in the mask is 255 (indicating foreground), the corresponding pixel value in the original image remains unchanged.

[0069] As an example, the image generation model is used to generate at least one new background for the first foreground image based on the first foreground image and a preset text prompt. A stable diffusion-based neural network architecture (ControlNet) is a neural network architecture used to enhance the control capabilities of the image generation model. It can be used as an enhancement technique in conjunction with the image generation model to improve the controllability and accuracy of the generated image.

[0070] As an example, merging the first foreground image with each new background may include placing the first foreground image at a specific location on each new background to obtain corresponding merged images (in the following embodiments, the merged images may be referred to as composite images). If the size of the first foreground image does not match the size of each new background, the foreground image may be resized to match the size of the background image.

[0071] Figure 3 Schematic diagram of database construction of an exemplary embodiment of the present application is shown. Figure 3 As shown, the database construction method includes the following steps.

[0072] S301, obtaining a sample image.

[0073] In this step, the sample image can be an original picture of the reserved item provided by the merchant client, and can be obtained by the merchant taking a picture of the reserved item.

[0074] S302: Receive a preset text prompt.

[0075] In this step, the preset text prompt is used to describe the sample image, for example Figure 3 The "meat roll" in the.

[0076] For example, let's say the sample image is a meat roll uploaded by a merchant client. The default text prompt might be "meat roll." For example, let's say the sample image is a bowl of noodles placed on a wooden table uploaded by a merchant client. The default text prompt might be "A bowl of steaming noodles placed on a rustic wooden table." The default text can be understood as a descriptive term or prompt (TextPrompt) to describe the sample image.

[0077] S303: Generate each regeneration background.

[0078] Specifically, if Figure 3 As shown in "Background Segmentation Based on Background Subtraction Model" in the previous section, the background image in the sample image is removed using the Segment Anything Model (SAM) to obtain the foreground image. This background subtraction model uses its segmentation capabilities to generate high-precision masks that can be used to accurately isolate the main food from the background, resulting in the foreground image in the sample image. The foreground image is used to indicate the main part of the image.

[0079] like Figure 3 As shown in the "Neural Network Structure Based on Stable Diffusion" section of the paper, the foreground image of the sample image and a preset text prompt are input into the image generation model, which then outputs various regenerated background images. Specifically, the image generation model can generate a background image that matches the foreground image based on the text prompt. In some scenarios, multiple different background images can be generated by adjusting the text prompt. The final output image contains the foreground image of the sample image and the newly generated background image, forming a complete image, namely the regenerated background image.

[0080] exist Figure 3 In the figure, a schematic diagram of the structure of the neural network based on stable diffusion used in the image generation model is also shown. Figure 3 The prompt word (c t) and time (t), the prompt word is the preset text prompt, and time represents the time parameter, which is used to guide the model to perform background separation. Figure 3 The conditions in (c f ), this input condition is used to further refine or adjust the model's output. For example, this input condition is used to guide the model in generating additional information for specific image types. This information can include additional textual hints, image category labels, or descriptions of image style, depending on the model design and application scenario. This information enables the image generation model to generate images that meet specific requirements, improving the accuracy and relevance of the generated images. Figure 3 Two encoders are also shown, each processing different inputs. The role of the encoder is to convert the input data into an internal representation that the model can process. Figure 3 A decoder is also shown in FIG. 4 , which corresponds to the encoder and is used to convert the output of the encoder back to the original data format for further processing or output. Figure 3 As shown in the zero convolution in , it can be used to transfer information between the encoder and decoder without changing the distribution or characteristics of the data. t ,t,c t , c f ), ∈θ can represent the function of the encoder, zt can represent the time-related features or states, t is the time parameter, c t and c f They correspond to the prompt words and the conditions entered above respectively.

[0081] It should be understood that the encoder output is an intermediate step in the model's processing, providing essential feature information for subsequent processing, but it is not the model's final output. The final output is the individual reconstructed background images (including the foreground image in the sample image and the newly generated background image) obtained through the decoder and other processing steps, representing the final result of the model's processing of the input data.

[0082] S304, image encoding.

[0083] In this step, the sample image is encoded to obtain its image features. These image features are in the form of feature vectors, also known as image embeddings. Specifically, image embeddings map the sample image into a low-dimensional vector space, where each point (vector) represents a certain feature representation of the image. Embeddings are often used to measure similarity between images.

[0084] S305: The image is stored in the database.

[0085] In this step, the image database is stored in the form of key-value pairs, using the image embedding representation of the sample image as a key and the respective regenerated background images corresponding to the sample image as a value.

[0086] Through steps S301-S305, a retrieval database containing a large number of background images is pre-built. For each received predetermined image, image segmentation is performed using a background removal model to obtain a foreground image within the sample image. The foreground image, along with the corresponding text prompt, is then input into the image generation model architecture to generate several (e.g., 2-5) different background images as regenerated background images. Each regenerated background image, containing a foreground image and a different background image, is stored as a value in the image database.

[0087] In some embodiments, the step of merging the foreground image in the predetermined image with the background images in each similar image to obtain each merged image in the above-mentioned step S202 may specifically include: extracting the foreground image from the predetermined image to obtain a second foreground image; extracting the background image from each similar image to obtain each background image; performing feature matching on the second foreground image with each background image to determine the matching position of the second foreground image in each background image; and performing image fusion on the second foreground image and each background image based on the matching position of the second foreground image in each background image to obtain each merged image.

[0088] As an example, the method of extracting a foreground image from a predetermined image may refer to the method of performing background removal on a sample image described in the above embodiment, which will not be described in detail here.

[0089] As an example, the matching position refers to: the position where the second foreground image is placed in the background image so that the two can blend naturally. The feature matching of the second foreground image and any background image can be performed by a feature matching algorithm (Scale-Invariant Feature Transform, SIFT). The feature matching algorithm can detect key points in the second foreground image and the background image, and extract descriptors for each extracted key point. By comparing the key point descriptors of the two images (the second foreground image and the background image), matching key points are found and the matching key point positions are recorded. The matching key point positions are used as the matching positions of the second foreground image in each background image.

[0090] As an example, the feature matching between the second foreground image and any background image can also be performed using the Speeded Up Robust Features (SURF) algorithm. This algorithm is similar to the feature matching algorithm described above, but is faster in computation and robust to scale and rotation changes.

[0091] As an example, the second foreground image and the background image can each be segmented into small blocks, the similarities between these blocks compared, and this information used to determine the image alignment. Alternatively, a deep learning model can be pre-trained to identify the location of the foreground image within the background image, and sample images from an image database can be used as training data for the deep learning model. The training process can be implemented through supervised learning.

[0092] It should be understood that in actual applications, the specific implementation method of feature matching can be selected according to actual needs. This embodiment of the present application does not make specific limitations.

[0093] In this embodiment, an image encoder can be used to extract image features (e.g., image embedding representation) from a received predetermined image, and an image retriever can be used to search a retrieval database for at least one previous background image with similar features to the foreground image of the predetermined image. Feature matching and image fusion can achieve seamless fusion between the foreground image and each background image in the predetermined image. Feature matching technology can accurately determine the position of the foreground within the background, improving fusion accuracy. The entire process can be automated, reducing manual intervention and improving processing efficiency.

[0094] In some embodiments, after obtaining each merged image, the method may further include: for any merged image, generating mask information of the merged image, the mask information being used to mark the connection position between the second foreground image and the background image in the merged image; inputting the merged image and the mask information into a pre-trained image diffusion model, and the image diffusion model using the merged image and the mask information to perform denoising processing at the connection position; and updating the merged image using the denoised image.

[0095] For example, a mask can be automatically generated using image segmentation techniques. Alternatively, after the feature matching is completed, mask information for the merged image can be generated based on the matching results. The mask is typically a binary image, where the foreground area is marked as 1 (or 255) and the background area is marked as 0.

[0096] For example, an image diffusion model is a type of generative model that can generate images by simulating a diffusion process. The diffusion process refers to a process that gradually destroys the structural information in an image (such as any of the merged images mentioned above) and eventually converts it into random noise. The goal of the diffusion model is to learn how to reverse this process, that is, to gradually recover the original data from the random noise. For ease of understanding, the working principle of the diffusion model is briefly described below. The diffusion model usually includes two processes: a forward diffusion process and a reverse diffusion process. The forward diffusion process refers to gradually adding Gaussian noise to the image until the image becomes completely random noise. The reverse diffusion process refers to starting from random noise, gradually removing noise, and finally generating a denoised image. This process can be achieved by using a neural network to predict the noise that needs to be removed at each step, thereby generating high-quality and diverse image data.

[0097] For example, the junction location indicates the boundary region where the second foreground image visually merges with the background image in the merged image. This region requires noise reduction to ensure a natural and smooth transition between the foreground and background, without noticeable seams or discontinuities. Mask information is generated to mark the junction location between the second foreground image and the corresponding background image, which helps the model identify the image region requiring processing.

[0098] In this embodiment, mask information is generated, an image diffusion model is used to perform denoising based on the merged image and the mask information, and the merged image is updated using the denoised image. Denoising can significantly improve the quality of the merged image, particularly at the junction of the foreground and background, reducing unnatural transitions and noise. The processed image is visually smoother and more natural, enhancing the overall visual effect of the image and achieving a good fusion of the foreground and background. The mask information allows precise control of the denoising area, ensuring that only the area requiring processing is optimized without affecting other parts of the image. In this embodiment, image processing technology based on mask information and an image diffusion model is beneficial for improving the quality of the merged image, enhancing the visual effect, and increasing the flexibility and practicality of image processing.

[0099] Figure 4 Schematic diagram showing the processing process of image retrieval and background replacement of an exemplary embodiment of the present application. Figure 4 In [1], the image retrieval and background replacement process includes the following steps.

[0100] like Figure 4As shown in “S401, removing background”, the predetermined image is segmented to remove the background image in the predetermined image to obtain a second foreground image of the predetermined image.

[0101] like Figure 4 As shown in “S402, image encoding”, the predetermined image is encoded to obtain a first image feature.

[0102] like Figure 4 As shown in "S403, Image Retrieval", each target reconstructed background image is retrieved from the reconstructed background images corresponding to each sample image included in the image database to obtain each similar image to the predetermined image.

[0103] like Figure 4 As shown in “S404, removing foreground”, image segmentation processing is performed on any retrieved similar image to remove the foreground image in the similar image and obtain the background image in the similar image.

[0104] like Figure 4 As shown in "S405, Feature Matching" in the figure, the second foreground image is feature matched with the background image in the similar image to determine a matching position of the second foreground image in each background image. Based on the matching position, the second foreground image and the background image in the similar image can be matched together to obtain a merged image.

[0105] like Figure 4 As shown in “S406, generating an image to be masked”, mask information is generated according to the matching position, and the mask information is used to mark the connection position between the second foreground image and the background image in the merged image.

[0106] like Figure 4 As shown in “S407, Noise Reduction Processing”, the merged image and mask information are input into a pre-trained image diffusion model, and the image diffusion model uses the merged image and mask information to perform denoising at the connection position to obtain a denoised merged image.

[0107] As an example, Figure 4 The figure also shows the structure of the neural network structure based on stable diffusion used in the image diffusion model. Figure 3 The neural network structure based on stable diffusion shown in is the same or equivalent and will not be described here. The difference is that Figure 4 The input of the neural network structure includes: the merged image and mask information, and the output includes: the merged image after denoising.

[0108] Through the above steps S401-S407, for a given predetermined image, such as a food image, its image features (such as image embedding representation) are extracted using an image encoder, and an image retriever is used to search for at least one previous background image with similar features to the foreground image of the predetermined image in the retrieval database. A foreground image is extracted from the predetermined image, and a background image is extracted from the retrieved image. The foreground image and the background image are then matched together. A pre-trained diffusion model is used to fill in the noise information generated at the connection position after the two are matched, and a final denoised composite image is generated. This processing process not only helps to ensure the high quality of the background image, but also further enhances the naturalness and beauty of the composite image through the generation capability of the diffusion model, thereby achieving high-fidelity background replacement.

[0109] In some embodiments, the above-mentioned step S203 may specifically include: obtaining feature information of the target item, where the target item and the item in the foreground image are the same type of item; inputting each merged image and the feature information into a pre-trained click-through rate prediction model to predict the click-through rate of each merged image through the click-through rate prediction model to obtain a click-through rate prediction result for each merged image.

[0110] For example, the same type of items refers to the target item being of the same type as the item in the foreground image, including but not limited to the same type of food (ramen, ice cream, etc.), the same type of clothing, the same type of household items, etc.

[0111] For example, a click-through rate (CTR) prediction model is a machine learning model used to predict the probability of a user clicking on an image or recommended content. CTR prediction models are widely used in fields such as online advertising and recommendation systems. They work by using a large amount of user behavior data (e.g., user click history, browsing history, search history, etc.) and feature information of displayed objects as input. The model then learns the relationships between these data to predict the probability of a user clicking on the image or recommended content.

[0112] As a specific example, the click-through rate prediction model can be any one of a logistic regression model, a gradient boosting decision tree (GBDT), and a deep learning model. Specifically, the logistic regression model is a simple and effective linear model, which is often used as a baseline model for click-through rate prediction. The gradient boosting decision tree is an integrated learning model that can effectively handle nonlinear features and interactions between features. The deep learning model can be, for example, a deep neural network (DNN), a wide and deep (Wide & Deep) model, a deep factor decomposition machine (DeepFM) model, etc., which can automatically learn the complex relationships between features and achieve better prediction results.

[0113] In this embodiment, the image selection technology based on target item feature information and click-through rate prediction model can realize automated and precise image processing and selection, thereby improving the click-through rate of the images to be displayed.

[0114] In some embodiments, in step S204, based on the click-through rate prediction results of each merged image, a target image to be displayed is selected from each merged image, including: based on the probability value contained in the click-through rate prediction results of each merged image, selecting the merged image whose probability value meets the predetermined sorting condition as the target image of the display object.

[0115] Exemplarily, the probability values ​​satisfying the predetermined sorting condition include: after sorting the probability values ​​from high to low, selecting at least one probability value in the sorted results in descending order of probability values. For example, selecting the merged image with the maximum predicted probability value as the target image.

[0116] In some embodiments, the feature information includes item attribute information of the target item and store attribute information of the store to which the target item belongs; the step of inputting each merged image and feature information into a pre-trained click-through rate prediction model to predict the click-through rate of each merged image through the click-through rate prediction model may specifically include: using the click-through rate prediction model to perform feature fusion on the input merged images, item attribute information and store attribute information, and performing click-through rate prediction based on the fused features to obtain the click-through rate prediction results of each merged image.

[0117] For example, item attribute information refers to the attribute information of the item itself. Examples include, but are not limited to, at least one of the following: item name, item category, item size, item color, item weight, item origin, and item purpose. Store attribute information is used to identify and differentiate different stores. Examples include, but are not limited to, at least one of the following: store name, store category, store business model, store city name, store address, store rating, and store payment method.

[0118] Exemplarily, the item attribute information and store attribute information in the characteristic information can be pre-obtained from a customer relationship management system. This system can record the interaction history between users and stores, including purchase records and customer feedback information. The characteristic information can be obtained from the purchase records and customer feedback information. Exemplarily, the item attribute information and store attribute information on the e-commerce platform can also be obtained from the application programming interface provided by the e-commerce platform. Exemplarily, a data integration tool or platform can also be used to integrate data from different sources into a unified system to facilitate the management and analysis of the above-mentioned item characteristic information.

[0119] Exemplarily, the item attribute information and store attribute information of each item pre-acquired based on any of the above methods can be stored in a predetermined information database. The information database can be a local database or a cloud database. When it is necessary to obtain the characteristic information of the target item, the item attribute information and store attribute information of the target item of the same type as the predetermined item (the item in the foreground image of the predetermined image) can be searched from the item attribute information and store attribute information of each item contained in the information database. In the information database, the text similarity between the item name of the target item of the same type and the item name of the predetermined item meets the predetermined text similarity threshold. The predetermined text similarity threshold can be customized according to actual needs, and is not specifically limited in the embodiments of the present application.

[0120] In this embodiment, by combining the feature information of each merged image and the object in the image to perform feature fusion, the click-through rate prediction model can more comprehensively capture the factors that affect user behavior, thereby helping to improve the accuracy of the prediction.

[0121] Figure 5 The following is a schematic diagram showing the working principle of the click rate prediction model of the exemplary embodiment of the present application. Figure 5 In the process of the click-through rate prediction model, the processing includes the following steps.

[0122] like Figure 5 As shown in “S501, image encoding”, image encoding is performed on any merged image to obtain an image embedding representation of the merged image.

[0123] like Figure 5 As shown in "S502, text encoding", the acquired item attribute information of the target item and the store attribute information of the store to which the target item belongs are input into the text encoder to obtain the text embedding representation output by the text encoder.

[0124] For example, the item in the foreground image is beef rice and the target item is beef rice bowl. The item attribute information (also called product features) includes but is not limited to Figure 5 Shown in: Food name and food category. Store attribute information (also known as store characteristics) includes but is not limited to Figure 5 Shown in: store name, store category, business model and city name.

[0125] Exemplarily, the text encoder can be a pre-trained language model (A Robustly Optimized BERT Approach, RoBERTa). This pre-trained language model is a pre-trained language model based on bidirectional encoder representations from transformers (BERT), which is primarily used for natural language processing tasks to improve the performance and robustness of the model. When performing click-through rate prediction, this pre-trained language model can serve as a text feature extractor, providing high-quality text features for the click-through rate prediction model, thereby improving the accuracy of click-through rate prediction.

[0126] It should be understood that the text encoder can also be any of the following: a transformer-based bidirectional encoder representation model, a shallow neural network-based word embedding model (Word2Vec), or a matrix decomposition-based word embedding model (GloVe). In actual scenarios, the text encoder can be reasonably selected and used based on actual needs and resource conditions, and this embodiment of the application does not specifically limit it.

[0127] like Figure 5 As shown in "S503, Feature Mining", the click-through rate prediction model can include 4 transformer (Transformer*4) layers. Multi-layer transformers can more deeply explore the complex relationship between click-through rate and item attribute information and store attribute information, thereby outputting feature representations containing rich interactive information. The transformer in the click-through rate prediction model can adopt a multi-head attention mechanism to capture the relationship between features from different angles. Specifically, it can divide the input into multiple heads, each head performs self-attention calculation independently, and then splices the outputs of all heads together, and then integrates them through a linear transformation. Multi-head attention allows the model to learn the relationship between features in different subspaces, thereby capturing the interactive information of features more comprehensively.

[0128] like Figure 5 As shown in "S504, Feature Fusion" in the CTR prediction model, the fully connected layer is typically located after the transformer layer. It further processes the features output by the multiple transformer layers and converts them into CTR prediction results. The CTR prediction results may include, for example, probability values ​​directly related to the CTR prediction.

[0129] Specifically, the feature representation output by the transformer layer is a high-dimensional vector containing rich feature interaction information. The fully connected layer fuses this feature information through a series of linear transformations and nonlinear activation functions. The output of the fully connected layer typically passes through an activation function (sigmoid), mapping the activation function to the interval [0, 1] to obtain the click-through rate prediction value. Through the output of the fully connected layer, the model can convert complex feature interaction information into a specific probability value, thereby obtaining the click-through rate prediction result.

[0130] Through steps S501-S504 above, the click-through rate (CTR) of each merged image can be evaluated using a pre-trained multimodal CTR prediction model. This model can comprehensively analyze the visual and textual features of each merged image to determine which image is most likely to attract a user click, and then select at least one image with the highest CTR (an attractive image) as the target image to be displayed. This model helps increase the likelihood of achieving a high CTR for the target image to be displayed, thereby optimizing the overall effectiveness of the image generation process for deployment in the real world and having high practicality.

[0131] According to the method of the embodiment of the present application, images similar to the predetermined image can be retrieved from the image database, which is beneficial to ensuring the diversity and relevance of the background image. The foreground image of the predetermined image is merged with the background image of the retrieved similar image to generate a plurality of combined images. Then, the target image to be displayed is screened from each combined image through click-through rate prediction. Compared with the image restoration technology in the related art, the method of the embodiment of the present application can generate a plurality of combined images by retrieving similar images and merging the foreground and background, thereby increasing the diversity of the image; and click-through rate prediction is beneficial to directly evaluate the effect of the generated image, ensuring that the final selected image has a high click-through rate, which is beneficial to improving the efficiency and practicality of the image processing method.

[0132] Figure 6 Schematic diagram of the processing flow of the image display method of the embodiment of the present application. Figure 6 As shown, the method may include steps S601 to S605.

[0133] S601: Receive a product image of a reserved product uploaded by a merchant client of a reserved product service platform.

[0134] S602: Retrieve images similar to the product image from a predetermined image database.

[0135] S603: Merge the foreground image of the product image with the background images of the similar images to obtain merged images.

[0136] S604 , predicting the click rate of each merged image to obtain a click rate prediction result corresponding to each merged image. The click rate prediction result is used to represent the probability of the merged image being clicked by the user after the merged image is displayed.

[0137] S605 : Based on the click rate prediction results of each merged image, a main image of the product to be displayed is selected from each merged image.

[0138] S606: Display the main image of the product on the target page of the product reservation service platform.

[0139] According to the image display method of the embodiment of the present application, a plurality of merged images with different backgrounds are generated by merging the foreground image of the product image with the background image of a similar image retrieved from the image database. The click-through rate prediction model is used to predict the click-through rate of each merged image, thereby accurately evaluating the impact of different merged images on user click behavior, and providing a scientific basis for the subsequent selection of the main product image. Selecting the main product image to be displayed from each merged image based on the click-through rate prediction results is conducive to ensuring that the main product image displayed on the target page has a higher probability of being clicked by users, which is conducive to improving the efficiency of selecting the main product image and increasing the accuracy of the click-through rate of the displayed main product image.

[0140] Illustratively, the product reservation service platform includes but is not limited to at least one of the following application systems: a food delivery platform, an advertising information management platform, an e-commerce application platform, a food delivery application platform, etc.

[0141] For example, the processing flow of the reserved product in S602-S605 can refer to the specific processing process of the image processing method for the reserved image in S201-S204 in the above embodiment, which will not be repeated here.

[0142] In some embodiments, after step S606 , the method may further include: pushing the target page containing the main image of the product to the user terminal corresponding to the target user group.

[0143] Exemplarily, the target user group includes: at least some users in the commodity reservation service platform.

[0144] In this embodiment, by pushing the target page containing the main image of the product to the user terminal corresponding to the target user group, the product information can be accurately delivered to potential interested users, avoiding invalid information push.

[0145] In some embodiments, to more accurately evaluate the click-through rate of different main images, after step S606, the following step may be performed: optimizing the product main image to obtain an optimized product main image. The target user group is divided into two groups, one group displaying the product main image and the other group displaying the optimized product main image. By comparing the click-through rate data of the two groups of users, the click-through rates of the product main image before optimization (the original main image) and the optimized product main image are compared. Based on the comparison results, the product main image with the higher click-through rate is selected and displayed.

[0146] As an example, the optimization process includes at least one of the following: cropping and resizing to adapt to the size of the target page, color adjustment to obtain an image that better meets visual requirements, image compression to speed up loading, sharpening and blurring to highlight the foreground image. If the difference between the click-through rate of the optimized product main image and the click-through rate of the original main image is greater than or equal to the preset difference threshold, then the optimized product main image can be displayed on the target page of the predetermined product service platform. If the difference between the click-through rate of the optimized product main image and the click-through rate of the original main image is less than the preset difference threshold, then the product main image (original main image) can be kept displayed on the target page of the predetermined product service platform.

[0147] The image display method according to the embodiment of the present application is conducive to ensuring that the main image of the product displayed on the target page has a higher probability of being clicked by users, which is conducive to improving the efficiency of selecting the main image of the product and improving the accuracy of the click-through rate of the displayed main image of the product.

[0148] Figure 7 The schematic diagram of the architecture of the image processing system of the exemplary embodiment of the present application is shown. Figure 7 In the example, the image processing system includes: an input module 710, a retriever 720, a generator 730 and a prediction model 740.

[0149] The input module 710 is configured to receive an original image of a predetermined item, item attribute information of the target item, and store attribute information of the store to which the target item belongs. The original image of the predetermined item may come from a merchant client.

[0150] Specifically, the predetermined item can be any commodity. The target item is an item of the same type as the predetermined item.

[0151] The retriever 720 is configured to retrieve similar images of a predetermined original image from a predetermined image database.

[0152] Specifically, if Figure 7As shown, the retriever 720 receives an original image. The original image may be an original image of a predetermined product uploaded by a merchant client. A predetermined number of target regenerated background images are retrieved from the regenerated background images corresponding to each sample image contained in a predetermined image database as similar images to the original image.

[0153] Among them, the similarity between the image features of each target regenerated background image and the first image features of the original image of the predetermined product meets the predetermined similarity condition; for any regenerated background image stored in the image database, the sample image corresponding to the regenerated background image contains the same foreground image and a different background image.

[0154] This retrieval process can use a pre-trained image encoder to vectorize the image, so as to perform similarity retrieval based on the image vector, ensuring that the selected background is consistent with the style and theme of the food body, and ensuring that the retrieved target regenerated background image has a high-quality background image. Taking the type of the predetermined product as food as an example, the image database in this example can be a food image database. Figure 7 In the example, the retriever 720 outputs multiple similar images, for example, the most similar k images, where k is an integer greater than or equal to 1.

[0155] The generator 730 is configured to replace the background of the original image according to the retrieved similar images.

[0156] Specifically, the generator 730 can be used to perform retrieval enhancement generation. The implementation process is as follows: an image diffusion model is used in combination with a neural network structure based on stable diffusion to perform image background replacement generation. Specifically, with the background image of each retrieved similar image as a control condition, the foreground image in the original image is merged with the background image in each similar image to obtain each merged image as each merged image with a new background image. Specifically, in the image generation process, mask information of any merged image can be generated by feature matching. The mask information is used for the image connection position in the merged image. The image diffusion model can use the merged image and the mask information to perform denoising processing at the connection position. The denoised image is conducive to ensuring the integrity of the main pixels of the object in the foreground image and the visual consistency of the background.

[0157] The prediction model 740 is used to predict the click-through rate of each merged image and obtain a click-through rate prediction result corresponding to each merged image.

[0158] Specifically, the prediction model 740 can receive each merged image obtained by background replacement, the item attribute information of the target item, and the store attribute information of the store to which the target item belongs, perform feature fusion on the received information, and perform click-through rate prediction based on the fused features to obtain the click-through rate prediction results of each merged image. The click-through rate prediction results can be click-through rate prediction values. Figure 7 In FIG, the predicted click-through rates of the merged images are shown to be 0.12, 0.10, and 0.09, respectively. The merged image corresponding to the largest predicted click-through rate value of 0.12 is used as the target image to be displayed.

[0159] In an embodiment of the present application, a framework based on retrieval-augmented generation (RAG) is provided. This is a natural language processing framework that combines information retrieval and text generation. The core idea is that before generating text, relevant information is retrieved from a large external knowledge base (such as an image database), and then this information is provided as context to the generation model to generate more accurate, richer, and more knowledgeable images. The framework includes a retriever and a generator. The retriever is responsible for retrieving similar images of the input image from an external image database. This usually uses some information retrieval techniques, such as search based on vector similarity. The generator can receive an input query and retrieved documents as input, and then generate a target text (such as a preset text prompt). This usually uses some pre-trained language models, such as a sequence-to-sequence (Bidirectional and Auto-regressive Transformers, BART) model or a transformer-based sequence-to-sequence model (Text-to-Text Transfer Transformer, T5). According to the natural language processing framework that combines information retrieval and text generation in an embodiment of the present application, by searching an image database, the accuracy of similar images retrieved to a predetermined image can be improved, reducing generation errors and model hallucinations. Based on each retrieved similar image, a composite image containing a richer background image can be generated, which helps increase the information content of each composite image and improve the image quality of each composite image. In some scenarios, the processing framework allows users to view the retrieved similar images to understand the basis for generating the image.

[0160] According to the image processing system of the embodiment of the present application, an image synthesis framework (Framework for Retrieval-Augmented Diffusion Model, FoRAGe) is provided for generating synthetic images with high click-through rates. In this framework, retrieval technology can be used to find background images similar to predetermined images (such as product images of predetermined products) from an image database, and then the product foreground is fused with these background images through a diffusion model to generate multiple synthetic images (i.e., each merged image). The click-through rates of the multiple synthetic images are predicted, and then the target image to be displayed is selected from each synthetic image based on the click-through rate prediction results of each synthetic image. Based on this framework, it is beneficial to generate high-quality synthetic images to improve the click-through rate of the target image to be displayed.

[0161] In view of the fact that the product images generated in the related art are insufficient in visual effects and thus cannot ensure a high click-through rate of the product images, an image processing method is provided to improve the visual appeal of the image and help improve the user click-through rate of the generated image in actual applications.

[0162] After the predetermined image is processed according to the image processing method of the embodiment of the present application and the target image to be displayed is obtained, the target image can be displayed on the target page of the predetermined platform. Online experiments in actual applications have shown that the click-through rate of the target image among the user group has been greatly improved. In the case where the predetermined image is a product image of a predetermined product, the predetermined platform is a predetermined product service platform. The target page may include but is not limited to at least one of the following: a product display page (such as a product list page, a product details page, etc.), a product search result page, and an order confirmation page. The target image includes the main image of the product to be displayed. The successful application of this image display method on the predetermined product service platform demonstrates the practical value of the embodiment of the present invention in actual application scenarios.

[0163] The present application also provides an image processing method. In some embodiments, the method can be applied to a merchant client. Figure 8 FIG. 1 shows a flow chart of an image processing method according to another embodiment of the present application. Figure 8 As shown, the method may include steps S801 to S802.

[0164] S801, sending the reservation image to the server of the reservation product service platform.

[0165] S802: Receive a target image sent by the server. The target image is an image obtained by the server performing the image processing method of the above embodiment on a predetermined image.

[0166] For example, the predetermined image may be a product image of a predetermined product, and the target image may be a main product image obtained by the server executing the image display method of the above embodiment on the product image of the predetermined product.

[0167] According to the method of the embodiment of the present application, the merchant client can send a predetermined image to the server of the predetermined product service platform. The server of the predetermined product service platform can retrieve images similar to the predetermined image from the image database, and merge the foreground image of the predetermined image with the background image of the retrieved similar images to generate multiple combinations of merged images, and then filter out the target image to be displayed from each merged image through click-through rate prediction. Through this method, after the merchant client uploads the predetermined image to the server, subsequent image retrieval, merging, screening and other operations are automatically completed by the server, which greatly simplifies the processing flow and helps reduce the merchant's manpower and time costs in image processing.

[0168] The image processing method performed on the predetermined image in the embodiment of the present application can refer to the corresponding description of the image processing method applied to the server described in the above embodiment, and has corresponding beneficial effects, which will not be repeated here.

[0169] The present application also provides an image display method. In some embodiments, the method can be applied to a user client. Figure 9 A flow chart of an image display method according to another embodiment of the present application is shown. Figure 9 As shown, the method may include step S901.

[0170] S901, displaying a target page. The display content of the target page includes a target image. The target image is an image obtained by the server according to the image processing method of the above embodiment.

[0171] For example, the target image can be the main product image to be displayed, obtained by the server executing the image display method of the above embodiment on the product image of the reserved product. For example, the user client can display the user interface of the reserved product service platform, which includes an interactive entry element for the target page. In response to an operation instruction on the interactive entry element, the target page is displayed.

[0172] According to the image display method of the embodiment of the present application, the user client can display the target page through the user operation interface of the reserved product service platform and obtain the target image displayed on the target page. The target image is obtained by the server according to the image processing method of the above embodiment, so the image has diversity and a high click-through rate.

[0173] The image processing method performed on the predetermined image in the embodiment of the present application can refer to the corresponding description of the image processing method applied to the server described in the above embodiment, and has corresponding beneficial effects, which will not be repeated here.

[0174] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides an image processing device. Figure 10 A schematic diagram of the structure of an image processing device according to an embodiment of the present application is shown, which can be used to execute the image processing method applied to the server provided in any of the above embodiments, such as Figure 10 As shown, the image processing device includes:

[0175] The retrieval module 1010 is configured to retrieve similar images of a predetermined image from a predetermined image database.

[0176] The merging module 1020 is configured to merge the foreground image in the predetermined image with the background images in each of the similar images to obtain merged images.

[0177] The prediction module 1030 is used to predict the click-through rate of each merged image and obtain a click-through rate prediction result corresponding to each merged image; the click-through rate prediction result is used to represent the probability of the merged image being clicked by the user after the merged image is displayed.

[0178] The selection module 1040 is configured to select a target image to be displayed from each merged image based on the click rate prediction results of each merged image.

[0179] In some embodiments, the retrieval module 1010 is specifically used to: encode a predetermined image to obtain a first image feature; retrieve at least one target regenerated background image from the regenerated background images corresponding to each sample image contained in the image database, and the similarity between the image feature of the target regenerated background image and the first image feature meets a predetermined similarity condition; any regenerated background image and the corresponding sample image contain the same foreground image and a different background image; and use each target regenerated background image as a similar image of the predetermined image.

[0180] In some embodiments, the image processing device also includes: a storage module, which is used to encode a pre-acquired sample image before retrieving at least one target regenerated background image from the regenerated background images corresponding to each sample image contained in the image database to obtain image features of the sample image; generate at least one regenerated background image corresponding to the sample image based on the sample image and a preset text prompt, the preset text prompt is used to describe the sample image through text; and store the sample image and the corresponding regenerated background images in a predetermined image database.

[0181] In some embodiments, when the storage module is used to generate at least one regenerated background image corresponding to the sample image based on the sample image and a preset text prompt, it is specifically used to: remove the background of the sample image to obtain a first foreground image; use a preset image generation model to generate at least one new background for the first foreground image based on the first foreground image and a preset text prompt; merge the first foreground image with each new background respectively to obtain each regenerated background image corresponding to the sample image.

[0182] In some embodiments, the merging module 1020 is specifically used to: extract a foreground image from a predetermined image to obtain a second foreground image; extract a background image from each similar image to obtain each background image; perform feature matching on the second foreground image with each background image to determine the matching position of the second foreground image in each background image; and perform image fusion on the second foreground image and each background image based on the matching position of the second foreground image in each background image to obtain each merged image.

[0183] In some embodiments, the image processing device also includes: an updating module, which is used to generate mask information of the merged image for any merged image after obtaining each merged image, and the mask information is used to mark the connection position between the second foreground image and the background image in the merged image; input the merged image and the mask information into a pre-trained image diffusion model, and the image diffusion model uses the merged image and the mask information to perform denoising processing at the connection position; and use the denoised image to update the merged image.

[0184] In some embodiments, the prediction module 1030 is specifically used to: obtain feature information of the target item, where the target item and the item in the foreground image are of the same type; input each merged image and feature information into a pre-trained click-through rate prediction model to predict the click-through rate of each merged image through the click-through rate prediction model to obtain a click-through rate prediction result for each merged image.

[0185] In some embodiments, the feature information includes item attribute information of the target item and store attribute information of the store to which the target item belongs; when the prediction module 1030 is used to input each merged image and feature information into a pre-trained click-through rate prediction model to predict the click-through rate of each merged image through the click-through rate prediction model, it is specifically used to: use the click-through rate prediction model to perform feature fusion on the input merged images, item attribute information and store attribute information, and perform click-through rate prediction based on the fused features to obtain the click-through rate prediction results of each merged image.

[0186] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0187] Corresponding to the application scenario and method of the method provided in the embodiment of the present application, the embodiment of the present application also provides an image display device. Figure 11 A schematic diagram of the structure of an image display device according to an embodiment of the present application is shown, and the device can be used to execute the image display method applied to the server provided in any of the above embodiments, such as Figure 11 As shown, the image display device includes:

[0188] The receiving module 1110 is configured to receive product images of reserved products uploaded by a merchant client of the reserved product service platform.

[0189] The retrieval module 1120 is configured to retrieve similar images of the product image from a predetermined image database.

[0190] The merging module 1130 is configured to merge the foreground image of the product image with the background images in each of the similar images to obtain merged images.

[0191] The prediction module 1140 is used to predict the click-through rate of each merged image and obtain a click-through rate prediction result corresponding to each merged image. The click-through rate prediction result is used to represent the probability of the merged image being clicked by the user after the merged image is displayed.

[0192] The selection module 1150 is configured to select a main product image to be displayed from each merged image based on the click rate prediction results of each merged image.

[0193] The display module 1160 is used to display the main image of the product on the target page of the product reservation service platform.

[0194] In some embodiments, the image display device further includes: a sending module for pushing the target page containing the main image of the product to the user terminal corresponding to the target user group after displaying the main image of the product on the target page of the reserved product service platform.

[0195] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0196] An embodiment of the present application also provides an image processing device. Figure 12 FIG. 1 shows a schematic diagram of the structure of an image processing device according to an embodiment of the present application. The device can be used to execute the image processing method applied to a merchant client provided in any of the above embodiments, such as Figure 12 As shown, the image processing device includes:

[0197] The sending module 1210 is used to send the reservation image to the server of the reservation product service platform;

[0198] The receiving module 1220 is configured to receive a target image sent by the server. The target image is an image obtained by the server performing an image processing method on a predetermined image.

[0199] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0200] An embodiment of the present application also provides an image display device. Figure 13 FIG. 1 shows a schematic diagram of the structure of an image display device according to an embodiment of the present application. The device can be used to execute the image display method applied to a user client provided in any of the above embodiments, such as Figure 13 As shown, the image display device includes:

[0201] The display module 1310 is used to display the target page. The display content of the target page includes a target image. The target image is an image obtained by the server according to the above-mentioned image processing method applied to the server.

[0202] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0203] Figure 14 FIG. 1 is a block diagram of an electronic device for implementing an embodiment of the present application. Figure 14 As shown, the electronic device includes: a memory 1401 and a processor 1402. The memory 1401 stores a computer program that can be executed on the processor 1402. When the processor 1402 executes the computer program, the method of the above embodiment is implemented. The number of memory 1401 and processor 1402 can be one or more. In a specific implementation, the electronic device may also include a communication interface 1403 for communicating with external devices and exchanging data.

[0204] In a specific implementation, if the memory 1401, the processor 1402, and the communication interface 1403 are implemented independently, the memory 1401, the processor 1402, and the communication interface 1403 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 14Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0205] Optionally, in a specific implementation, if the memory 1401 , the processor 1402 , and the communication interface 1403 are integrated on a chip, the memory 1401 , the processor 1402 , and the communication interface 1403 may communicate with each other through an internal interface.

[0206] An embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the image processing method or image display method provided in the embodiment of the present application.

[0207] An embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the image processing method or image display method provided in the embodiment of the present application.

[0208] An embodiment of the present application also provides a chip, which includes a processor for calling and executing instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the image processing method or image display method provided in the embodiment of the present application.

[0209] An embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the image processing method or image display method provided in the embodiment of the application.

[0210] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0211] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).

[0212] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0213] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0214] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0215] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.

[0216] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.

[0217] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0218] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0219] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An image processing method, characterized in that: include: Retrieving each similar image to the predetermined image from a predetermined image database; Merging the foreground image in the predetermined image with the background images in each of the similar images to obtain merged images; Predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image; the click-through rate prediction result is used to represent the probability of the merged image being clicked by the user after the merged image is displayed; Based on the click rate prediction results of the merged images, a target image to be displayed is selected from the merged images.

2. The method according to claim 1, characterized in that The retrieving similar images of the predetermined image from the predetermined image database includes: Encoding the predetermined image to obtain a first image feature; Retrieving at least one target regenerated background image from the regenerated background images corresponding to the sample images contained in the image database, wherein the similarity between the image features of the target regenerated background image and the first image features satisfies a predetermined similarity condition; and any of the regenerated background images and the corresponding sample image include the same foreground image and a different background image; The background images of the respective objects are reproduced as similar images to the predetermined image.

3. The method according to claim 2, characterized in that Before retrieving at least one target regenerated background image from the regenerated background images corresponding to the sample images contained in the image database, the method further includes: Encoding a pre-acquired sample image to obtain image features of the sample image; generating at least one regenerated background image corresponding to the sample image based on the sample image and a preset text prompt, wherein the preset text prompt is used to describe the sample image in words; The sample images and the corresponding regenerated background images are stored in the predetermined image database.

4. The method according to claim 3, characterized in that The step of generating at least one regenerated background image corresponding to the sample image based on the sample image and the preset text prompt includes: Performing background removal on the sample image to obtain a first foreground image; generating at least one new background for the first foreground image based on the first foreground image and the preset text prompt using a preset image generation model; The first foreground image is merged with each new background respectively to obtain each regenerated background image corresponding to the sample image.

5. An image display method, characterized in that: The method comprises: Receiving product images of the reserved products uploaded by the merchant client of the reserved product service platform; Retrieving each similar image of the product image from a predetermined image database; Merging the foreground image of the product image with the background images of the similar images to obtain merged images; Predicting the click-through rate of each merged image to obtain a click-through rate prediction result corresponding to each merged image, wherein the click-through rate prediction result is used to represent the probability of the merged image being clicked by a user after the merged image is displayed; Selecting a main product image to be displayed from the merged images based on the click-through rate prediction results of the merged images; The main image of the product is displayed on the target page of the reserved product service platform.

6. An image processing method, characterized in that: The method comprises: Sending the reservation image to the service end of the reservation product service platform; Receive a target image sent by the server, where the target image is an image obtained by the server executing the method according to any one of claims 1 to 4 on the predetermined image.

7. An image display method, characterized in that: The method comprises: The target page is displayed, where the displayed content of the target page includes a target image, and the target image is an image obtained by the server according to the method according to any one of claims 1 to 4.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor implements the method according to any one of claims 1 to 4, claim 5, claim 6, or claim 7 when executing the computer program.

9. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the method of any one of claims 1 to 4, claim 5, claim 6, or claim 7.

10. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4, claim 5, claim 6 or claim 7.