Image processing method and device, storage medium, equipment and program product

By using splitting and splicing image processing technology combined with a style transfer model, high-quality target images are generated, which solves the problem of insufficient style fusion and improves the effect and quality of subject-background synthesis.

CN120852157APending Publication Date: 2025-10-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202410526457.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Current technology suffers from insufficient style fusion during the style transfer process of subject-background synthesis, resulting in unnatural generated images. This is especially true in advertising, where the product and background are not sufficiently integrated, resulting in a poor visual experience.

Method used

By obtaining the prompt word, original image and background image, splitting processing is performed to obtain the target subject area map, mask map and edge map, and then splicing processing is performed. Combined with the trained style transfer model, style transfer processing is performed to generate the target image, ensuring a high degree of fusion of the target subject and target background styles.

Benefits of technology

A high degree of integration between the target subject and the target background style is achieved, and the generated image has higher quality, meets the specific style and theme requirements, and improves the effectiveness of advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852157A_ABST
    Figure CN120852157A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device, a storage medium, equipment and a program product, which are applied to a style migration scene. The method comprises the following steps: obtaining a cue word, an original image containing a target subject and a background image with a target background style, wherein the cue word is used for prompting to generate an image containing the target subject and with the target background style; splitting the original image to obtain a target main body area image, a mask image and an edge image; carrying out splicing processing on the target main body area graph, the mask graph and the edge graph to obtain a spliced image; according to the spliced image, the cue word and the background image, style migration processing is carried out, a target image matched with the cue word is obtained, and the target image comprises a target body and has a target background style. Through double guidance of the cue word and the background image, the original image is subjected to style migration to generate the high-quality target image, and the target main body in the target image and the target background style have a high fusion degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing technology, specifically to an image processing method, apparatus, storage medium, device, and program product. Background Art

[0002] Generative Artificial Intelligence Generated Content (AIGC) plays a crucial role in subject-background synthesis, primarily relying on the concept of local mapping and utilizing diffusion models for image generation and style transfer. The diffusion model generates images by training a parameterized Markov chain that progressively denoises the data. This process includes a diffusion process, where noise is added to corrupt the data, while a de-diffusion process generates samples through denoising.

[0003] In terms of style transfer, current techniques typically modify the inference strategy, changing the starting point for model denoising from random noise to the noisy state of the target image, and then performing denoising from that starting point to control the subsequent generation process. This strategy, known as "pad image," enables the style of the resulting image to resemble the background image, thus achieving style transfer.

[0004] However, current technologies suffer from insufficient style fusion during the style transfer process of subject-background synthesis. Summary of the Invention

[0005] This application provides an image processing method, apparatus, storage medium, device, and program product. Through dual guidance of prompts and background images, style transfer is performed on the original image to generate a high-quality target image. The target subject in the target image has a high degree of fusion with the target background style, which significantly improves the style transfer effect.

[0006] On one hand, embodiments of this application provide an image processing method, the method comprising:

[0007] The process involves obtaining a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style. The original image is then split to obtain a target subject region image, a mask image, and an edge image. The mask image is used to cover the target subject region image, and the edge image is used to depict the edge contours of the target subject region image. The target subject region image, the mask image, and the edge image are then stitched together to obtain a stitched image. Style transfer processing is then performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

[0008] On the other hand, embodiments of this application provide an image processing apparatus, the apparatus comprising:

[0009] The acquisition unit is used to acquire a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style.

[0010] A splitting unit is used to split the original image to obtain a target subject region map, a mask map, and an edge map. The mask map is used to cover the target subject region map, and the edge map is used to depict the edge contour of the target subject region map.

[0011] The stitching unit is used to stitch the target subject area map, the mask map, and the edge map together to obtain a stitched image;

[0012] The processing unit is configured to perform style transfer processing on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word, wherein the target image contains the target subject and has the target background style.

[0013] In some embodiments, the processing unit is specifically used to: input the stitched image, the prompt word, and the background image into a trained style transfer model for style transfer processing to obtain a target image; wherein, the trained style transfer model is trained based on multiple sample prompt words and sample images matching each sample prompt word.

[0014] In some embodiments, the trained style transfer model includes a local drawing module, an image cue adapter module, and a diffusion module; the processing unit includes: a first processing subunit, used to input the stitched image and the cue word into the local drawing module for embedding processing to obtain an image embedding vector corresponding to the stitched image and a text embedding vector corresponding to the cue word; a second processing subunit, used to input the cue word and the background image into the image cue adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image; and a third processing subunit, used to input the image embedding vector, the text embedding vector, the first cross-attention score, and the second cross-attention score into the diffusion module for style transfer processing to obtain a target image.

[0015] In some embodiments, the first processing subunit is specifically used to: obtain a first weight coefficient corresponding to the local drawing module, wherein the first weight coefficient is used to control the degree of fusion between the edge contour of the target subject in the original image and the spatial relationship in the background image; input the first weight coefficient, the stitched image and the prompt word into the local drawing module for embedding processing to obtain the updated image embedding vector corresponding to the stitched image and the updated text embedding vector corresponding to the prompt word.

[0016] In some embodiments, the second processing subunit is specifically used to: obtain a second weight coefficient corresponding to the image cue adapter module, the second weight coefficient being used to characterize the similarity between the target image and the background image; input the second weight coefficient, the cue word, and the background image into the image cue adapter module for attention decoupling processing, to obtain an updated first cross-attention score corresponding to the cue word and an updated second cross-attention score corresponding to the background image.

[0017] In some embodiments, the image prompt adapter module includes a text encoding module, an image encoding module, and a decoupled cross-attention adaptation module; the second processing subunit is specifically used for:

[0018] The prompt words are input into the text encoding module for encoding processing to obtain text features;

[0019] The background image is input into the image encoding module for encoding processing to obtain image features;

[0020] The text features and the image features are input into the decoupled cross-attention adaptation module for attention decoupling processing to obtain the first cross-attention score corresponding to the prompt word and the second cross-attention score corresponding to the background image.

[0021] In some embodiments, the image processing apparatus further includes a training unit, configured to: acquire a plurality of sample cue words and sample images matching each of the sample cue words; train the local drawing module and the diffusion module in the style transfer model based on a first sample cue word among the plurality of sample cue words and a first sample image matching the first sample cue word, and determine a first loss function; train the image cue adapter module and the diffusion module in the style transfer model based on the first loss function, a second sample cue word among the plurality of sample cue words, a second sample image matching the second sample cue word, and a third sample image matching a third sample cue word among the plurality of sample cue words, and determine a second loss function, wherein the second sample image and the third sample image correspond to the same sample subject but have different background styles; update the parameters of the style transfer model based on the first loss function and the second loss function to obtain the trained style transfer model.

[0022] In some embodiments, when the training unit trains the local drawing module and the diffusion module in the style transfer model based on the first sample prompt word among the plurality of sample prompt words and the first sample image matching the first sample prompt word, and determines the first loss function, it specifically performs the following steps: splitting the first sample image to obtain a first sample main body region image, a first sample mask image, and a first sample edge image, wherein the first sample mask image is used to cover the first sample main body region image, and the first sample edge image is used to depict the edge contour of the first sample main body region image; stitching the first sample main body region image, the first sample mask image, and the first sample edge image together to obtain a first sample stitched image; inputting the first sample stitched image and the first sample prompt word into the local drawing module for embedding to obtain a first sample image embedding vector corresponding to the first sample stitched image and a first sample text embedding vector corresponding to the first sample prompt word; inputting the first sample prompt word, the first sample image, the first sample image embedding vector, and the first sample text embedding vector into the diffusion module for style transfer to obtain a first generated image; and determining the first loss function based on the mean squared error loss function between the first generated image and the first sample image.

[0023] In some embodiments, when the training unit trains the image cue adapter module and the diffusion module in the style transfer model based on the first loss function, the second sample cue word among the plurality of sample cue words, the second sample image matching the second sample cue word, and the third sample image matching the third sample cue word among the plurality of sample cue words, and determines the second loss function, it is specifically used for: updating the parameters of the style transfer model according to the first loss function to obtain an updated style transfer model; inputting the second sample cue word and the third sample image into the image cue adapter module in the updated style transfer model for attention decoupling processing to obtain a first cross-attention score corresponding to the second sample cue word and a second cross-attention score corresponding to the third sample image; inputting the second sample cue word, the third sample image, the first cross-attention score corresponding to the second sample cue word, and the second cross-attention score corresponding to the third sample image into the diffusion module in the updated style transfer model for style transfer processing to obtain a second generated image; and determining the second loss function based on the mean squared error loss function between the second generated image and the second sample image.

[0024] In some embodiments, the splitting unit is specifically used for: performing image matting on the original image to obtain the mask image; performing composite processing on the mask image and the original image to obtain the target subject region image; and performing edge detection on the mask image to obtain the edge image.

[0025] On the other hand, an embodiment of this application provides a computer-readable storage medium storing a computer program adapted for loading by a processor to perform the image processing method as described in any of the above embodiments.

[0026] On the other hand, an embodiment of this application provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor executes the image processing method described in any of the above embodiments by calling the computer program stored in the memory.

[0027] On the other hand, an embodiment of this application provides a computer program product, including computer instructions that, when executed by a processor, implement the image processing method as described in any of the above embodiments.

[0028] This embodiment of the application obtains a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to guide the generation of an image containing the target subject and having the target background style. The original image is split to obtain a target subject region image, a mask image, and an edge image. The mask image is used to cover the target subject region image, and the edge image is used to depict the edge contour of the target subject region image. The target subject region image, mask image, and edge image are stitched together to obtain a stitched image. Style transfer processing is performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style. This embodiment of the application achieves precise control over the subject-background synthesis by comprehensively utilizing the prompt word, the original image, and the background image with the target background style. The prompt word guides the generation process, ensuring that the generated image meets specific style and theme requirements. The original image provides the target subject to be synthesized, while the background image provides the target background style. The combination of the two ensures that the generated image retains the characteristics of the target subject while incorporating the target background style. By splitting the original image, a target subject region map, a mask map, and an edge map are obtained. These parts are then stitched together, allowing for more precise control over the position and shape of the target subject during style transfer. The mask map covers the target subject region map, ensuring the integrity of the subject during compositing. The edge map depicts the edge contours of the target subject region map, enhancing the blending between the subject and the background. Furthermore, by combining the stitched image, the cue word, and the background image for style transfer processing, a target image matching the cue word can be generated. This processing method not only results in a higher quality image but also ensures a higher degree of blending between the target subject and the target background style. This embodiment achieves precise control over the style transfer process through the dual guidance of cue words and background images. This processing method overcomes the problem of insufficient style transfer blending in current technologies, improving the effect and quality of subject-background compositing. Attached Figure Description

[0029] Figure 1 This is a schematic diagram of the first application scenario of a diffusion model.

[0030] Figure 2 This is a schematic diagram of a second application scenario for a diffusion model.

[0031] Figure 3 This is a schematic diagram illustrating an application scenario of the image processing system provided in an embodiment of this application.

[0032] Figure 4 This is a schematic flowchart of the image processing method provided in an embodiment of this application.

[0033] Figure 5This is a schematic diagram of a first application scenario of the image processing method provided in the embodiments of this application.

[0034] Figure 6 This is a schematic diagram of a second application scenario of the image processing method provided in the embodiments of this application.

[0035] Figure 7 This is a schematic diagram of a third application scenario of the image processing method provided in the embodiments of this application.

[0036] Figure 8 This is another schematic flowchart of the image processing method provided in the embodiments of this application.

[0037] Figure 9 This is a schematic diagram of a fourth application scenario of the image processing method provided in the embodiments of this application.

[0038] Figure 10 This is a schematic diagram of the fifth application scenario of the image processing method provided in the embodiments of this application.

[0039] Figure 11 This is a schematic diagram of the structure of the image processing apparatus provided in the embodiments of this application.

[0040] Figure 12 A schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0041] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0042] This application provides an image processing method, apparatus, storage medium, device, and program product. Exemplarily, the image processing method of this application can be executed by a computer device, which can be a terminal or a server, etc. The terminal can be a smartphone, tablet, laptop, desktop computer, smart TV, smart speaker, wearable smart device, personal computer (PC), smart vehicle terminal, etc. The terminal may also include a client, which can be a video client, shopping application client, reading application client, browser client, or instant messaging client, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms.

[0043] The embodiments of this application can be applied to scenarios such as artificial intelligence, machine learning, image processing, subject-background synthesis, and style transfer.

[0044] First, some of the nouns or terms that appear in the description of the embodiments of this application are explained as follows:

[0045] Artificial Intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to have perception, reasoning, and decision-making capabilities. AI technology is a comprehensive discipline involving a wide range of fields, encompassing both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can be widely applied to downstream tasks in various areas of AI after fine-tuning. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0046] Computer Vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and further processes images to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision technology. Pre-trained models in the vision field, such as Swin-transformer, ViT, V-MOE, and MAE, can be quickly and widely applied to downstream tasks after fine-tuning. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0047] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.

[0048] Image generation: This involves using prompt words as input to an AI model to generate corresponding images, falling under the category of generative artificial intelligence (AIGC).

[0049] AIGC Subject Background Synthesis: In text-based image applications, one type of application involves changing the background of the subject image. Users provide the subject image and a prompt for the desired background, which is then input into the AIGC background synthesis model. The model can then relatively naturally replace the subject with the appropriate background.

[0050] Current AIGC (Artificial Intelligence Generative Design) technology plays a crucial role in subject-background synthesis, primarily relying on the idea of ​​inpainting and utilizing a diffusion model for image generation and style transfer. Specifically, it employs a specific sampling method, such as using a denoising diffusion probabilistic model (DDPM), to train the model, enabling it to accurately draw the desired image under the guidance of prompts. DDPM, as an advanced diffusion model for image generation, is trained using a parameterized Markov chain to progressively denoise the data.

[0051] like Figure 1 As shown, the amplification model mainly covers two core steps: First, the diffusion process, which is to gradually transform the original image (Data) on the left into noise on the right. This transformation process is to destruct the original image by gradually adding noise. Second, the reverse diffusion process, which is to gradually restore the original image (Data) on the left from the noise on the right. This process is to gradually denoise to recover the original image from the Gaussian noise.

[0052] Building upon this, current techniques for style transfer typically modify the inference strategy, changing the starting point for model denoising from random noise to the state of the target image after adding noise at a specific number of steps. Denoising is then performed from this new starting point, thus precisely controlling the subsequent image generation process. This strategy, known as "pad image," effectively guides the style of the generated image to approximate the background image, thereby achieving the effect of style transfer.

[0053] However, current technology still suffers from insufficient blending in style transfer during subject-background compositing. For example, the compositing effect often appears unnatural or exhibits disordered spatial relationships.

[0054] like Figure 2 As shown, when the configuration of adding noise for 800 steps and removing noise for 100 steps is used to synthesize the main image a and the background image b, in the resulting composite image c, the main image a appears to be directly pasted onto the background image b, the positional relationship appears messy and the visual effect is not natural, and the integration of the main image (product) with the background is obviously insufficient.

[0055] For clients whose goal is advertising, this composite effect is clearly unsatisfactory. The product itself doesn't blend well with the background, resulting in a poor visual experience and making it unsuitable as an advertising image.

[0056] In view of this, this application proposes an innovative solution to the problem of insufficient style transfer fusion in current AIGC subject-background synthesis. The aim is to generate target images with higher subject-background fusion and a superior visual experience, thereby meeting the needs of a wide range of customers and improving the effectiveness of advertising.

[0057] The solutions provided in this application relate to technologies such as image generation using artificial intelligence, and are specifically illustrated through the following embodiments. Detailed descriptions are provided below. It should be noted that the order of description in the following embodiments is not intended to limit the priority of the embodiments.

[0058] Please see Figure 3 , Figure 3 This is a schematic diagram illustrating an application scenario of the image processing system provided in an embodiment of this application. The system can implement an image processing method. The image processing system may include a terminal 1000, a server 2000, and a network 3000, and the terminal 1000 and server 2000 can interact with each other via the network 3000.

[0059] The server 2000 is used to acquire a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style. The original image is split to obtain a target subject region image, a mask image, and an edge image. The mask image is used to cover the target subject region image, and the edge image is used to depict the edge contour of the target subject region image. The target subject region image, mask image, and edge image are stitched together to obtain a stitched image. Style transfer processing is performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

[0060] Terminal 1000 is used to obtain the target image from server 2000 and display the target image.

[0061] This application provides an image processing method, which can be executed by a terminal or a server, or by both a terminal and a server. This application uses the example of an image processing method executed by a server to illustrate the method.

[0062] Please see Figures 4 to 10 , Figure 4 and Figure 8 This is a schematic flowchart of the image processing method provided in the embodiments of this application. Figures 5 to 7 ,as well as Figures 9 and 10 These are all schematic diagrams illustrating application scenarios of the image processing method provided in the embodiments of this application. The method may include the following steps 110 to 140:

[0063] Step 110: Obtain the prompt word, the original image containing the target subject, and the background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and with the target background style.

[0064] In this context, a prompt is a descriptive text, typically entered by the user, that provides key information to guide the image generation process. The prompt specifically specifies the content or style expected to be seen in the generated image. Prompts usually include descriptions of the target subject and background style. For example, such as... Figure 5 As shown, the prompt is "a bottle of shower gel, placed on a table, light green wallpaper." This prompt clearly identifies the target subject as "a bottle of shower gel" and provides a description of the background style: "placed on a table, light green wallpaper." In this way, the style transfer model 10 can understand the specific requirements of the synthesized image, namely the placement of the main object, the environmental background, and its detailed features. This description provides clear guidance for subsequent image processing, ensuring that the generated image meets the user's expectations.

[0065] The original image is the image that contains the target subject. For example, such as Figure 5 As shown, the original image is a clear picture of a bottle of shower gel. This image provides specific features such as the shape, color, and texture of the target subject. This original image can be a high-quality image uploaded by the user or already existing in the system library. Its main purpose is to provide accurate visual information about the target subject so as to maintain the subject's realism and detail integrity in subsequent processing.

[0066] The background image is an image with the target background style. For example, such as... Figure 5 As shown, the background image is "a potted plant on a stand, light green wallpaper." This image demonstrates the specific characteristics of the target background style, including the style of the stand, the color and texture of the wallpaper, etc. The purpose of this background image is to provide the target background style, so that the generated target image not only includes the specified target subject, but also reflects the specific atmosphere of the described environment.

[0067] In step 110, the concept of the target subject (or subject) is extremely broad, encompassing multiple different categories, including but not limited to goods, items, people, buildings, virtual characters, natural landscapes, works of art, text, or symbols. Each of these subjects has unique characteristics and attributes, which can be specially processed in subsequent image processing and style transfer.

[0068] In this context, merchandise serves as the primary element in marketing and visual presentation, such as a bottle of shower gel, a piece of clothing, or a piece of jewelry; these are typically used for product displays on e-commerce websites. Items, on the other hand, can be any concrete object, such as a mobile phone, car, home appliance, furniture, tool, or any household item. These items can be used for user manuals, advertising, or other commercial purposes. The main merchandise and items usually have a distinct shape and color; during style transfer, it's crucial to ensure that these characteristics are preserved while seamlessly integrating with the new background style.

[0069] The subjects can include portraits, celebrities, artists, models, etc., suitable for magazine covers, social media content, story illustrations, etc. The subjects typically involve complex facial features, body postures, and clothing details. When transferring style, special care must be taken to maintain the distinctive features of the subject while considering the harmony between the subject and the background.

[0070] The buildings included can be historical buildings, skyscrapers, residential areas, traditional dwellings, etc., and are suitable for real estate promotion, urban planning displays, tourism promotion, etc. The main body of the building has a unique structure and style. When integrating it with a new background, the perspective and lighting effects of the building need to be considered to ensure the realism of the composite image.

[0071] Virtual characters can encompass characters or objects from novels, comics, animations, movies, and games, and are suitable for creating promotional materials, game scene design, illustration, and more. These virtual characters possess distinct personalities and characteristics, and it is necessary to maintain their original stylistic features during style transfer while ensuring harmony with the new background style.

[0072] Natural landscapes can include mountains, waterfalls, lakes, forests, and other natural scenery, suitable for tourism brochures, environmental education materials, geography education, and other fields. Natural landscapes have vast spaces and rich details. When transferring styles, it is necessary to consider the layering and changes in light and shadow of natural landscapes in order to create realistic and artistic composite images.

[0073] The artworks included can be paintings, sculptures, photographs, etc., and are suitable for art exhibition previews, catalogue publications, etc. Each artwork possesses a unique artistic style and expressive power. During style transfer, it is necessary to respect and preserve the original style of these artworks while integrating them with the new context to create new artistic effects.

[0074] When text or symbols are used as the target subject, it may involve text design, logos, or patterns, such as trademarks and emblems, which are suitable for brand identification and information dissemination. In style transfer, it is necessary to maintain the clarity and recognizability of these texts or symbols while integrating them with the background style to enhance the overall visual effect.

[0075] Step 120: The original image is split to obtain a target subject region map, a mask map, and an edge map. The mask map is used to cover the target subject region map, and the edge map is used to depict the edge contour of the target subject region map.

[0076] In some embodiments, the original image is split to obtain a target subject region map, a mask map, and an edge map, including: performing image matting on the original image to obtain a mask map; performing composite processing on the mask map and the original image to obtain a target subject region map; and performing edge detection on the mask map to obtain an edge map.

[0077] In step 120, the original image needs to undergo fine-grained segmentation. The purpose of this segmentation is to separate the target subject from the original background and simultaneously extract edge and masking information related to the target subject. This step provides the necessary materials and conditions for subsequent style transfer and image compositing.

[0078] First, the original image undergoes background removal processing. Background removal is an image processing technique designed to precisely separate a specific part of an image (i.e., the target subject) from the background. For example... Figure 5 and Figure 6 As shown, the goal of image matting is to extract the shower gel bottle as the target subject. Image matting can be achieved using various algorithms, such as threshold-based segmentation, edge-based detection, and region-based segmentation. A suitable matting algorithm can be selected based on the characteristics of the original image and the features of the target subject. After matting, a mask image is obtained. This mask image is a binary image where the target subject region is marked as white or black, while the background region is marked as another color (usually black or white). The purpose of the mask image is to cover the target subject region, ensuring that only the target subject is preserved in subsequent compositing processes, while the background is replaced or covered.

[0079] Next, the mask image is composited with the original image. The goal of this compositing process is to separate the extracted subject from the background of the original image, resulting in a region map containing only the subject. During compositing, the mask image can be used as a mask to perform pixel-by-pixel operations on the original image. Specifically, white or black areas in the mask image can be blended or replaced with corresponding pixels in the original image. Through this operation, the background in the original image is removed, leaving only the subject, thus obtaining the subject region map. The subject region map is a color image that fully preserves the color, texture, and detail information of the subject. This image will serve as the base material for subsequent style transfer and image compositing.

[0080] Next, edge detection processing is performed on the mask image to obtain an edge map. Edge detection is an image processing technique used to detect edge information in an image, that is, the boundary line between the target subject and the background. In the edge detection process, edge detection algorithms are used to process the mask image. These algorithms are usually based on features such as image grayscale changes, gradient information, or filter responses to detect edges. For example, edge detection algorithms such as the Canny edge detection operator, the Sobel operator, and the Laplacian operator are used to extract the edge contours of the target subject. After edge detection processing, an edge map is obtained. This edge map is a single-channel image that highlights the edge contours of the target subject, describing the shape and structure of the target subject in the form of lines. The edge map plays an important role in the subsequent subject-background synthesis. It can help the style transfer model 10 to more accurately locate the position of the target subject in the synthesized image (target image) and ensure that the blending between the target subject and the background is more natural and smooth. At the same time, the edge map can also enhance the contrast between the target subject and the background, improving the overall visual effect of the synthesized image (target image).

[0081] For example, the original image can be split into five channels to obtain a target subject region map, a mask map, and an edge map. Here, "five channels" refers to a multi-channel image set containing a target subject region map (red, green, and blue, RGB), a mask map (1 channel), and an edge map (1 channel). This split allows each layer to be processed and modified independently, thus providing finer control over the final image generation.

[0082] like Figure 5 and Figure 6 As shown, the original image undergoes five-channel input splitting to obtain a target subject region map, a mask map, and an edge map. Specifically, the original image is first processed to obtain a 1-channel mask map. This mask map can be used to cover the target subject region map. For example, the black areas of the mask map can reveal the target subject region map "shower gel bottle," while the white areas can cover the original background area outside the target subject region map "shower gel bottle." This mask map is the mask. Then, the mask map and the original image are composited to obtain an RGB 3-channel target subject region map, which is the cutout of the target subject "shower gel bottle." Next, edge detection is performed on the mask map to obtain a 1-channel edge map. The edge map is used to depict the edge contour of the target subject region map; that is, the edge map is the line drawing describing the target subject "shower gel bottle."

[0083] Step 130: The target subject area map, mask map and edge map are stitched together to obtain a stitched image.

[0084] Step 130 involves stitching together the target subject region map, mask map, and edge map to obtain the final stitched image. This step is crucial for ensuring the integrity, accuracy, and visual quality of the image.

[0085] For example, in specific operations, splicing can be performed in different ways:

[0086] Channel-based stitching: This is the most common stitching method, especially when processing multi-channel images. For example, a 3-channel RGB target area image, a 1-channel mask image, and a 1-channel edge image can be stacked according to the channel dimension. If the original target area image is in RGB color space, and the mask image and edge image are single-channel grayscale images, then the grayscale information of the mask image and edge image can be added after the three color channels of the target area image, respectively, to form a new multi-channel image. In this way, the new image will have added transparency (defined by the mask image) and edge details (defined by the edge image) on top of the original color information.

[0087] Spatial stitching: In some cases, it may be necessary to align and overlay different image layers in space. This typically requires image registration techniques to ensure that each layer is perfectly aligned in space. This method is particularly useful when dealing with complex scenes or situations requiring high-precision alignment.

[0088] Hybrid stitching: In some complex scenarios, it may be necessary to stitch together both channel and spatial dimensions simultaneously. This requires more detailed adjustments and blending of each layer to achieve the best visual effect.

[0089] For example, taking channel-level stitching as an example, during the stitching process, the consistency of each image in size and resolution is first ensured so that they can be seamlessly stitched together. Then, the stitching operation is performed at the channel level. This means merging the three RGB channels of the target subject region image, one channel of the mask image, and one channel of the edge image into a single multi-channel image containing all the information. During this stitching process, the mask image accurately indicates the position of the target subject in the image, ensuring that the target subject is not obscured or distorted during stitching. Simultaneously, the edge image enhances the outline of the target subject, making it more prominent and clear in the stitched image.

[0090] Through stitching, a complete stitched image containing the target subject, mask, and edge information was obtained. This image not only preserves the original features of the target subject but also enhances its visual vividness and three-dimensionality through masking and edge enhancement.

[0091] For example, with Figure 6For example, the target subject area image, mask image, and edge image obtained through five-channel input splitting are stitched together along the channel dimension to form a stitched image with rich information. This image not only shows the details and colors of the target subject, but also enhances its overall visual effect through the addition of masking and edges.

[0092] Step 140: Perform style transfer processing on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

[0093] In some embodiments, style transfer processing is performed on the stitched image, the cue words, and the background image to obtain the target image, including: inputting the stitched image, the cue words, and the background image into a trained style transfer model for style transfer processing to obtain the target image; wherein, the trained style transfer model is trained based on multiple sample cue words and sample images matching each sample cue word.

[0094] Step 140 is a crucial step in the entire image processing workflow. It involves inputting the stitched image, the prompt word, and the background image into the trained style transfer model for style transfer processing, thereby obtaining a target image that matches the prompt word. This step not only requires the model to accurately capture and fuse the style features of different images, but also to ensure that the target image visually meets the user's expectations while maintaining the original features of the target subject.

[0095] During style transfer processing, the model first analyzes the style information of the target subject in the stitched image and the background image, while also parsing the style description in the prompts. Then, based on the prompts, the model extracts corresponding target background style features from the background image, such as color, texture, and composition. Next, the model fuses these target background style features with the target subject, ensuring that the target subject retains its original features while maintaining harmony with the new target background style.

[0096] Through processing by the trained style transfer model, the final target image matching the prompt word was obtained. This image not only contains the target subject but also has the target background style, fulfilling the user's dual requirements for image style and content. Furthermore, because the model learned from a large amount of sample data during training, it is able to handle various complex style transfer tasks and generate high-quality target images.

[0097] like Figure 5 and Figure 6 As shown, the stitched image, prompt word and background image are input into the trained style transfer model 10 for style transfer processing to obtain the target image.

[0098] In some embodiments, the trained style transfer model includes a local drawing module, an image cue adapter module, and a diffusion module. The process of inputting a stitched image, a cue word, and a background image into the trained style transfer model for style transfer processing to obtain a target image includes: inputting the stitched image and the cue word into the local drawing module for embedding processing to obtain an image embedding vector corresponding to the stitched image and a text embedding vector corresponding to the cue word; inputting the cue word and the background image into the image cue adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image; and inputting the image embedding vector, text embedding vector, first cross-attention score, and second cross-attention score into the diffusion module for style transfer processing to obtain the target image.

[0099] like Figure 5 and Figure 6 As shown, the trained style transfer model 10 includes a local drawing module 1, an image cue adapter module 2, and a diffusion module 3. The stitched image and the cue word are input into the local drawing module 1 for embedding processing, obtaining the image embedding vector corresponding to the stitched image and the text embedding vector corresponding to the cue word. The cue word and the background image are input into the image cue adapter module 2 for attention decoupling processing, obtaining the first cross-attention score corresponding to the cue word and the second cross-attention score corresponding to the background image. The image embedding vector, text embedding vector, first cross-attention score, and second cross-attention score are input into the diffusion module for style transfer processing to obtain the target image.

[0100] like Figure 6 As shown, the amplification module 3 can adopt a U-shaped network (U Net) structure, which can include multiple U-shaped network encoders 31 (UNetEncoder) and multiple U-shaped network decoders 32 (UNetDecoder).

[0101] UNet is a neural network architecture commonly used for image segmentation tasks. It employs an encoder-decoder structure and fuses features between the decoding and encoder stages through skip connections to better recover detailed information in the image.

[0102] like Figure 6 As shown, the local drawing module 1 may include multiple control encoders 11, and the output of the local drawing module 1 is connected to the U-shaped network decoder 32 of the amplification module 3.

[0103] like Figure 6As shown, the image prompt adapter module 2 may include a text encoding module 21, an image encoding module 22, and a decoupled cross-attention adaptation module 23. The decoupled cross-attention adaptation module 23 includes a first cross-attention layer 231 and a second cross-attention layer 232. The output of the text encoding module 21 is connected to the input of the first cross-attention layer 231, and the output of the first cross-attention layer 231 is connected to the U-shaped network decoder 32 of the amplification module 3. The output of the image encoding module 22 is connected to the input of the second cross-attention layer 232, and the output of the second cross-attention layer 232 is connected to the U-shaped network decoder 32 of the amplification module 3.

[0104] like Figure 6 As shown, the augmentation module 3 generates the target image under the dual guidance of prompt words and background images through the process of adding and removing noise. The time step refers to the number of steps required for noise removal. The larger the time step, the cleaner the noise removal and the higher the quality of the generated target image, but the corresponding time consumption will increase.

[0105] For example, the local painting module 1 can be an inpaint plugin, which can be trained using a ControlNet plugin. The local painting module 1 can include multiple control encoders 11, and a ControlNet plugin can be trained on the control encoder 11 side. Since the output of the local painting module 1 is injected into the U-shaped network decoder 32 of the amplification module 3 in the form of weights, the local painting module 1 and the amplification module 3 are essentially decoupled. During inference, the local painting module 1 and the amplification module 3 interact, and the weights of the local painting module 1 and the amplification module 3 are additive. Therefore, the local painting module 1 does not change the weight distribution of the U-shaped network decoder 32 in the amplification module 3.

[0106] ControlNet is a plugin for controlling image generation in artificial intelligence (AI). Based on Conditional Generative Adversarial Networks (CGANs), it allows users fine-grained control over the generated images. By providing additional input conditions to the model, ControlNet can precisely guide the model to generate images that meet user requirements, thereby improving the flexibility and controllability of image generation.

[0107] The multiple control encoders 11 in the local plotting module 1 can be understood as stacked fully connected networks. The input to the local plotting module 1 is a stitched image and a prompt word. The output of the local plotting module 1 is the text embedding vector corresponding to the prompt word and the image embedding vector corresponding to the five-channel stitched image. During the processing of the stitched image and prompt word, the multiple control encoders 11 may also include various cross-attention calculations to obtain the output weights of the local plotting module 1. The text embedding vector and the image embedding vector corresponding to the stitched image output by the local plotting module 1 can be input to the input side of the U-shaped network decoder 32 of the augmentation module 3 for decoding. The output weights of the local plotting module 1 can be superimposed on the output side of the U-shaped network decoder 32, jointly affecting the output image of the U-shaped network decoder 32.

[0108] In the background compositing process, the local drawing module 1 controls the degree of fusion between the edge contours of the target subject in the original image and the spatial relationship in the background image. The image cue adapter module 2 controls the similarity between the target image and the background image. The weights of the local drawing module 1 and the image cue adapter module 2 can be balanced by adjusting the first weight coefficient corresponding to the local drawing module 1 and / or the second weight coefficient corresponding to the image cue adapter module 2. In practical use, a first interface (adjustment range 0-1) can be provided to the user to adjust the first weight coefficient corresponding to the local drawing module 1, allowing the user to customize the desired fusion level. Similarly, a second interface (adjustment range 0-1) can be provided to the user to adjust the second weight coefficient corresponding to the image cue adapter module 2, allowing the user to customize the desired similarity level.

[0109] In some embodiments, the stitched image and the prompt word are embedded into the local drawing module to obtain an image embedding vector corresponding to the stitched image and a text embedding vector corresponding to the prompt word, including:

[0110] Obtain the first weight coefficient corresponding to the local drawing module. The first weight coefficient is used to control the degree of fusion between the edge contour of the target subject in the original image and the spatial relationship in the background image.

[0111] The first weighting coefficient, the stitched image, and the prompt word are input into the local drawing module for embedding processing to obtain the updated image embedding vector corresponding to the stitched image and the updated text embedding vector corresponding to the prompt word.

[0112] For example, a first interface (within a range of 0 to 1) can be provided to users to adjust the first weight coefficient corresponding to the local drawing module 1, allowing users to customize the desired blending effect. For instance, the first interface could be a slider or input box between 0 and 1. Users can adjust the first weight coefficient through the first interface as needed, thereby fine-tuning the composite effect. For example, if they want the target subject to blend more closely with the background, forming a more realistic overall visual effect, they might increase the value of the first weight coefficient. Conversely, if they want to maintain the sharpness of the target subject's edges so that it stands out against the background, they might decrease the value of the first weight coefficient.

[0113] This adjustable first weighting coefficient provides flexibility, allowing users to customize the results according to their specific needs and preferences. The first weighting coefficient controls the degree to which the edge contours of the target subject in the original image blend with the spatial relationship in the background image. For example, when displaying products in e-commerce, it may be necessary to place the product (target subject) in a spatial location within a background image (such as a table), and it is desirable for the product to blend naturally into the scene while maintaining its clear recognizability. By appropriately adjusting the first weighting coefficient, it is possible to ensure that the product edges are clearly visible while also guaranteeing the realism and consistency of background details.

[0114] For example, in the background compositing process, cosmetics (the target subject) are placed on a table (the spatial position in the background image). The spatial relationship between the edge contour of the target subject and the background image refers to the ability to blend the edges of the cosmetics and the table naturally, while making the background look realistic. Generally speaking, the larger the first weight coefficient, the more closely the target subject and the background are combined, and the better the background is drawn. However, the edge constraints will be weakened. Therefore, in order to balance the relationship between the two, an appropriate setting of the first weight coefficient is needed.

[0115] The first weighting coefficient can be used to update the model parameters of the local drawing module 1, thereby changing the output of the local drawing module 1. Specifically, the first weighting coefficient, the stitched image, and the prompt are input into the local drawing module for embedding processing, resulting in an updated image embedding vector corresponding to the stitched image and an updated text embedding vector corresponding to the prompt. "Updated" means that the generation of these vectors has been influenced by the user-specified first weighting coefficient, thus determining the degree of fusion between the target subject and the background.

[0116] In some embodiments, the input image cue adapter module is subjected to attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image. This includes: obtaining a second weight coefficient corresponding to the image cue adapter module, the second weight coefficient being used to characterize the similarity between the target image and the background image; and inputting the second weight coefficient, the cue word, and the background image into the image cue adapter module for attention decoupling processing to obtain an updated first cross-attention score corresponding to the cue word and an updated second cross-attention score corresponding to the background image.

[0117] To regulate the operation of the image tooltip adapter module 2, a second weighting coefficient is introduced. This second weighting coefficient characterizes the similarity between the target image and the background image; that is, how consistent the style and content features of the background should be with the original background image in the final output target image. For example, a second interface (within a range of 0 to 1) can be provided to the user to adjust the second weighting coefficient corresponding to the image tooltip adapter module 2, allowing the user to customize the desired similarity effect. For example, the second interface could be a slider or input box between 0 and 1. 0 indicates that the features of the background image are not considered at all, while 1 indicates that the similarity with the background image is maintained as much as possible. Adjusting this second weighting coefficient directly affects the degree to which the style transfer model applies background image features when generating the target image.

[0118] The second weighting coefficient can be used to update the model parameters of the image prompt adapter module 2, thereby changing the output of the image prompt adapter module 2.

[0119] For example, if a user wants the generated image to retain more of the background style, they can set the second weight coefficient to a higher value, such as 0.8. This will cause the model to adopt more of the style and content of the background image, making the generated target image visually closer to the selected background. In contrast, if the user wants the subject to stand out more, or wants the background style to be less obvious, they might choose a lower weight value, such as 0.5.

[0120] like Figure 7 As shown, given Figure 7 The original image shown in (a) Figure 7 (b) shows the background image. When the second weight coefficient of the input is 0.5, the style transfer model 10 finally generates the image shown. Figure 7 The target image shown in (c).

[0121] like Figure 7 As shown, given Figure 7 The original image shown in (a) Figure 7(b) shows the background image. When the second weight coefficient of the input is 0.8, the style transfer model 10 finally generates the image shown. Figure 7 The target image shown in (d) Figure 7 The target image shown in (d) is compared to Figure 7 (c) The target image is more similar to the background image.

[0122] In this way, the image tooltip adapter module and its second weighting coefficient provide users with flexible style control, enabling them to create target images with different background similarities as needed. This personalized adjustment not only helps meet diverse user needs but also increases the fun and engagement of the interactive experience.

[0123] In some embodiments, the image cue adapter module includes a text encoding module, an image encoding module, and a decoupled cross-attention adaptation module. The cue word and background image are input into the image cue adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image. This includes: inputting the cue word into the text encoding module for encoding processing to obtain text features; inputting the background image into the image encoding module for encoding processing to obtain image features; and inputting the text features and image features into the decoupled cross-attention adaptation module for attention decoupling processing to obtain the first cross-attention score corresponding to the cue word and the second cross-attention score corresponding to the background image.

[0124] like Figure 6 As shown, the image prompt adapter module 2 may include a text encoding module 21, an image encoding module 22, and a decoupled cross-attention adaptation module 23. The decoupled cross-attention adaptation module 23 includes a first cross-attention layer 231 and a second cross-attention layer 232. The output of the text encoding module 21 is connected to the input of the first cross-attention layer 231, and the output of the first cross-attention layer 231 is connected to the U-shaped network decoder 32 of the amplification module 3. The output of the image encoding module 22 is connected to the input of the second cross-attention layer 232, and the output of the second cross-attention layer 232 is connected to the U-shaped network decoder 32 of the amplification module 3. The prompt word is input into the text encoding module 21 for encoding processing to obtain text features; the background image is input into the image encoding module 22 for encoding processing to obtain image features; the text features are input into the first cross-attention layer 231 of the decoupled cross-attention adaptation module 23 for attention decoupling processing to obtain the first cross-attention score corresponding to the prompt word; and the image features are input into the second cross-attention layer 232 of the decoupled cross-attention adaptation module 23 for attention decoupling processing to obtain the second cross-attention score corresponding to the background image.

[0125] For example, a linear transformation module and a layer normalization (LN) module can be set between the image encoding module 22 and the second cross-attention layer 232. Before being input into the second cross-attention layer 232, the image encoding module 22 output by the image encoding module 22 can undergo feature processing through the linear transformation module and the LN module. The linear transformation module can be used to perform linear transformation on image features, adjusting their dimensions and shape to adapt to the input requirements of the second cross-attention layer 232. The LN module can be used to normalize image features to eliminate the dimensional differences between different features and improve the stability and robustness of the model.

[0126] In some embodiments, in the above Figure 4 Based on the corresponding embodiments, this application also provides another optional embodiment of the image processing method, such as... Figure 8 As shown, the steps for training the style transfer model include the following steps S81 to S84:

[0127] S81, acquire multiple sample prompt words and sample images that match each sample prompt word.

[0128] In training the style transfer model, a rich sample dataset needs to be constructed, containing multiple sample prompts and corresponding sample images for each prompt. Sample prompts can be descriptive text expressing the background style, content (subject), or theme of the image the user desires. Sample images are existing real images that match the sample prompts, representing the desired image effect.

[0129] These sample data can be obtained through various means. For example, sample images that meet the criteria can be selected from publicly available image databases. Additionally, user-uploaded sample images or user-created sample images can be collected through crowdsourcing platforms. When collecting sample data, attention must be paid to both data quality and diversity. On the one hand, it is essential to ensure a high degree of matching between sample prompts and sample images, meaning that each sample prompt should have a matching sample image. On the other hand, it is also crucial to ensure the diversity of the sample data, covering images with different subjects and background styles, so that the model can learn richer features and information.

[0130] S82, based on the first sample prompt word among multiple sample prompt words and the first sample image matching the first sample prompt word, train the local drawing module and diffusion module in the style transfer model, and determine the first loss function.

[0131] In some embodiments, based on a first sample prompt word from a plurality of sample prompt words and a first sample image matching the first sample prompt word, a local drawing module and a diffusion module in the style transfer model are trained to determine a first loss function, including: splitting the first sample image to obtain a first sample main body region map, a first sample mask map, and a first sample edge map, wherein the first sample mask map is used to cover the first sample main body region map, and the first sample edge map is used to depict the edge contour of the first sample main body region map; stitching the first sample main body region map, the first sample mask map, and the first sample edge map together to obtain a first sample stitched image; inputting the first sample stitched image and the first sample prompt word into the local drawing module for embedding to obtain a first sample image embedding vector corresponding to the first sample stitched image and a first sample text embedding vector corresponding to the first sample prompt word; inputting the first sample prompt word, the first sample image, the first sample image embedding vector, and the first sample text embedding vector into the diffusion module for style transfer to obtain a first generated image; and determining the first loss function based on the mean squared error loss function between the first generated image and the first sample image.

[0132] In training the local plotting and diffusion modules of the style transfer model, multiple sample data pairs were selected. Each sample data pair contains a sample cue word and its corresponding sample image (i.e., a real image). The sample cue word typically contains a detailed description of the content of the image to be generated, while the sample image (i.e., the real image) is a visual instance of that description. For example, the input sample data pair for each training round is the first sample cue word and its associated first sample image.

[0133] For example, such as Figure 9 As shown, the first sample prompt word and the first sample image are obtained. For example, the first sample prompt word is "a bottle of soda, placed on a table, with a blurred background of light spots", and the first sample image displays the image content of "a bottle of soda, placed on a table, with a blurred background of light spots".

[0134] To more precisely control the quality and detail of the generated images, the first sample image undergoes meticulous preprocessing. Specifically, through a splitting process, the original image is divided into three parts: a first sample main body region map (i.e., the core content of the image), a first sample mask map (used to distinguish and highlight the main body region), and a first sample edge map (depicting the outline of the main body region). The purpose of this is to enable the style transfer model 10 to better understand and learn the key elements of the image structure.

[0135] Next, these three components are stitched together to form the first sample stitched image. This stitching structure helps the model understand and fuse image features from multiple perspectives.

[0136] After image stitching is completed, the local drawing module 1 receives the first sample stitched image and the corresponding first sample prompt word as input. The local drawing module 1 converts the first sample stitched image into a first sample image embedding vector through embedding technology, and at the same time converts the first sample prompt word into a first sample text embedding vector. This step enables the model to map language description and visual information into the same semantic space.

[0137] Subsequently, these embedded vectors are input into diffusion module 3. Diffusion module 3 uses a process of progressive diffusion and de-diffusion to simulate the image generation process, adjusting model parameters to achieve style transfer based on cue word descriptions. The final output first generated image should retain the main content features of the first sample image while also demonstrating consistency with the style described by the cue words in the first sample image.

[0138] To quantify the difference between the first generated image and the first sample image, and to guide the training of the style transfer model 10 accordingly, a first loss function based on the mean squared error (MSE) loss function is defined. This function measures the pixel-level difference between the first generated image and the first sample image. By minimizing this loss function, the style transfer model 10 is prompted to iteratively update its model parameters, thereby more accurately transferring the target style while preserving the content, thus improving the quality and realism of the generated image. For example, when the first sample prompt is "a bottle of soda, placed on a table, with a blurred background of light spots," the style transfer model 10 should be able to generate an image with corresponding scene and style features based on this description, and the visual difference between the first generated image and the actual corresponding first sample image should be as small as possible.

[0139] S83, based on the first loss function, the second sample cue word among multiple sample cue words, the second sample image matching the second sample cue word, and the third sample image matching the third sample cue word among multiple sample cue words, train the image cue adapter module and the diffusion module in the style transfer model to determine the second loss function, wherein the second sample image and the third sample image correspond to the same sample subject but have different background styles.

[0140] In some embodiments, the image cue adapter module and the diffusion module in the style transfer model are trained based on a first loss function, a second sample cue word among multiple sample cue words, a second sample image matching the second sample cue word, and a third sample image matching the third sample cue word among multiple sample cue words, and the second loss function is determined by: updating the parameters of the style transfer model according to the first loss function to obtain an updated style transfer model; inputting the second sample cue word and the third sample image into the image cue adapter module in the updated style transfer model for attention decoupling processing to obtain a first cross-attention score corresponding to the second sample cue word and a second cross-attention score corresponding to the third sample image; inputting the second sample cue word, the third sample image, the first cross-attention score corresponding to the second sample cue word, and the second cross-attention score corresponding to the third sample image into the diffusion module in the updated style transfer model for style transfer processing to obtain a second generated image; and determining the second loss function based on the mean squared error loss function between the second generated image and the second sample image.

[0141] For example, the image cue adapter module 2 can be an image cue adapter (IP Adapter) plugin, providing an image guidance interface for the augmentation module 3, and simultaneously fine-tuning the output weights of the image cue adapter module 2 and the weights of the augmentation module 3 during training. During training, the image cue adapter module 2 uses a pre-trained IP Adapter plugin, and the second loss function is determined through joint training of the image cue adapter module 2 and the augmentation module 3.

[0142] For example, such as Figure 10 As shown, a second sample prompt, a second sample image, and a third sample image are obtained. For example, the second sample prompt is "a bottle of shower gel, placed on a table, light green wallpaper"; the second sample image displays the image content of "a bottle of shower gel, placed on a table, light green wallpaper"; and the third sample image displays the image content of "a bottle of shower gel, placed on a stone, grass background". The subject of the second sample image is "a bottle of shower gel" with a "light green wallpaper" background style, and the subject of the third sample image is "a bottle of shower gel" with a "grass background" background style. The second and third sample images correspond to the same subject but have different background styles.

[0143] After initially training the local drawing module 1 and diffusion module 3 in the style transfer model 10 to determine the first loss function, the image cue adapter module 2 and diffusion module 3 in the style transfer model 10 are further jointly trained. Specifically, the parameters of the style transfer model 10 are updated according to the first loss function to obtain the updated style transfer model 10. This step is based on the model's existing training results (the first loss function) and adjusts the model weights through gradient backpropagation to ensure that the model can be further optimized based on previous learning results in subsequent fine-tuning. Then, the second sample cue word and the third sample image are input into the updated image cue adapter module 2 in the style transfer model 10 for attention decoupling processing, obtaining the first cross-attention score corresponding to the second sample cue word and the second cross-attention score corresponding to the third sample image. The function of the image cue adapter module 2 is to decouple the attention distribution of text cue and image content, enabling the model to process them independently. Then, the second sample prompt, the third sample image, the first cross-attention score corresponding to the second sample prompt, and the second cross-attention score corresponding to the third sample image are input into the diffusion module of the updated style transfer model for style transfer processing, resulting in a second generated image. This second generated image should retain the original sample subject (a bottle of shower gel), but the background style is transformed from the "grass background" style in the original input third sample image to the "light green wallpaper" style corresponding to the description of the sample prompt. By comparing the difference between the second generated image (sample subject is a bottle of shower gel, background style is light green wallpaper) and the second sample image (sample subject is a bottle of shower gel, background style is light green wallpaper), the second loss function is calculated using the mean squared error loss function (MSE Loss). This loss function measures the accuracy of the model in background style transfer, aiming to make the second generated image as close as possible to the background style specified by the second sample prompt while retaining the subject unchanged.

[0144] S84, based on the first loss function and the second loss function, the parameters of the style transfer model are updated to obtain the trained style transfer model.

[0145] The first loss function typically focuses on basic content reconstruction and fidelity, ensuring that the model can accurately capture and reproduce the core elements of the input image or generate reasonable image content based on text prompts. The second loss function, on the other hand, focuses on the specific effect of style transfer; for example, in the scenario above, it reflects the model's ability to successfully transfer different background styles while keeping the subject unchanged.

[0146] In actual training, based on gradient descent or other optimization algorithms, the results of both loss functions are considered together to update all parameters of the style transfer model. Using backpropagation or other gradient backpropagation methods, the gradients of the first and second loss functions with respect to the model parameters are calculated separately, and the model's weights and biases are updated based on these gradients. This means that in each training iteration, it is necessary not only to reduce the content reconstruction error based on the first loss function but also to reduce the style transfer error based on the second loss function.

[0147] As the training process progresses, the model parameters undergo a series of fine-tuning and optimizations, gradually enhancing its performance. Once the model reaches the preset convergence criterion or the number of training epochs, the training is considered complete. The resulting style transfer model possesses robust style transfer capabilities, preserving high-quality images that retain the original subject features while also reflecting the desired style.

[0148] In each batch of iterative training phases, the training processes of the local drawing module 1 and the image cue adapter module 2 in the style transfer model 10 are decoupled. When the local drawing module 1 and the diffusion module 3 are jointly trained, the image cue adapter module 2 is inactive; when the image cue adapter module 2 and the diffusion module 3 are jointly trained, the local drawing module 1 is inactive.

[0149] During the inference phase, the local drawing module 1 and the image prompt adapter module 2 in the trained style transfer model 10 can be activated simultaneously. The trained style transfer model 10 can perform style transfer on the original image based on the prompt words input by the user, the original image containing the target subject, and the background image with the target background style. Through the dual guidance of the prompt words and the background image, the original image can be generated to contain the target subject and have the target background style. The target subject and the target background style in the target image have a high degree of fusion, which significantly improves the style transfer effect.

[0150] This application's embodiments can be applied to advertising systems with AIGC capabilities. A crucial application scenario for subject-background composites is providing merchants with a vast amount of creative advertising materials to assist them in ad placement and traffic attraction. Utilizing the image processing method provided in this application's embodiments, merchants can easily "copy" similar background styles (such as the background of a popular advertisement), generating high-quality subject-background composite images. This is a powerful way to facilitate the implementation of AIGC in advertising systems.

[0151] This application embodiment supports multimodal input of both text prompts and background images, accepts dual guidance from text prompts and background images, performs style transfer on the original image to generate a target image containing the target subject and having the target background style, and can provide users with more possibilities for background synthesis.

[0152] In practical use, this application embodiment can provide users with a first interface to adjust the first weight coefficient corresponding to the local drawing module, allowing users to customize the desired blending effect. It can also provide users with a second interface to adjust the second weight coefficient corresponding to the image tooltip adapter module, allowing users to customize the desired similarity effect. By adjusting the first and / or second weight coefficients, the weights of the local drawing module and the image tooltip adapter module are balanced.

[0153] All of the above technical solutions can be combined in any way to form optional embodiments of this application, and will not be described in detail here.

[0154] This embodiment of the application obtains a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to guide the generation of an image containing the target subject and having the target background style. The original image is split to obtain a target subject region image, a mask image, and an edge image. The mask image is used to cover the target subject region image, and the edge image is used to depict the edge contour of the target subject region image. The target subject region image, mask image, and edge image are stitched together to obtain a stitched image. Style transfer processing is performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style. This embodiment of the application achieves precise control over the subject-background synthesis by comprehensively utilizing the prompt word, the original image, and the background image with the target background style. The prompt word guides the generation process, ensuring that the generated image meets specific style and theme requirements. The original image provides the target subject to be synthesized, while the background image provides the target background style. The combination of the two ensures that the generated image retains the characteristics of the target subject while incorporating the style of the target background. By splitting the original image, a target subject region map, a mask map, and an edge map are obtained. These parts are then stitched together, allowing for more precise control over the position and shape of the target subject during style transfer. The mask map covers the target subject region map, ensuring the integrity of the subject during compositing. The edge map depicts the edge contours of the target subject region map, enhancing the blending between the subject and the background. Furthermore, by combining the stitched image, the cue word, and the background image for style transfer processing, a target image matching the cue word can be generated. This processing method not only results in a higher quality image but also ensures a higher degree of blending between the target subject and the target background style. This embodiment achieves precise control over the style transfer process through the dual guidance of cue words and background images. This processing method overcomes the problem of insufficient style transfer blending in current technologies, improving the effect and quality of subject-background compositing.

[0155] To facilitate better implementation of the image processing method of this application embodiment, this application embodiment also provides an image processing apparatus. Please refer to... Figure 11 , Figure 11 This is a schematic diagram of the structure of an image processing apparatus provided in an embodiment of this application. The image processing apparatus 200 may include:

[0156] The acquisition unit 210 is used to acquire a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style.

[0157] The splitting unit 220 is used to split the original image to obtain a target subject region map, a mask map and an edge map. The mask map is used to cover the target subject region map and the edge map is used to depict the edge contour of the target subject region map.

[0158] The stitching unit 230 is used to stitch together the target subject area map, the mask map, and the edge map to obtain a stitched image;

[0159] The processing unit 240 is used to perform style transfer processing based on the spliced ​​image, the prompt word and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

[0160] In some embodiments, the processing unit 240 is specifically used to: input the stitched image, the cue words and the background image into the trained style transfer model for style transfer processing to obtain the target image; wherein, the trained style transfer model is trained based on multiple sample cue words and sample images matching each sample cue word.

[0161] In some embodiments, the trained style transfer model includes a local drawing module, an image cue adapter module, and a diffusion module; the processing unit 240 includes: a first processing subunit, used to input the stitched image and the cue word into the local drawing module for embedding processing to obtain an image embedding vector corresponding to the stitched image and a text embedding vector corresponding to the cue word; a second processing subunit, used to input the cue word and the background image into the image cue adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image; and a third processing subunit, used to input the image embedding vector, the text embedding vector, the first cross-attention score, and the second cross-attention score into the diffusion module for style transfer processing to obtain the target image.

[0162] In some embodiments, the first processing subunit is specifically used to: obtain a first weight coefficient corresponding to the local drawing module, the first weight coefficient being used to control the degree of fusion between the edge contour of the target subject in the original image and the spatial relationship in the background image; input the first weight coefficient, the stitched image, and the prompt word into the local drawing module for embedding processing to obtain the updated image embedding vector corresponding to the stitched image and the updated text embedding vector corresponding to the prompt word.

[0163] In some embodiments, the second processing subunit is specifically used to: obtain a second weight coefficient corresponding to the image cue adapter module, the second weight coefficient being used to characterize the similarity between the target image and the background image; input the second weight coefficient, the cue word, and the background image into the image cue adapter module for attention decoupling processing to obtain an updated first cross-attention score corresponding to the cue word and an updated second cross-attention score corresponding to the background image.

[0164] In some embodiments, the image cue adapter module includes a text encoding module, an image encoding module, and a decoupled cross-attention adaptation module; the second processing subunit is specifically used for: inputting the cue word into the text encoding module for encoding processing to obtain text features; inputting the background image into the image encoding module for encoding processing to obtain image features; inputting the text features and image features into the decoupled cross-attention adaptation module for attention decoupling processing to obtain a first cross-attention score corresponding to the cue word and a second cross-attention score corresponding to the background image.

[0165] In some embodiments, the image processing apparatus 200 further includes a training unit, configured to: acquire a plurality of sample cue words and sample images matching each sample cue word; train a local drawing module and a diffusion module in the style transfer model based on a first sample cue word among the plurality of sample cue words and a first sample image matching the first sample cue word, and determine a first loss function; train an image cue adapter module and a diffusion module in the style transfer model based on the first loss function, a second sample cue word among the plurality of sample cue words, a second sample image matching the second sample cue word, and a third sample image matching a third sample cue word among the plurality of sample cue words, and determine a second loss function, wherein the second sample image and the third sample image correspond to the same sample subject but have different background styles; update the parameters of the style transfer model based on the first loss function and the second loss function to obtain the trained style transfer model.

[0166] In some embodiments, when the training unit trains the local drawing module and the diffusion module in the style transfer model based on the first sample prompt word among multiple sample prompt words and the first sample image matching the first sample prompt word, and determines the first loss function, the specific steps are as follows: The first sample image is split to obtain a first sample main body region map, a first sample mask map, and a first sample edge map, wherein the first sample mask map is used to cover the first sample main body region map, and the first sample edge map is used to depict the edge contour of the first sample main body region map; the first sample main body region map, the first sample mask map, and the first sample edge map are stitched together to obtain a first sample stitched image; the first sample stitched image and the first sample prompt word are input into the local drawing module for embedding to obtain a first sample image embedding vector corresponding to the first sample stitched image and a first sample text embedding vector corresponding to the first sample prompt word; the first sample prompt word, the first sample image, the first sample image embedding vector, and the first sample text embedding vector are input into the diffusion module for style transfer to obtain a first generated image; and the first loss function is determined based on the mean squared error loss function between the first generated image and the first sample image.

[0167] In some embodiments, when the training unit trains the image cue adapter module and the diffusion module in the style transfer model based on the first loss function, the second sample cue word among multiple sample cue words, the second sample image matching the second sample cue word, and the third sample image matching the third sample cue word among multiple sample cue words, and determines the second loss function, it specifically performs the following steps: updating the parameters of the style transfer model according to the first loss function to obtain an updated style transfer model; inputting the second sample cue word and the third sample image into the image cue adapter module in the updated style transfer model for attention decoupling processing to obtain a first cross-attention score corresponding to the second sample cue word and a second cross-attention score corresponding to the third sample image; inputting the second sample cue word, the third sample image, the first cross-attention score corresponding to the second sample cue word, and the second cross-attention score into the diffusion module in the updated style transfer model for style transfer processing to obtain a second generated image; and determining the second loss function based on the mean squared error loss function between the second generated image and the second sample image.

[0168] In some embodiments, the splitting unit 220 is specifically used for: performing image matting on the original image to obtain a mask image; performing composite processing on the mask image and the original image to obtain a target subject region image; and performing edge detection on the mask image to obtain an edge image.

[0169] It should be noted that the functions of each unit in the image processing apparatus 200 in this application embodiment can be referred to the specific implementation of any embodiment in the above method embodiments, and will not be repeated here.

[0170] Each unit in the above-described device can be implemented entirely or partially through software, hardware, or a combination thereof. Each unit can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each unit.

[0171] For example, the image processing device 200 may be integrated into a terminal or server that has storage and a processor and thus computing power, or the image processing device 200 may be the terminal or server.

[0172] In some embodiments, this application also provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0173] Figure 12 A schematic structural diagram of the computer device provided in the embodiments of this application, such as Figure 12 As shown, the computer device 300 may include: a communication interface 301, a memory 302, a processor 303, and a communication bus 304. The communication interface 301, memory 302, and processor 303 communicate with each other via the communication bus 304. The communication interface 301 is used for data communication between the device 300 and external devices. The memory 302 can be used to store software programs and modules, and the processor 303 runs the software programs and modules stored in the memory 302, such as the software programs for the corresponding operations in the aforementioned method embodiments.

[0174] In some embodiments, the processor 303 may invoke software programs and modules stored in the memory 302 to perform the following operations:

[0175] The process involves obtaining a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to generate an image that contains the target subject and has the target background style. The original image is split to obtain a target subject region image, a mask image, and an edge image. The mask image is used to cover the target subject region image, and the edge image is used to depict the edge contour of the target subject region image. The target subject region image, mask image, and edge image are then stitched together to obtain a stitched image. Style transfer processing is performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

[0176] In some embodiments, the computer device 300 may be integrated into a terminal or server that has storage and a processor and thus computing power, or the computer device 300 may be the terminal or server.

[0177] This application also provides a computer-readable storage medium for storing a computer program. This computer-readable storage medium can be applied to a computer device, and the computer program causes the computer device to execute the corresponding processes in the methods described above in the embodiments of this application; for brevity, further details are omitted here.

[0178] This application also provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0179] This application also provides a computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the corresponding processes in the methods described above in the embodiments of this application. For brevity, these details will not be elaborated further here.

[0180] It should be understood that the processor in the embodiments of this application may be an integrated circuit chip with signal processing capabilities. In implementation, the steps of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor described above can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0181] It is understood that the memory in the embodiments of this application can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0182] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0183] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0184] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0185] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0186] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0187] In addition, the functional units in the embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0188] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer or a server) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.

[0189] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method, characterized in that, The method includes: The system obtains a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style. The original image is split to obtain a target subject region map, a mask map, and an edge map. The mask map is used to cover the target subject region map, and the edge map is used to depict the edge contour of the target subject region map. The target subject region map, the mask map, and the edge map are stitched together to obtain a stitched image; Style transfer processing is performed on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word. The target image contains the target subject and has the target background style.

2. The image processing method as described in claim 1, characterized in that, The step of performing style transfer processing based on the stitched image, the prompt word, and the background image to obtain the target image includes: The stitched image, the prompt word, and the background image are input into the trained style transfer model for style transfer processing to obtain the target image. The trained style transfer model is trained based on multiple sample cue words and sample images that match each sample cue word.

3. The image processing method as described in claim 2, characterized in that, The trained style transfer model includes a local drawing module, an image cue adapter module, and a diffusion module. The step of inputting the stitched image, the prompt word, and the background image into a trained style transfer model for style transfer processing to obtain the target image includes: The stitched image and the prompt word are input into the local drawing module for embedding processing to obtain the image embedding vector corresponding to the stitched image and the text embedding vector corresponding to the prompt word; The prompt word and the background image are input into the image prompt adapter module for attention decoupling processing to obtain the first cross-attention score corresponding to the prompt word and the second cross-attention score corresponding to the background image; The image embedding vector, the text embedding vector, the first cross-attention score, and the second cross-attention score are input into the diffusion module for style transfer processing to obtain the target image.

4. The image processing method as described in claim 3, characterized in that, The step of embedding the stitched image and the prompt word into the local drawing module to obtain the image embedding vector corresponding to the stitched image and the text embedding vector corresponding to the prompt word includes: Obtain the first weight coefficient corresponding to the local drawing module. The first weight coefficient is used to control the degree of fusion between the edge contour of the target subject in the original image and the spatial relationship in the background image. The first weight coefficient, the stitched image, and the prompt word are input into the local drawing module for embedding processing to obtain the updated image embedding vector corresponding to the stitched image and the updated text embedding vector corresponding to the prompt word.

5. The image processing method as described in claim 3, characterized in that, The step of inputting the prompt word and the background image into the image prompt adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the prompt word and a second cross-attention score corresponding to the background image includes: Obtain the second weight coefficient corresponding to the image prompt adapter module. The second weight coefficient is used to characterize the similarity between the target image and the background image. The second weight coefficient, the prompt word, and the background image are input into the image prompt adapter module for attention decoupling processing to obtain the updated first cross-attention score corresponding to the prompt word and the updated second cross-attention score corresponding to the background image.

6. The image processing method as described in claim 3, characterized in that, The image prompt adapter module includes a text encoding module, an image encoding module, and a decoupled cross-attention adaptation module; The step of inputting the prompt word and the background image into the image prompt adapter module for attention decoupling processing to obtain a first cross-attention score corresponding to the prompt word and a second cross-attention score corresponding to the background image includes: The prompt words are input into the text encoding module for encoding processing to obtain text features; The background image is input into the image encoding module for encoding processing to obtain image features; The text features and the image features are input into the decoupled cross-attention adaptation module for attention decoupling processing to obtain the first cross-attention score corresponding to the prompt word and the second cross-attention score corresponding to the background image.

7. The image processing method as described in claim 3, characterized in that, The steps for training the style transfer model include: Obtain multiple sample prompt words and sample images that match each of the sample prompt words; Based on the first sample prompt word among the plurality of sample prompt words and the first sample image matching the first sample prompt word, the local drawing module and the diffusion module in the style transfer model are trained to determine the first loss function; Based on the first loss function, the second sample prompt word among the plurality of sample prompt words, the second sample image matching the second sample prompt word, and the third sample image matching the third sample prompt word among the plurality of sample prompt words, the image prompt adapter module and the diffusion module in the style transfer model are trained to determine the second loss function, wherein the second sample image and the third sample image correspond to the same sample subject but have different background styles; The style transfer model is updated with parameters based on the first loss function and the second loss function to obtain the trained style transfer model.

8. The image processing method as described in claim 7, characterized in that, The step of training the local drawing module and the diffusion module in the style transfer model based on the first sample prompt word among the plurality of sample prompt words and the first sample image matching the first sample prompt word, and determining the first loss function, includes: The first sample image is split to obtain a first sample main body region image, a first sample mask image, and a first sample edge image. The first sample mask image is used to cover the first sample main body region image, and the first sample edge image is used to depict the edge contour of the first sample main body region image. The first sample main area image, the first sample mask image and the first sample edge image are stitched together to obtain the first sample stitched image; The first sample stitched image and the first sample prompt word are input into the local drawing module for embedding processing to obtain the first sample image embedding vector corresponding to the first sample stitched image and the first sample text embedding vector corresponding to the first sample prompt word. The first sample prompt word, the first sample image, the first sample image embedding vector, and the first sample text embedding vector are input into the diffusion module for style transfer processing to obtain the first generated image. The first loss function is determined based on the mean squared error loss function between the first generated image and the first sample image.

9. The image processing method as described in claim 7, characterized in that, The step of training the image cue adapter module and the diffusion module in the style transfer model based on the first loss function, the second sample cue word among the plurality of sample cue words, the second sample image matching the second sample cue word, and the third sample image matching the third sample cue word among the plurality of sample cue words, and determining the second loss function, includes: The style transfer model is updated with parameters based on the first loss function to obtain the updated style transfer model. The second sample cue word and the third sample image are input into the image cue adapter module in the updated style transfer model for attention decoupling processing to obtain the first cross-attention score corresponding to the second sample cue word and the second cross-attention score corresponding to the third sample image. The second sample prompt word, the third sample image, the first cross-attention score corresponding to the second sample prompt word, and the second cross-attention score corresponding to the third sample image are input into the diffusion module in the updated style transfer model for style transfer processing to obtain the second generated image; The second loss function is determined based on the mean squared error loss function between the second generated image and the second sample image.

10. The image processing method as described in claim 1, characterized in that, The step of splitting the original image to obtain a target subject region map, a mask map, and an edge map includes: The original image is processed by image matting to obtain the mask image; The mask image is combined with the original image to obtain the target subject region image; Edge detection is performed on the mask image to obtain the edge image.

11. An image processing apparatus, characterized in that, The device includes: The acquisition unit is used to acquire a prompt word, an original image containing the target subject, and a background image with the target background style. The prompt word is used to prompt the generation of an image containing the target subject and having the target background style. A splitting unit is used to split the original image to obtain a target subject region map, a mask map, and an edge map. The mask map is used to cover the target subject region map, and the edge map is used to depict the edge contour of the target subject region map. The stitching unit is used to stitch the target subject area map, the mask map, and the edge map together to obtain a stitched image; The processing unit is configured to perform style transfer processing on the stitched image, the prompt word, and the background image to obtain a target image that matches the prompt word, wherein the target image contains the target subject and has the target background style.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program adapted for loading by a processor to perform the image processing method as described in any one of claims 1-10.

13. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program, and the processor executing the image processing method as described in any one of claims 1-10 by calling the computer program stored in the memory.

14. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the image processing method according to any one of claims 1-10.

Citation Information

Cited By

  • Editing digital images based on masks

    US12700157B2

  • Editing digital images based on masks

    US20250384601A1