Image processing method, device, equipment, medium and program product

Through the methods of cross-modal learning and multimodal learning, the concept injection module and edge control module are introduced, which solves the conceptual cognitive difference and edge deformation problems of the literary graphics model when changing the background, and generates high-quality images.

CN120472044APending Publication Date: 2025-08-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410179153.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

When the existing literary and artistic graphics model replaces the image background, there are problems with poor conceptual cognition and edge deformation of the specified objects in the generated new image, resulting in poor fusion between the image subject and the background.

Method used

Using cross-modal learning and multi-modal learning methods, concept injection module and edge control module are introduced to generate high-quality images through cross-modal attention feature fusion processing and edge optimization processing.

Benefits of technology

Ensure that the target object in the generated image matches the text description with high degree of matching, the edges are not deformed, the background and foreground are naturally blended, and the image quality is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472044A_ABST
    Figure CN120472044A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and device, equipment, a medium and a program product. The method comprises the steps of obtaining a to-be-processed text description and a first image; performing cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate cross-modal attention features; and performing edge optimization processing on the cross-modal attention feature according to the edge feature of the target object in the first image to obtain a second image. By adopting the embodiment of the invention, the quality of the target object in the generated image can be ensured while the image background is replaced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to the field of artificial intelligence, and specifically to an image processing method, an image processing apparatus, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Text-generated graphics refers to the process of generating images based on text content.

[0003] Currently, it is supported to use open source text-based image models to process input text to generate an image subject (such as a target object) and an image background, thereby composing a new image based on the image subject and the image background. However, when the existing background of a specified object is replaced based on the image background described in the text, there are problems with poor conceptual cognition and edge deformation of the specified object in the generated new image. For example, when using the open source text-based image models SD Inpainting and ControlNet Inpainting to replace the background of a specified object, there will be a poor match between the image subject in the new image and the image subject described in the text, and the edge of the subject in the new image will be deformed, resulting in poor fusion between the image subject and the image background. Summary of the Invention

[0004] The embodiments of the present application provide an image processing method, apparatus, device, medium, and program product, which can replace the image background while ensuring the quality of the target object in the generated image.

[0005] In one aspect, an embodiment of the present application provides an image processing method, the method comprising:

[0006] Obtaining a text description to be processed and a first image; the first image contains a target object, and the text description is used to describe the scene content containing the target object and the target background;

[0007] Performing cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; the cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background;

[0008] The cross-modal attention features are edge-optimized according to the edge characteristics of the target object in the first image to obtain the second image.

[0009] On the other hand, an embodiment of the present application provides an image processing device, the device comprising:

[0010] An acquisition unit is configured to acquire a text description to be processed and a first image; the first image includes a target object, and the text description is used to describe a scene content including the target object and a target background;

[0011] a processing unit, configured to perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; the cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background;

[0012] A processing unit is used to perform edge optimization processing on the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain a second image.

[0013] In one implementation, the processing unit is configured to perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate the cross-modal attention feature, specifically for:

[0014] Performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object;

[0015] Performing an attention operation on the target object features based on the attention mechanism to generate a first feature attention map; the first feature attention map is used to indicate the object characteristics of the target object; and

[0016] Based on the attention mechanism, an attention operation is performed on the scene content described in the text description to generate a second feature attention map; the second feature attention map is used to indicate the background characteristics of the target background;

[0017] The first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0018] In one implementation, the processing unit is configured to perform feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object, specifically to:

[0019] Acquire an object background mask image corresponding to the first image, where the mask area in the object background mask image is the area where the target object is located in the first image;

[0020] Performing feature extraction processing on the target object in the object background mask image to obtain reference object features corresponding to the target object;

[0021] Format mapping processing is performed on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object.

[0022] In one implementation, the processing unit is configured to perform an attention operation on the target object features based on the attention mechanism to generate the first feature attention map, specifically for:

[0023] Perform attention calculation on the target object features based on the attention mechanism to obtain the intermediate feature attention map;

[0024] Obtaining an object mask image corresponding to the first image;

[0025] The object mask image is used to perform feature value extraction on the intermediate feature attention map to generate the first feature attention map.

[0026] In one implementation, the processing unit is configured to fuse the first feature attention map and the second feature attention map to obtain a cross-modal attention feature, specifically configured to:

[0027] Obtain a control coefficient, which is used to indicate the fusion weight of the first feature attention map during the fusion process;

[0028] According to the fusion weight indicated by the control coefficient, the first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0029] In one implementation, the processing unit is configured to perform edge optimization processing on the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain the second image, specifically for:

[0030] Obtaining an edge feature map corresponding to the first image, where the edge feature map is used to characterize edge characteristics of a target object in the first image;

[0031] The boundary between the target object and the target background in the cross-modal attention feature is optimized according to the edge feature map to obtain a second image.

[0032] In one implementation, the processing unit, when configured to obtain the edge feature map corresponding to the first image, is specifically configured to:

[0033] Obtain the object background mask image, object mask image and object edge mask image corresponding to the first image; object edge mask image

[0034] performing splicing processing on the object background mask image, the object mask image and the object edge mask image to generate a spliced mask image;

[0035] The spliced mask image is edge enhanced to obtain an edge feature map; the edge feature map is used to indicate the edge characteristics of the target object.

[0036] In one implementation, the cross-modal fusion process and the edge optimization process are performed by calling a background replacement model; the background replacement model includes a concept injection module, an edge control module, and an attention module; the concept injection module includes an attention submodule, and the attention module includes an attention submodule;

[0037] Among them, the attention submodule in the concept injection module is located in the attention module, and the attention submodule in the concept injection module is created based on the original attention submodule in the attention module; the attention submodule belonging to the concept injection module in the attention module is used to: perform attention operations on the object features of the target object based on the attention mechanism, and the original attention submodule in the attention module is used to perform attention operations on the text description based on the attention mechanism.

[0038] In one implementation, the probability injection module further includes a feature extraction submodule and a format mapping submodule;

[0039] The feature extraction submodule is used to perform feature extraction processing on the target object in the object background mask image corresponding to the first image to obtain reference object features corresponding to the target object;

[0040] The format mapping submodule is used to perform format mapping processing on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object; the feature format of the target object features meets the feature input requirements of the attention module;

[0041] The attention submodule belonging to the concept injection module in the attention module is used to perform attention operation on the target object features based on the attention mechanism to obtain the intermediate attention map.

[0042] In one implementation, the edge control module is configured to perform edge enhancement processing on a stitched mask image; the stitched mask image is obtained by stitching an object background mask image, an object mask image, and an object edge mask image corresponding to the first image in an input channel dimension;

[0043] Among them, the object background mask image is used to enhance the learning of object characteristics of the target object in the first image; the object mask image is used to enhance the learning of the edge of the target object in the first image; and the object edge mask image is used to enhance the learning of the edge of the target object.

[0044] In one implementation, the attention module includes multiple network layers having different resolution levels; the edge characteristics of the target object are represented by an edge feature map; and the processing unit is configured to perform edge optimization processing on the cross-modal attention features based on the edge characteristics of the target object in the first image to obtain the second image, specifically for:

[0045] Embed the edge feature map output by the edge control module into the output position of each network layer in multiple network layers;

[0046] Perform splicing operation on the output result of each network layer and the edge feature map to obtain the edge splicing result corresponding to each network layer;

[0047] Based on the edge splicing results corresponding to each network layer, the boundary between the target object and the target background in the cross-modal attention feature is optimized to obtain the second image.

[0048] On the other hand, an embodiment of the present application provides a computer device, the computer device comprising:

[0049] a processor for loading and executing computer programs;

[0050] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the image processing method is implemented.

[0051] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. The computer program is suitable for being loaded by a processor and executing the above-mentioned image processing method.

[0052] On the other hand, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned image processing method.

[0053] In an embodiment of the present application, the user is required to provide a first image containing a target object, and a text description is also required to be obtained, which is used to describe the scene content containing the target object and the target background. In this way, a cross-modal attention feature can be generated with the help of the scene content described by the text description and the target object in the first image. The cross-modal attention feature has both the background characteristics of the target background in the scene content and the object characteristics of the target object; by injecting the target object features of the target object extracted from the first image in the process of generating the target background and the target object based on the text description, the target object features of the target object can be combined to assist the text description in generating the target object, that is, ensuring that the generated target object can match the content (or object, such as the target object) described by the text description, thereby improving the matching degree between the generated target object and the target object in the text description. The embodiment of the present application also supports edge optimization processing of the cross-modal attention feature based on the edge characteristics of the target object in the first image to generate a second image; by optimizing the boundary edge between the foreground (i.e., the target object) and the background (i.e., the target background) in the cross-modal attention feature, it is ensured that the edge of the generated target object is not deformed, thereby ensuring the naturalness of the fusion between the foreground and the background, and improving the image quality of the generated second image. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0055] Figure 1 is a schematic diagram of a foreground image and a background image corresponding to an image provided by an exemplary embodiment of the present application;

[0056] Figure 2 is a schematic diagram of the architecture of an image processing system provided by an exemplary embodiment of the present application;

[0057] Figure 3 is a flowchart of an image processing method provided by an exemplary embodiment of the present application;

[0058] Figure 4 This is a schematic diagram of an interface for uploading a first image and inputting a text description in a service interface of an application provided by an exemplary embodiment of the present application;

[0059] Figure 5 is a flowchart of a computer device performing cross-module fusion processing on a first image and a text description, provided by an exemplary embodiment of the present application;

[0060] Figure 6 This is a schematic diagram of a model framework of a background replacement model provided by an exemplary embodiment of the present application;

[0061] Figure 7 This is a schematic diagram of an exemplary embodiment of the present application, which provides a method of segmenting a first image using a segmentation model to obtain an object background mask image, an object mask image, and an object edge mask image corresponding to the first image;

[0062] Figure 8 1 is a schematic diagram of a fusion of a second feature attention map corresponding to a text description and a first feature attention map corresponding to a first image, provided by an exemplary embodiment of the present application;

[0063] Figure 9 is a flowchart of another image processing method provided by an exemplary embodiment of the present application;

[0064] Figure 10 1 is a schematic diagram of a network layer for injecting an edge feature map into a decoder side of a Unet network, provided by an exemplary embodiment of the present application;

[0065] Figure 11is a structural diagram of an image processing device provided by an exemplary embodiment of the present application;

[0066] Figure 12 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0067] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0068] In the embodiments of the present application, an image processing solution is proposed, specifically a solution for realizing image generation based on a cultural graph model. The following is a brief introduction to the technical terms and related concepts involved in the image processing solution provided in the embodiments of the present application, including:

[0069] 1. Artificial Intelligence (AI)

[0070] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. In other words, AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0071] As mentioned above, the image processing solutions provided in the embodiments of this application involve large models and machine learning in the field of artificial intelligence. The following is a brief introduction to the relevant contents of large models and machine learning, including:

[0072] (1) The large model is specifically a large text-graph model. The large text-graph model is also called a text-graph model, a text-graph diffusion model, etc. It is specifically an application based on a diffusion model (Diffusion Model), which supports inputting text to control the generation of pictures (or images) during the pre-training diffusion model. Among them, there are many types of open source text-graph models, which can include but are not limited to: SD (Stable Diffusion) Inpainting model and ControlNet Inpainting model. Among them: ① The SD Inpainting model diffuses in the latent space rather than the pixel space, and combines the text semantic feedback from the Transformer to generate images; its model structure mainly consists of three parts, namely the text encoder TextEncoder, the image encoder VAE Encoder and the denoising model Unet. During pre-training of the SD Inpainting model, the VAE Encoder maps the image to a latent space and then adds noise to the acquired image latent code. The Unet denoising model continuously performs denoising based on the input noisy image latent code and the text features encoded by the Text Encoder. The VAE Decoder then restores the original image (or images). The ControlNet Inpainting model uses ControlNet as a plugin for the text-based graph model. Based on a provided conditional control graph, it injects control condition features into the Unet network, thereby controlling the output of the text-based graph model to follow the conditional graph format.

[0073] (2) Machine learning is a multidisciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithmic complexity theory, and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.

[0074] The image processing solution provided in the embodiment of the present application specifically relates to multimodal learning (MMML) and cross-modal learning in the field of machine learning. The modality here, or modal information, can be understood as a form of data or information, specifically a way of expressing or perceiving things; modal information may include but is not limited to video, image, text, voiceprint, point cloud, etc. Among them: ① Multimodal learning refers to the simultaneous use of multiple modal information for machine learning to ensure the ability to learn models from multiple modalities (such as images, speech, and text, etc.); for example, multimodal learning such as image-to-text (i.e., generating text from images) and text-to-image (i.e., generating images from text). The key to multimodal learning is to integrate and analyze different modal information to obtain a more comprehensive and in-depth source of information than a single modal information, which can help to better understand complex problems in the real world and improve the accuracy and generalization ability of the model. Cross-modal learning is a learning process that transfers and understands information between different modalities. It specifically involves extracting information from one modality (such as text) and using that information to understand or enhance the content of another modality (such as images). The core of cross-modal learning lies in exploring and leveraging the correlations and complementarities between different modalities, so that information extracted from one modality can enhance understanding of the content of another modality.

[0075] 2. Image.

[0076] An image can be a representation of visual information obtained by observing the objective world in different forms (or means) using various observation devices (or systems). In the computer field, an image can also refer to a bitmap image or other image format input through an input device (such as a scanner or digital camera). The embodiments of this application do not limit the source of the image.

[0077] An image often contains one or more image elements, and an image element refers to an object or content in the image; for example, an image is obtained by photographing a physical environment with a camera, so the image obtained contains objects in the physical environment that are under the camera lens, that is, the image elements contained in the image are objects in the physical environment that are under the camera lens. In order to facilitate understanding of the importance of the many image elements in an image, the image elements that are considered important in the image can often be called the main content (or simply the main body), and the image separated from the image and containing the main content is called the foreground image; conversely, the other image elements in the image except the main content are called background content, and the image separated from the image and containing the background content is called the background image. An exemplary schematic diagram of an image containing main content and background content can be seen in Figure 1 ;like Figure 1As shown, assuming that the image element "product 102" in image 101 is considered to be an important element, image segmentation of image 102 can obtain a foreground image 103 containing the main content "product 102", and the image obtained by stripping the main content from image 101 is called a background image 104.

[0078] In practical applications, cross-modal learning is supported for model training of the aforementioned text-based graph model, so that the trained text-based graph model can draw the input text to generate an image. Specifically, during the pre-training diffusion model stage, the text feature information of the input text (such as text converted from a sampled speech signal (such as a signal sampled during a user's speech)) is injected into the cross-modal attention layer of the denoising model to extract information from the text modality and use the extracted information to generate an image in the image modality. In detail, the process of drawing the input text to generate an image using the diffusion model can include two main processes: forward diffusion and backward denoising. In the forward diffusion stage, the text is gradually contaminated by noise until it becomes completely random noise. In the backward process, a series of Markov chains are used to gradually remove the prediction noise at each time step (or timestamp, such as 1 second), thereby recovering the data from the Gaussian noise to generate an image.

[0079] In practice, it has been found that using traditional open-source text-based image models to replace the background of a specified object based on input text can lead to problems such as conceptual errors and poor edge control, which seriously degrades the image quality of the newly generated image. For example, when using the open-source SD Inpainting model and ControlNet Inpainting model to process input text to generate a background for a specified object, there is a low degree of fit between the main content of the newly generated image (i.e., the specified object) and the main content described in the text (or a low matching degree, i.e., the main content in the generated image and the main content described in the text are poorly similar). There is also a problem of unnatural connection between the main content and background content in the generated image (such as deformation of the main content's edges).

[0080] Based on this, the embodiment of the present application combines cross-modal learning and multimodal learning to propose an image processing solution based on a text graph model. The solution can specifically introduce a concept injection module and an edge control module into the text graph model to generate different backgrounds based on text descriptions for different target objects (such as products in the e-commerce field) to obtain a new image containing the target object and the background described by the text description. Among them: ① The concept injection module has the ability of concept recognition; concept recognition can mean that when a text graph model is given a text description (such as the text mentioned above) for describing the main content and background content, the text graph model can generate the main content and background content described by the text description, specifically, the objects generated by the text graph model and the object description words in the text description (such as the words used to describe the object in the text description, such as "water bottle", "big tree", etc.) can match (i.e., have a high similarity). ② The edge control module has the ability of edge control, specifically, it assists the text graph model in controlling the edges of the main content (such as the outline edge of the object, i.e., the boundary between the object and the background) when generating the background for the main content. It prevents deformation and distortion, so as to ensure the natural connection between the main content and the background content. In this way, the text-based graph model with the concept injection module and the edge control module only needs to receive a first image containing a target object (such as the main content) provided by the user, and can automatically replace the background in the first image according to the text description (that is, the description of the scene content of the new image to be generated), thereby generating a new image containing the target object and the new background, and can ensure that the details of the target object in the new image are consistent with the details of the target object in the first image, and the edges of the target object will not be deformed or distorted.

[0081] For the sake of distinction, the embodiment of the present application refers to the text graph model embedded with the concept injection module and the edge control module as a background replacement model. Specifically, after the background replacement model is trained, it can be deployed in a computer device; in this way, when the user has a need to generate a background for the main content (such as a specified product) to obtain a new image, the trained background replacement model deployed in the computer device can be called to quickly replace the background. Among them, the general process of calling the trained background replacement model deployed in the computer device to implement the image processing solution provided by the embodiment of the present application may include: obtaining a first image provided by the user, the first image containing a target object as the main content of the new image to be generated (which may be called a second image in the embodiment of the present application); and also obtaining a text description, the text description is used to describe the scene content containing the target object and the target background. In layman's terms, the text description describes the image content contained in the new image to be generated. Both the first image and the text description are input as input information into the trained background replacement model. By adopting a multimodal learning approach (i.e., learning the text modality and the image modality simultaneously), richer information about the main content can be extracted from the image dimension and the text dimension. In this way, the background replacement model can perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image (such as cross-modal fusion between the text description in the text modality and the first image in the image modality). Specifically, the text features for the text description and the object features for the target object are fused to generate a cross-modal attention feature. It can be seen that the cross-modal attention feature is a rich feature learned from the multimodal perspective of the text modality and the image modality. The cross-modal attention feature combines the object characteristics of the target object and the background characteristics of the target background. The background replacement model also analyzes the edge characteristics of the target object in the first image and optimizes the cross-modal attention feature based on the edge characteristics of the target object to control the edge of the target object in the cross-modal attention feature from deformation and other problems, thereby obtaining a second image. The second image contains the target object in the first image and the scene content described by the text description.

[0082] As can be seen from the above description, on the one hand, the embodiment of the present application only requires the user to provide a first image containing a target object, and then the trained background replacement module can be called to use the text description to quickly generate a second image containing the user-specified target object and a new background (i.e., the target background mentioned above); not only can multimodal learning be used, specifically, the text-based graph model is learned based on both text modality and image modality to extract richer features, which is conducive to improving the performance and generalization ability of the model, but also cross-modal learning can be used, specifically, information is extracted from the text description input by the user (i.e., text modality), and the extracted information is used to generate a second image (i.e., image modality), fully identifying and utilizing the intrinsic connection between different modal information; for users, it avoids users from learning new operations while greatly improving the efficiency of generating new backgrounds for target objects. On the other hand, the embodiments of the present application support generating the scene content (such as the target object and the target background) described by the text description based on the text description, specifically generating the target object and the target background based on the text description, combining the target object features of the target object in the first image to guide the model to accurately match the target object described by the text description, which can ensure the consistency of the details of the target object generated by the model and the details of the target object described by the text description, ensure the high quality of the newly generated main content "target object", and avoid conceptual cognitive errors. On the other hand, the embodiments of the present application also support the model to generate cross-modal attention features based on the text description and the first image, and the auxiliary model strictly controls the edges of the target object to prevent deformation, distortion, and expansion, ensure the natural fusion of the target object and the target background, and ensure the image quality of the second image while generating a new background (i.e., the target background) for the target object.

[0083] The image processing solution provided in the embodiment of the present application is a solution that can automatically generate a corresponding new image based on a given target object and text description, so that the image processing solution provided in the embodiment of the present application can be applied to any application scenario that requires the generation of specified scene content. Application scenarios may include but are not limited to: ① Background replacement scenario: The image processing solution can replace the new background (i.e., target background) for the target object in the first image given by the user based on the text description, and obtain a second image containing the target background and the target object, so as to achieve the replacement of the new background for the target object in the first image. ② Image generation scenario: The image processing solution can generate a second image containing the target object based on the target object given by the user (specifically, it can be a first image containing only the target object, such as the background content of the first image is blank) and the text description.

[0084] Furthermore, the application scenarios described above may belong to different application fields, including but not limited to: e-commerce field, advertising field and model training field, etc., which are not limited in the embodiments of the present application. For example, the background replacement scenario belongs to the e-commerce field, and in the e-commerce field, there is often a need to automatically replace the product background for the same product to obtain product promotion images (or product posters) containing different backgrounds. In this way, by using the image processing solution provided by the embodiment of the present application, only text descriptions describing different product backgrounds need to be provided, and the background of the product image can be replaced based on different text descriptions to obtain a second image containing different product backgrounds and the same product, so as to place the same product in the desired scene (i.e., the target background); accelerate the entire creation process of product promotion images, and greatly improve the efficiency of background replacement for products in the e-commerce field. Another example is the background replacement scenario in the advertising field, where different backgrounds may be needed for the same advertising content. In this case, the advertiser only needs to provide a first image containing the advertising content and a text description of the desired background. Multimodal learning can then be used to extract richer features from the text description in the text modality and the first image in the image modality. Specifically, cross-modal fusion processing is performed on the advertising background (i.e., scene content) described in the text description in the text modality and the advertising content (e.g., the product to be advertised) in the first image in the image modality (i.e., cross-modal fusion of background features in the text modality and features of the advertising content in the image modality) to generate a cross-modal attention feature. This cross-modal attention feature is a fused feature learned from the text description in the text modality and the first image in the image modality, and can be used to indicate the content characteristics of the advertising content in the first image and the background characteristics of the advertising background in the text description. It also supports learning edge characteristics of the advertising content in the first image (e.g., edge characteristics of the product) and using these edge characteristics of the advertising content to perform edge optimization processing on the cross-modal attention feature to generate a second image containing the advertising content in the first image and a new advertising background. This shows that it supports the generation of a second image in an image modality based on a text description in a text modality, so as to fully utilize and identify the intrinsic connection between the text modality and the image modality through this cross-modal learning method. For example, the background replacement scenario belongs to the field of model training. In the field of model training where the sample data is an image, a large number of sample images are often required. In this case, the embodiment of the present application can be used to perform multiple background replacements on a first image, so that multiple sample images with different backgrounds can be quickly obtained, thereby enriching the training data while improving the efficiency of collecting training data.

[0085] It should be noted that the above description is merely an example of the application scenarios and application fields provided in the embodiments of this application, and does not limit the application scenarios and fields of the image processing solutions provided in the embodiments of this application; the subsequent embodiments of this application use the application scenario of background replacement as an example to illustrate. The image processing solutions provided in the embodiments of this application can provide efficient and accurate image generation services in various application scenarios and application fields related to image generation, and demonstrate high value and practicality in various application scenarios and application fields related to image generation.

[0086] To facilitate understanding of the image processing solution provided in the embodiment of the present application, the following Figure 2 The schematic diagram of the scenario shown in FIG. 1 briefly introduces the application scenario involved in the embodiment of the present application; Figure 2 As shown, the image generation system includes a terminal 201 and a server 202. The embodiment of the present application does not limit the number and naming of the terminal 201 and the server 202.

[0087] Among them, terminal 201 may refer to a terminal device held by a user for inputting a first image. The terminal device may include but is not limited to: smart phones (such as smart phones deploying the Android system, or smart phones deploying the Internetworking Operating System (IOS)), tablet computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, MID), vehicle-mounted devices, head-mounted devices, smart chat robots and aircraft and other smart devices. The embodiments of this application do not limit the type of terminal devices, which are explained here. Server 202 is a server corresponding to terminal 201, which is used to interact with terminal 201 for data to provide computing and application service support for terminal 201. Server 202 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 201 and the server 202 may be connected directly or indirectly via wired or wireless communication, which is not limited in this application.

[0088] The image processing solution provided by the embodiment of the present application can be executed by a computer device, which is equipped with a background replacement model trained by the embodiment of the present application; thus, in the model inference stage, the computer device can be used to call the background replacement model to replace the background of the target object in the first image. The computer device can be Figure 2 The terminal or server in the system shown, that is, the embodiment of the present application supports the execution of the image processing solution by either the terminal and the server, or the terminal and the server together.

[0089] Take the example of the terminal and server jointly executing the image processing solution, such as Figure 2 As shown, it is assumed that a trained background replacement model is deployed in the server 202, and an application with a background replacement function (such as an application in the form of a client, a mini-program, or a web page) is installed in the terminal 201. For the user, when there is a need to replace the background of the target object in the first image (or the need to generate a new background for the target object), the user can open the application with a background replacement function through the terminal 201 and input a specified first image containing the target object in the application. The terminal 201 (specifically, the application running in the terminal 201) can upload the first image to the server 202. After receiving the first image provided by the user, the server 202 can perform the following steps: ① Preprocess the first image to obtain an object background mask image 203, an object mask image 204, and an object edge mask image 205 corresponding to the first image; specifically, extract the object background mask image 203 and the object mask image 204 corresponding to the first image through a segmentation model, and calculate the canny edge of the boundary of the target object based on the object background mask image 203 and the object mask image 204 to generate an object edge mask image 205. ② Input the object background mask image 203, object mask image 204, and object edge mask image 205 corresponding to the first image, as well as the text description (from the user or the text description stored in the server) into the background replacement model deployed in the server 202, so as to call the background replacement model to automatically replace the target object background in the first image based on the object background mask image 203, object mask image 204, and object edge mask image 205 corresponding to the first image, and output a second image that combines the target object in the first image and the target background described by the text description. In this way, the server 202 can return the second image with the replaced background to the terminal 201, so that the terminal 201 can display the second image with the replaced background to the user.

[0090] It should be noted that: ① Figure 2The diagram shown is merely a schematic diagram of the architecture of an exemplary image processing system provided by an embodiment of the present application; in actual applications, the architecture can be adaptively modified. For example, when a trained background replacement model is deployed in a terminal 201 held by a user, the image processing solution provided by an embodiment of the present application can be executed by the terminal 201, that is, the aforementioned execution subject computer device is the terminal 201; in this implementation, after receiving the first image uploaded by the user and the determined text description, the terminal 201 can directly call the background replacement model to generate a new background for the target object in the first image based on the text description, without having to send the first image and text description to the server for related processing. For example: In order to improve the resource utilization of terminal 201 (such as the utilization of CPU or GPU) and reduce the workload of server 202, the embodiment of the present application also supports the terminal 201 to perform preprocessing operations on the first image (that is, the process of generating object background mask image 203, object mask image 204 and object edge mask image 205); under this implementation method, terminal 201 can send the object background mask image 203, object mask image 204 and object edge mask image 205 corresponding to the first image to server 202. For server 202, there is no need to perform preprocessing on the first image, which effectively saves the resources of server 202 and avoids problems such as lag of server 202 due to resource occupancy.

[0091] ② It can be understood that an image is the basic unit of a video, and a video is a multimedia form presented by playing a plurality of images in chronological order, and each image constituting a video can be called a video frame. The first image to be processed involved in the embodiment of the present application can be any video frame of a video frame. Then, when the number of the first images to be processed is a plurality of video frames in a video, the image processing scheme provided by the embodiment of the present application can be used to replace the background of a plurality of video frames; thereby, in certain application scenarios where a large number of backgrounds need to be replaced (such as filming a film or TV series to avoid the high cost of reshooting), the effect of quickly changing the background of a large number of video frames can be achieved (such as the image processing scheme provided by the embodiment of the present application can complete the function of replacing a background in an average of 3 seconds through practice), which not only improves the efficiency of background replacement or replacement, but also reduces the cost of background replacement.

[0092] ③ The collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information requires the knowledge or consent of the individual subject (or the legal basis for obtaining the information), and subsequent data use and processing shall be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, when obtaining the first image, it is necessary to obtain the permission or consent of the user holding the first image, and the collection, use and processing of relevant data (such as the collection and release of the barrage posted by the object, etc.) shall comply with the relevant laws, regulations and standards of the relevant region.

[0093] Based on the image processing scheme described above, the embodiment of the present application proposes a more detailed image processing method. The image processing method proposed in the embodiment of the present application will be described in detail below with reference to the accompanying drawings.

[0094] Figure 3 A flowchart of an image processing method provided by an exemplary embodiment of the present application is shown; the image processing method may be executed by a computer device in the aforementioned system, such as a server; the image processing method may include but is not limited to steps S301-S303:

[0095] S301: Acquire a text description to be processed and a first image.

[0096] Among them, the first image is an image provided by the user containing the target object. For example, in the e-commerce field, the first image provided by the e-commerce company may be an image containing a certain commodity (such as a certain beverage bottle); the commodity serves as the main content of the first image, and the embodiment of the present application does not limit the background content in the first image, such as the background content can be blank or non-blank. In a specific implementation, when the user has a need to replace the background (or background change) of the first image, the user can use the terminal to open an application with a background replacement function, and input the first image through the upload entrance provided by the application, and then determine that the terminal obtains the first image provided by the user. The embodiment of the present application does not limit the image source of the first image uploaded by the user; for example, the first image can be any image selected by the user from the local storage space of the terminal; for another example, the first image is an image downloaded directly from the Internet by the user; for another example, the first image is an image copied or downloaded by the user from other applications deployed in the terminal (such as chat images in social applications, music covers in music applications, etc.); and so on.

[0097] The text description is a text that describes the scene content represented by the new image to be generated (i.e., the second image), and is expressed in the form of a character string. Considering that the embodiment of the present application wants to change the background of the target object in the first image, the text description can be understood as a sentence that describes the scene content including the target object and the target background. For example, the target object in the first image is a "water bottle", and the user wants to place the water bottle on the water surface, that is, the new target background is the water surface, then the scene content including the text description of the water bottle placed on the water surface can be expressed as: a bottle of water placed on the water surface; the object description words in the text description include "a bottle of water" and "water surface". In the embodiment of the present application, the text description is mainly used to generate the background content of the new image; that is, all the background content of the new image is generated based on the text description; the way of creating the background through text description is not only convenient for the image designer to express the desired target background using text language, but also can provide the image designer with more background design ideas.

[0098] Similar to the acquisition method of the first image described above, the text description can be provided by the user according to his or her own background generation needs. Optionally, the text description can be directly edited and input by the user in an application with a background replacement function deployed in the terminal; in this way, the user can input the text description in a personalized way according to his or her own background replacement needs, which meets the user's need to customize a new background and improves the user experience. Optionally, the text description can also be obtained by converting the voice signal collected by the terminal; for example, a voice collection and conversion function is set in the application. When the voice collection and conversion function in the application is turned on, the microphone of the terminal can be turned on to collect the human voice in the physical environment of the user, and the voice signal corresponding to the collected human voice is converted into text form, thereby obtaining a text description in text form; in this way, the ease of inputting text descriptions is improved in certain scenarios (such as directly inputting text descriptions by voice in driving scenarios), and the user's input experience is improved,

[0099] For example, the user uploads the first image and enters the text description in the service interface provided by the application program. Figure 4 .like Figure 4As shown, the service interface 401 provided by the application includes an image upload area 4011 for realizing image upload; in response to a triggering operation on any display position in the image upload area 4011, indicating that the user has a need to upload a first image, an image source window 4012 can be output, and the image source window 4012 includes entries from different sources (such as an entry from a local source, an entry from an Internet source, and an entry from a third-party application source, etc.); the user can perform a selection operation on the entry of the corresponding image source in the image source window 4012 according to his or her own image source needs, so that the user can jump from the service interface 401 to the corresponding image display interface (such as a search interface provided by an Internet application, and then a folder interface corresponding to a local storage space, etc.) to select the first image. Similarly, the service interface 401 also includes a text input area 402. The user can trigger the input area 402 to call up a virtual keyboard and use the virtual keyboard to input text descriptions.

[0100] It should be noted that: ① The embodiment of the present application does not limit the style and display position of the page elements contained in the service interface provided by the application for uploading the first image and inputting the text description. For example, if the text description is converted through the voice collection and conversion function described above, then the service interface 401 needs to include a voice collection option (or button, component), etc. For another example, the application can provide one or more candidate text descriptions in the service interface, so that the user can select a candidate text description from one or more candidate text descriptions as the text description of the model input.

[0101] ②The above Figure 4 The example in which the interfaces for uploading the first image and inputting the text description are the same service interface is used for introduction; in actual applications, the interfaces for uploading the first image and inputting the text description can also be different interfaces, such as the application provides an upload interface for uploading the first image, and provides an input interface independent of the upload interface for inputting the text description.

[0102] ③ In addition to being provided by the user, the text description can also be a default one. In other words, the computer device can store one or more text descriptions by default. Then, for any first image received, the background in the first image is replaced with the target background described by the default stored text description (such as replaced with the target background described by a specified text description, or replaced with the target background described by a randomly selected text description from multiple text descriptions). This method of replacing the background with a specified text description can save the user from entering a text description in some specific scenarios (such as promoting a certain background, or when the user wants to experience a background blind box), enrich the input diversity of the text description, and increase the fun of background replacement.

[0103] S302: Perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature.

[0104] After obtaining the text description and the first image provided by the user, the computer device can analyze the target object in the first image and the scene content described by the text description respectively, so as to obtain the detailed information of the target object from the first image and the scene information of the scene content from the text description; and then perform cross-modal fusion processing on the detailed information of the target object and the scene information of the scene content, so as to generate a cross-modal attention feature with the object characteristics of the target object and the background characteristics of the target background. This method of analyzing relevant information from the image dimension (or image modality) and the text dimension (or text modality) respectively can obtain more comprehensive, rich and accurate information, thereby improving the image quality of the generated cross-modal attention feature when performing cross-modal fusion processing based on comprehensive information.

[0105] The following combination Figure 5 A specific implementation process of a computer device performing cross-modal fusion processing on a first image and a text description is introduced in detail. The implementation process may include but is not limited to steps s11 to s14; wherein:

[0106] s11: Perform feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object. In a specific implementation, in order to facilitate accurate positioning of the target object from the first image, the embodiment of the present application supports obtaining the object background mask image corresponding to the first image for feature extraction processing. Among them, the mask image (which can be simply referred to as mask, mask, mask, etc.) is generally used to specify a certain area in the image, which is the area of interest, so that certain specific operations can be performed on the area without affecting other parts of the image. The object background mask image corresponding to the first image is obtained by the computer device pre-processing the first image after acquiring the first image. Exemplarily, the computer device can specifically perform mask processing on the first image by using a segmentation model to mask the area other than the target object in the first image, and determine the display area of the target object in the first image as the area of interest, thereby obtaining the object background mask image corresponding to the first image. After the computer device acquires the object background mask image corresponding to the first image, it can perform feature extraction processing on the object background mask image corresponding to the first image to obtain the reference object features corresponding to the target object. As mentioned above, the area other than the target object in the object background mask image is masked, which makes it possible to directly perform feature extraction processing on the target object in the background mask image when performing feature extraction processing on the object background mask image, thereby extracting more comprehensive and detailed reference object features of the target object.

[0107] The cross-modal fusion processing mentioned above in the embodiment of the present application specifically includes feature extraction processing for the target object in the object background mask image, which is performed by the computer device calling the trained background replacement model. Among them, the model framework diagram of the background replacement model provided in the embodiment of the present application can be found in Figure 6 ;like Figure 6 As shown in Figure 2, the background replacement model includes a concept injection module and an attention module.

[0108] 1) The attention module belongs to the Unet network (or module, which is a denoising model). The Unet module is an encoder-decoder network structure based on a convolutional neural network (specifically a U-shaped network structure); in the diffusion model, it can be used to implement feature extraction and feature fusion to complete the denoising of the hidden code of the noisy image. In detail, the Unet module includes a cross-modal attention layer, which is mainly used to receive the text features obtained by feature extraction for the text description. In this way, the cross-modal attention mechanism of the cross-modal attention layer (or simply cross-modal attention) can be used to perform a linear transformation on the text features corresponding to the text description to obtain the second feature attention map corresponding to the text description (i.e. Figure 6 The feature maps shown are shown in Figure 2. Among them, cross-modal attention can use the attention mechanism to help the model focus on important information in the modal information during cross-modal learning; for example, when processing the input text description, the attention mechanism can be used to learn the key information in the text, so that the key information in the text description under the text modality can guide the generation of the feature attention map of the image modality (i.e., the second feature attention map mentioned above). In more detail, the cross-modal attention layer in the Unet network can be further divided into three linear layers, namely the first linear layer Key (key value), the second linear layer Value (value) and the third linear layer Query (query). The output obtained by linearly transforming the text features corresponding to the input text description in the first linear layer Key is represented as K; the output obtained by linearly transforming the text features corresponding to the input text description in the second linear layer Value (value) is represented as V; the output obtained by linearly transforming the text features corresponding to the input text description in the third linear layer Query (query) is represented as Q.

[0109] 2) The concept injection module is composed of Figure 6The framework shown is composed of a feature extraction submodule, a format mapping submodule and an attention submodule. Among them: ① The feature extraction submodule is a module that performs feature extraction processing on images; the embodiment of the present application does not limit the model type of the feature extraction submodule, such as the feature extraction submodule can be an image encoder in a CLIP (Contrastive Language-Image Pre-Training) model that does not perform gradient update. Among them, the main structure of the CLIP model is an image encoder and a text encoder; the image encoder can be used to encode the image (i.e., feature extraction processing) to obtain a vector representation of the image, and similarly, the text encoder can be used to encode the text (i.e., feature extraction processing to obtain a vector representation of the text. The embodiment of the present application really utilizes the image encoder in the CLIP model to have the ability to encode images, and utilizes the image encoder in the CLIP model to implement feature extraction processing of the object background mask image corresponding to the first image.

[0110] ②The format mapping submodule is a linear network, such as Figure 6As shown in the model framework, the format mapping submodule can be named Mapper network; it can be used to realize feature mapping, specifically mapping the features extracted from the feature extraction submodule (such as the image encoder described above) to a more suitable feature space. ③ The attention submodule in the concept injection module is located in the attention module, and the attention submodule in the concept injection module is created based on the original linear layer in the attention module; specifically, the original first linear layer Key and second linear layer Value in the attention module are copied, and the parameters of the copied linear layer Key and linear layer Value are updated during the model training process. In short, in order to take advantage of the attention mechanism to focus on the content of interest in the image, the embodiment of the present application supports the creation of linear layer Key and linear layer Value for attention operation on the image in the cross-modal attention layer in the Unet model, which can be used to help the Unet model understand the injected features about the first image; wherein, the creation of linear layer Key and linear layer Value in the cross-modal attention layer in the Unet model is specifically the copying of the linear layer Key and linear layer Value originally used for text processing in the cross-modal attention layer. Therefore, the Unet model provided in the embodiment of the present application not only includes the linear layer of its original cross-modal attention layer but also includes the attention sub-module (specifically the linear layer Key and the linear layer Value) belonging to the concept injection module after being copied; in this way, the original cross-modal attention layer in the Unet model is used to perform attention operations on the text description based on the attention mechanism; the attention sub-module (i.e., the linear layer Key and the linear layer Value) belonging to the concept injection module in the attention module of the Unet model is used to perform attention operations on the object features of the target object in the first image based on the attention mechanism from the image dimension.

[0111] Based on the above introduction to the concept injection module in the background replacement model, an exemplary process for a computer device to call the background replacement model to extract features of a target object in an object background mask image may include: a feature extraction submodule in the concept injection module receives an object background mask image corresponding to a first image, and the feature extraction submodule is used to perform feature extraction on the target object in the object background mask image corresponding to the first image to obtain reference object features corresponding to the target object; the reference object features may include, but are not limited to, texture features, color features, contour features, and shape features of the target object. Furthermore, considering that the spatial distribution of the reference object features output by the feature extraction submodule does not meet the feature distribution requirements of the Unet denoising model, it is necessary to input the reference object features (or feature information) of the target object extracted by the feature extraction submodule into a format mapping submodule (i.e., the aforementioned linear network) so that the format mapping submodule performs format mapping on the reference object features corresponding to the target object to obtain target object features corresponding to the target object; the feature spatial distribution of the target object features meets the feature distribution requirements of the Unet denoising model.

[0112] s12: Perform attention operation on the target object features based on the attention mechanism to generate the first feature attention map.

[0113] As can be seen from the description of step s11 above, after the computer device calls the background replacement model to perform feature extraction and format mapping on the object background mask image of the input first image, the target object feature of the target object can be obtained; the target object feature can be used to characterize the object characteristics of the target object, and the feature format (or data format) of the target object feature is in accordance with Figure 6 The input data requirements of the Unet module in the model shown are as follows. In this way, the target object features of the target object can be input into the Unet module, so that the Unet module can assist in the generation of the target object in the process of generating the target object and the target background based on the scene content described by the text description, combining the target object features of the target object extracted from the first image, thereby improving the accuracy of conceptual cognition (i.e., improving the similarity between the target object generated by the model and the target object generated by the object description word (or the aforementioned object description word) about the target object in the text description).

[0114] The process of the Unet module combining the target object features of the target object extracted from the first image to assist in generating the target object and the target background based on the scene content described by the text description may include:

[0115] ① Input the target object features of the target object (i.e., the output of the format mapping submodule) into the attention submodule (i.e., the linear layer Key and linear layer Value) of the concept injection module in the Unet module, so that the attention submodule of the concept injection module can perform attention calculation on the target object features of the target object based on the attention mechanism to obtain an intermediate feature attention map; the intermediate feature attention map is used to characterize the object characteristics of the target object in the first image. Among them, the attention submodule of the concept injection module in the cross-modal attention layer of the Unet module calculates the intermediate feature attention map based on the target object features of the injected target object based on the attention mechanism. The formula is as follows:

[0116]

[0117] Where Q = XW q , X is the noisy image hidden code associated with the first image input to the Unet module, W q K is the linear layer parameter of the original cross-modal attention layer in the Unet module (the linear layer parameter does not need to be updated during the model training phase, that is, it is fixed). c =TW k , T is the output of the format mapping submodule, that is, the target object feature of the target object; W k is the linear layer parameter of the concept injection module, that is, the linear layer parameter of the K linear layer in the concept injection module. The linear layer parameter W k During the model training phase, it needs to be continuously updated (i.e., the parameters that need to be trained). c =TW v , W v is the linear layer parameter of the concept injection module, that is, the linear layer parameter of the V linear layer in the concept injection module. The linear layer parameter W V During the model training phase, it needs to be continuously updated (i.e., the parameters that need to be trained).

[0118] ② In order to achieve more accurate concept cognition, that is, the auxiliary model generates a target object that is more matched with the object description words in the text description (words used to describe the object to be generated); the embodiment of the present application also supports the use of the mask of the target object to extract the activation value (or feature value) in the mask area of the intermediate feature attention map, aiming to extract more accurate object features from the intermediate feature attention map, thereby facilitating the subsequent auxiliary model to generate a more accurate target object. In a specific implementation, the process of using the mask of the target object to extract the activation value in the mask area of the intermediate feature attention map can include: obtaining the object mask image corresponding to the first image; similar to the aforementioned process of determining the object background mask image corresponding to the first image, the object mask image corresponding to the first image is also obtained by segmenting the first image using a segmentation model. A schematic diagram of an exemplary use of a segmentation model to segment the first image can be seen. Figure 7 After acquiring the first image, the computer device can input the first image into the segmentation model, so that the segmentation model uses different masks to block different areas in the first image, and obtains the object background mask image 701 and the object mask image 702 corresponding to the first image. After acquiring the object mask image corresponding to the first image, the object mask image is used to perform feature value extraction processing on the intermediate feature attention map output by the attention submodule in the concept injection module to generate a first feature attention map. Taking mask to represent the object mask image corresponding to the first image, and the intermediate feature attention map is represented as A′ (i.e., the aforementioned formula (1)) as an example, the feature value extraction process can be expressed as:

[0119] A′*mask (2)

[0120] That is, the first feature attention map can be expressed as A′*mask.

[0121] s13: Based on the attention mechanism, an attention operation is performed on the scene content described by the text description to generate a second feature attention map.

[0122] Steps s11-s12 utilize the background replacement model to perform correlation processing on the first image to extract a first feature attention map of the target object, the main content of the first image. Step s13 primarily utilizes the Unet module within the background replacement model to process the text description to extract a second feature attention map corresponding to the text description. This second feature attention map is used to indicate the background characteristics of the target background within the scene content described by the text description; that is, the text description in this embodiment of the present application can be used to generate the background of a new image.

[0123] In a specific implementation, after obtaining the text description, the computer device can first perform feature extraction processing on the text description, such as using the text encoder in the CLIP model to perform feature extraction processing on the text description to obtain the text features corresponding to the text description. Then, the text features corresponding to the text description are injected into the Unet model in the background replacement model, so that the Unet model uses the original cross-modal attention layer to perform attention calculation on the text features, and obtains a second feature attention map (in the embodiment of the present application, A is used to represent the second feature attention map). Specifically, the three matrices KQV in the cross-modal attention layer are used to calculate the mutual dependence between the tokens in the text description. The token here can be understood as a string (including one or more characters) in the text description; wherein K and Q can be used to calculate the similarity between the current token and other tokens, and this similarity is used as a weight to perform weighted summation on V. The result of the weighted summation can be used as the token of the next attention layer to obtain the second feature attention map. Based on this, the attention mechanism layer in the Unet module allows the model to deeply feel the characteristics represented by the text description, thereby extracting more accurate background characteristics of the target background and ensuring that the extracted background characteristics are purer.

[0124] s14: Fuse the first feature attention map and the second feature attention map to obtain cross-modal attention features.

[0125] Based on the relevant processing of the first image and the text description in the aforementioned steps, after obtaining the first feature attention map corresponding to the first image and the second feature attention map corresponding to the text description, the embodiment of the present application supports the background replacement model (specifically the Unet module in the background replacement model) to perform fusion processing on the two feature attention maps; the fusion processing here may include: in the process of generating the target object and the target background based on the object description words in the text description, the target object features of the target object extracted from the first image are used to guide the generation of the target object, which not only improves the generation quality of the target object, but also ensures that the target object and the target background are naturally connected. It can be seen that in the process of generating the target background and the target object based on the second feature attention map (generated based on the text description), the first feature attention map related to the target object is embedded, which improves the quality of the generated target object while ensuring the fusion between the target object and the target background, that is, the target object can be harmoniously integrated into the target background, so that the fusion between the target object and the target background will not produce a sense of disobedience, that is, no conceptual cognitive errors will occur.

[0126] The embodiment of this application supports the use of A new To represent the cross-modal attention feature after fusion processing; the fusion processing obtains the cross-modal attention feature A newThe specific implementation process may include: obtaining a control coefficient, which is used to indicate the fusion weight of the first feature attention map during the fusion process. Then, according to the fusion weight indicated by the control coefficient, the first feature attention map and the second feature attention map are fused to obtain a cross-modal attention feature; specifically, the control coefficient and the first feature attention map are multiplied to obtain the first feature attention map after the multiplication operation, and then the first feature attention map and the second feature attention map after the multiplication operation are spliced to obtain a cross-modal attention feature, which has the background characteristics of the target background and the object characteristics of the target object. The above fusion process can be found in Figure 8 , and can be expressed by the following formula (3):

[0127] A new =A+λ(A′*mask) (3)

[0128] Among them, A is the second feature attention map obtained by cross-modal processing of text description using the Unet network. The second feature attention map can be called a cross-modal feature attention map; A′*mask is the first feature attention map that characterizes the object characteristics of the target object; λ is the control coefficient, which can be between 0.3 and 0.5.

[0129] Based on the above steps s11-s14, the embodiment of the present application supports the introduction of a concept injection module in the text graph model, specifically adding a feature extraction submodule and a format mapping submodule to realize feature extraction and format mapping of the target object in the first image. A K, V linear layer is also created in the Unet module to perform attention operation on the target object features of the target object to extract and help the model learn the injected target object features to learn richer and more accurate object characteristics of the target object. In order to improve the accuracy of concept cognition, the first feature attention map related to the target object learned by the K, V linear layer and the second feature attention map learned for the text description are also fused, so as to assist the model in accurately generating a target object that matches the object description word in the text description in the process of generating a background based on the second feature attention map corresponding to the text description, thereby improving the naturalness of the fusion between the target object and the target background in the generated new image.

[0130] S303: Perform edge optimization processing on the cross-modal attention features according to the edge characteristics of the target object in the first image to obtain a second image.

[0131] Among them, the edge of the target object can be understood as the edge of the object contour of the target object, that is, the boundary between the target object and the image background. In the traditional image generation process, the edge of the target object often suffers from problems such as deformation or distortion, and in terms of visual effect, part of the edge of the target object is blocked by the background. However, in the task of replacing / replacing the background of the target object in the first image, it is often hoped that the contour edge of the target object will not be deformed when the target object and the new background are merged, so as to ensure the naturalness of the fusion of the target object and the new background. To this end, the embodiment of the present application adopts edge control as an auxiliary model to strictly control the edge of the target object from being deformed when generating a new background, so as to optimize the overall performance of the model in accurately identifying the target object and precisely controlling its edge. In the subsequent embodiments, the relevant technical content of edge control will be introduced in detail in conjunction with the background replacement model.

[0132] In summary, the embodiment of the present application can add a concept injection module to the background replacement model to realize feature extraction and attention operation of the first image provided by the user, thereby extracting the object characteristics of the target object in the first image. At the same time, the original Unet module can also be used to analyze the text description used to describe the target object and the target background to extract the background characteristics of the scene content of the new image (i.e., the second image). In this way, the object characteristics of the target object (specifically, the first feature attention map corresponding to the first image) and the background characteristics corresponding to the text description (specifically, the second feature attention map corresponding to the text description) are fused to generate a cross-modal attention feature; the cross-modal attention feature has both the background characteristics of the target background in the scene content and the object characteristics of the target object; by injecting the target object characteristics of the target object extracted from the first image in the process of generating the second image based on the text description, the target object characteristics of the target object are combined to assist in accurately matching the content described in the text description (or object, such as the target object), thereby improving the matching degree between the generated target object and the target object in the text description. In addition, it also supports edge optimization processing of cross-modal attention features based on the edge characteristics of the target object in the first image; by optimizing the boundary edges between the foreground (i.e., target object) and the background (i.e., target background) in the cross-modal attention features, the edge of the target object is controlled not to be deformed during the generation of the second image, thereby ensuring the natural fusion between the foreground and background and improving the image quality of the second image.

[0133] above Figure 3The embodiment shown mainly introduces the relevant contents of the concept injection involved in the embodiment of the present application in detail, namely, it mainly introduces: in the background replacement model, the concept injection module is used to perform feature extraction and attention operation on the first image to obtain an intermediate feature attention map; in order to improve the accuracy of concept cognition, the object mask image corresponding to the first image is also used to extract the feature value of the intermediate feature attention map to obtain the first feature attention map; further, the second feature attention map and the first feature attention map obtained by processing the text description by the Unet module are integrated in the Unet module to obtain a cross-modal attention feature that has both the object characteristics of the target object and the background characteristics of the target background. Figure 9 The illustrated embodiment introduces the specific implementation process of "edge control", another important point involved in the image processing method provided in the embodiment of the present application.

[0134] Figure 9 A flowchart of another image processing method provided by an exemplary embodiment of the present application is shown; the image processing method can be executed by a computer device in the aforementioned system, such as a server; the image processing method may include but is not limited to steps S901-S904:

[0135] S901: Acquire a text description and a first image to be processed.

[0136] S902: Perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature.

[0137] It should be noted that the specific implementation process of the embodiment shown in steps S901-S902 is the same as the above Figure 3 The specific implementation process shown in steps S301-S302 in the illustrated embodiment is similar, and reference may be made to the relevant description shown in the aforementioned steps S301-S302, which will not be repeated here.

[0138] S903: Obtain an edge feature map corresponding to the first image.

[0139] The edge feature map corresponding to the first image can be used to characterize the edge characteristics of the target object in the first image, such as the shape of the edge of the target object. Figure 6 The background replacement model shown is used to obtain the edge feature map corresponding to the first image. Figure 6 The background replacement model provided in the embodiment of the present application also includes an edge control module (Edge-Control), which can be used to extract the edge characteristics of the target object in the first image.

[0140] The edge control module can be a ControlNet network. ControlNet is a conditional graph encoding network that can conditionally control the output of the denoising model Unet according to the conditional graph. Specifically, the ControlNet network can be used as a plug-in for the Wensheng graph model (such as the background replacement model in this application), providing some conditional control graphs, so that the output of the Wensheng graph model can be controlled in the form of the conditional graph. Computer vision (CV) is a science that studies how to make machines "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes to identify, detect, and measure targets, and further performs image processing to make the computer processing into an image that is more suitable for human observation or transmission to instruments for detection. Computer vision technology generally includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0141] In order to further enhance the recognition and edge control capabilities of the target object, the edge control module provided in the embodiment of the present application supports edge control by combining three different conditional information. Among them, the three different conditional information are: the object background mask image corresponding to the first image (that is, the image remaining after the first image is masked, in layman's terms, the image that retains the target object in the first image and removes the background), the object mask image corresponding to the first image, and the object edge mask image. Among them, the object background mask image can be used to enhance the injection of target object features of the target object, that is, to enhance the model's learning of the object characteristics of the target object in the first image; the object background mask image can provide the model with richer edge perception, helping the model to better learn the edges of the target object in the first image; the object edge mask image is used to further enhance the control of the edges of the target object, that is, to help the model enhance the learning of the edges of the target object. It can be seen that the embodiment of the present application fully combines the three mask images corresponding to the first image, which can optimize the overall performance of the background replacement model in accurately identifying objects and accurately controlling their edges, and achieve accurate control of the edges of the target object without deformation or distortion during the generation of the target background, that is, to ensure that the edges of the target object can be accurately controlled and improve the accuracy of the rendering of the target object.

[0142] In a specific implementation, after acquiring the first image, the computer device may perform segmentation processing on the first image to obtain the object background mask image, object mask image and object edge mask image corresponding to the first image. For example, the object edge mask image can refer to the aforementioned Figure 7 The mask image 703 in the schematic diagram shown in FIG. The masked area in the object edge mask image includes all areas in the first image excluding the area occupied by the edge of the target object's object contour. A computer device then performs a splicing process on the object background mask image, the object mask image, and the object edge mask image corresponding to the first image to generate a spliced mask image. Specifically, the splicing process is performed on the object background mask image, the object mask image, and the object edge mask image corresponding to the first image along the input channel dimension of the model. Finally, the spliced mask image is input into the background replacement model, specifically, into the edge control module within the background replacement model, causing the edge control module to perform edge enhancement processing on the spliced mask image to obtain an edge feature map.

[0143] S904: Perform edge optimization processing on the boundary between the target object and the target background in the cross-modal attention feature according to the edge feature map to obtain a second image.

[0144] Among them, edge optimization processing is achieved by calling the background replacement model provided by the embodiment of the present application. As described above, the embodiment of the present application embeds an edge control module in the background replacement model. This edge control module can combine three types of conditional information (i.e., the object background mask image corresponding to the first image, the object mask image, and the object edge mask image) to jointly improve the recognition ability and edge control ability for the target object, and obtain an edge feature map that contains both the edge characteristics of the target object and a small amount of object characteristics.

[0145] Furthermore, after obtaining the edge feature map based on the edge control module, the computer device can inject the edge feature map into the attention module in the background replacement model, specifically into the following Figure 6 In the Unet module shown, the Unet module performs edge optimization processing on the boundary between the target object and the target background in the cross-modal attention feature according to the edge feature map to obtain a second image. In detail, as mentioned above, the Unet module is an encoder-decoder U-shaped network structure; the encoder and decoder in the U-shaped network structure both contain multiple network layers, and the resolution levels of the multiple network layers are different (specifically reflected in the different resolutions of the input feature maps); each network layer can be composed of structures such as convolutional layers, pooling layers, convolutional connection layers, and attention mechanisms. Among them, the network layer on the encoder side is mainly used for feature extraction, and the network layer on the decoder side can be used for feature decoding.

[0146] In the embodiment of the present application, when the edge feature map output by the edge control module in the background replacement model is injected into the Unet module, the edge feature map output by the edge control module is embedded into the output position of each network layer on the decoding side of the multiple network layers, specifically, the output position of the network layer corresponding to the feature resolution on the decoder side; the schematic diagram of exemplary injection of the edge feature map output by the edge control module into multiple network layers on the decoder side can be seen in Figure 10 ,like Figure 10 As shown, the edge feature map output by the edge control module can be embedded in the output position of each network layer on the decoder side of the Unet module. In detail, the edge control module in the background replacement model provided in the embodiment of the present application can automatically generate edge feature maps of different resolutions (or called feature resolutions). Specifically, the edge feature maps of different resolutions are embedded in the output position of the network layer of the corresponding resolution on the decoder side of the Unet network to ensure that the output result of the network layer and the edge feature map can be added.

[0147] As can be seen from the above description, by injecting the edge feature map output by the Edge-Control module into each network block of different resolution levels on the decoder side, the output results of the network layers at different resolution levels are combined with the edge features represented by the edge feature map and a small amount of object features to generate a second image, which effectively enhances the decoder's perception of the concept and edge information of the target object in the Unet module. Among them, the specific process of injecting the edge feature map into each network block of different resolution levels on the decoder side can include: performing a splicing operation on the output result of each network layer and the edge feature map to obtain the edge splicing result corresponding to each network layer; then, based on the edge splicing result, performing edge optimization processing on the boundary between the target object and the target background in the cross-modal attention feature to generate a second image. In this way, through iterative refinement of time steps, the naturalness of the fusion between the target object and the target background and the accuracy of the edge control of the target object can be gradually achieved. In the above process, by embedding the edge feature map into the network layer of the decoder in the Unet module, each network layer on the decoder side can deeply feel the edge characteristics of the target object represented by the edge feature map, so that the decoder can combine the edge characteristics of the target object represented by the edge feature map in the process of generating the second image to control the edge of the target object to be clear enough and not deformed, so as to achieve the purpose of controlling the edge of the target object from being deformed; thereby ensuring that the target object and the target background in the second image finally output by the Unet module are naturally integrated, and the edge of the target object is well controlled.

[0148] In summary, on the one hand, the embodiment of the present application can add a concept injection module to the background replacement model to realize the feature extraction and attention operation of the first image provided by the user, thereby extracting the object characteristics of the target object in the first image. At the same time, the original Unet module can also be used to analyze the text description used to describe the target object and the target background to extract the background characteristics of the scene content of the new image (i.e., the second image). In this way, the object characteristics of the target object (specifically, the first feature attention map corresponding to the first image) and the background characteristics corresponding to the text description (specifically, the second feature attention map corresponding to the text description) are fused to generate a cross-modal attention feature; the cross-modal attention feature has both the background characteristics of the target background in the scene content and the object characteristics of the target object; by injecting the target object characteristics of the target object extracted from the first image in the process of generating the second image based on the text description, the target object characteristics of the target object are combined to assist in accurately matching the content described in the text description (or object, such as the target object), thereby improving the matching degree between the generated target object and the target object in the text description. On the other hand, the embodiment of the present application also supports the introduction of an edge control module in the Wensheng graph model, so that the edge control module can fully combine the three mask images corresponding to the first image, and accurately extract the edge feature map used to characterize the edge characteristics of the target object from the spliced mask image spliced by the three mask images; and inject the edge feature map into each network layer on the decoder side of the Unet module, so that each network layer can deeply learn the edge characteristics of the target object, thereby optimizing the edge of the target object in the process of generating the second image based on the aforementioned cross-modal attention features on the decoder side, ensuring that the edge of the target object can be precisely controlled.

[0149] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, accordingly, the device of the embodiment of the present application is provided below. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuit or memory) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the module or unit function.

[0150] Figure 11 A schematic diagram of the structure of an image processing device provided by an exemplary embodiment of the present application is shown; the image processing device can be used to perform Figure 3 and Figure 9 Some or all of the steps in the method embodiment shown. Figure 11, the device includes the following units:

[0151] An acquisition unit 1101 is configured to acquire a text description to be processed and a first image; the first image includes a target object, and the text description is used to describe a scene content including the target object and a target background;

[0152] Processing unit 1102 is configured to perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; the cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background;

[0153] The processing unit 1102 is configured to perform edge optimization processing on the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain a second image.

[0154] In one implementation, the processing unit 1102 is configured to perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature, specifically for:

[0155] Performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object;

[0156] Performing an attention operation on the target object features based on the attention mechanism to generate a first feature attention map; the first feature attention map is used to indicate the object characteristics of the target object; and

[0157] Based on the attention mechanism, an attention operation is performed on the scene content described in the text description to generate a second feature attention map; the second feature attention map is used to indicate the background characteristics of the target background;

[0158] The first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0159] In one implementation, the processing unit 1102 is configured to perform feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object, specifically to:

[0160] Obtaining an object background mask image corresponding to the first image, where the mask area in the object background mask image is the area where the target object is located in the first image;

[0161] Performing feature extraction processing on the target object in the object background mask image to obtain reference object features corresponding to the target object;

[0162] Format mapping processing is performed on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object.

[0163] In one implementation, the processing unit 1102 is configured to perform an attention operation on the target object features based on the attention mechanism to generate a first feature attention map, specifically for:

[0164] Perform attention calculation on the target object features based on the attention mechanism to obtain the intermediate feature attention map;

[0165] Obtaining an object mask image corresponding to the first image;

[0166] The object mask image is used to perform feature value extraction on the intermediate feature attention map to generate the first feature attention map.

[0167] In one implementation, the processing unit 1102 is configured to fuse the first feature attention map and the second feature attention map to obtain a cross-modal attention feature, specifically for:

[0168] Obtain a control coefficient, which is used to indicate the fusion weight of the first feature attention map during the fusion process;

[0169] According to the fusion weight indicated by the control coefficient, the first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0170] In one implementation, the processing unit 1102 is configured to perform edge optimization processing on the cross-modal attention feature based on the edge characteristics of the target object in the first image to obtain the second image, specifically for:

[0171] Obtaining an edge feature map corresponding to the first image, where the edge feature map is used to characterize edge characteristics of a target object in the first image;

[0172] The boundary between the target object and the target background in the cross-modal attention feature is optimized according to the edge feature map to obtain a second image.

[0173] In one implementation, the processing unit 1102, when configured to obtain the edge feature map corresponding to the first image, is specifically configured to:

[0174] Obtain the object background mask image, object mask image and object edge mask image corresponding to the first image; object edge mask image

[0175] performing splicing processing on the object background mask image, the object mask image and the object edge mask image to generate a spliced mask image;

[0176] The spliced mask image is edge enhanced to obtain an edge feature map; the edge feature map is used to indicate the edge characteristics of the target object.

[0177] In one implementation, the cross-modal fusion process and the edge optimization process are performed by calling a background replacement model; the background replacement model includes a concept injection module, an edge control module, and an attention module; the concept injection module includes an attention submodule, and the attention module includes an attention submodule;

[0178] Among them, the attention submodule in the concept injection module is located in the attention module, and the attention submodule in the concept injection module is created based on the original attention submodule in the attention module; the attention submodule belonging to the concept injection module in the attention module is used to: perform attention operations on the object features of the target object based on the attention mechanism, and the original attention submodule in the attention module is used to perform attention operations on the text description based on the attention mechanism.

[0179] In one implementation, the probability injection module further includes a feature extraction submodule and a format mapping submodule;

[0180] The feature extraction submodule is used to perform feature extraction processing on the target object in the object background mask image corresponding to the first image to obtain reference object features corresponding to the target object;

[0181] The format mapping submodule is used to perform format mapping processing on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object; the feature format of the target object features meets the feature input requirements of the attention module;

[0182] The attention submodule belonging to the concept injection module in the attention module is used to perform attention operation on the target object features based on the attention mechanism to obtain the intermediate attention map.

[0183] In one implementation, the edge control module is configured to perform edge enhancement processing on a stitched mask image; the stitched mask image is obtained by stitching an object background mask image, an object mask image, and an object edge mask image corresponding to the first image in an input channel dimension;

[0184] Among them, the object background mask image is used to enhance the learning of object characteristics of the target object in the first image; the object mask image is used to enhance the learning of the edge of the target object in the first image; and the object edge mask image is used to enhance the learning of the edge of the target object.

[0185] In one implementation, the attention module includes multiple network layers with different resolution levels; the edge characteristics of the target object are represented by an edge feature map; the processing unit 1102 is configured to perform edge optimization processing on the cross-modal attention features based on the edge characteristics of the target object in the first image, and when obtaining the second image, specifically to:

[0186] Embed the edge feature map output by the edge control module into the output position of each network layer in multiple network layers;

[0187] Perform splicing operation on the output result of each network layer and the edge feature map to obtain the edge splicing result corresponding to each network layer;

[0188] Based on the edge splicing results corresponding to each network layer, the boundary between the target object and the target background in the cross-modal attention feature is optimized to obtain the second image.

[0189] According to one embodiment of the present application, Figure 11 The various units in the image processing device shown can be individually or fully combined into one or several other units to form a whole, or one (or some) of the units can be further divided into multiple functionally smaller units to form a whole, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the image processing device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, the image processing device can be executed by running on a general-purpose computing device such as a computer that includes processing elements and storage elements such as a central processing unit (CPU), a random access storage medium (RAM), and a read-only storage medium (ROM). Figure 4 and Figure 8 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 11 The image processing apparatus shown in and the image processing method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0190] In an embodiment of the present application, the user is required to provide a first image containing a target object, and a text description is also required to be obtained, which is used to describe the scene content containing the target object and the target background. In this way, a cross-modal attention feature can be generated with the help of the scene content described by the text description and the target object in the first image. The cross-modal attention feature has both the background characteristics of the target background in the scene content and the object characteristics of the target object; by injecting the target object features of the target object extracted from the first image in the process of generating the target background and the target object based on the text description, the target object features of the target object can be combined to assist the text description in generating the target object, that is, ensuring that the generated target object can match the content (or object, such as the target object) described by the text description, thereby improving the matching degree between the generated target object and the target object in the text description. The embodiment of the present application also supports edge optimization processing of the cross-modal attention feature based on the edge characteristics of the target object in the first image to generate a second image; by optimizing the boundary edge between the foreground (i.e., the target object) and the background (i.e., the target background) in the cross-modal attention feature, it is ensured that the edge of the generated target object is not deformed, thereby ensuring the naturalness of the fusion between the foreground and the background, and improving the image quality of the generated second image.

[0191] Figure 12 FIG2 shows a schematic diagram of a computer device provided by an exemplary embodiment of the present application. Figure 12 , the computer device includes a processor 1201, a communication interface 1202 and a computer-readable storage medium 1203. The processor 1201, the communication interface 1202 and the computer-readable storage medium 1203 can be connected via a bus or other means. The communication interface 1202 is used to receive and send data. The computer-readable storage medium 1203 can be stored in the memory of the computer device. The computer-readable storage medium 1203 is used to store computer programs. The computer programs include program instructions. The processor 1201 is used to execute the program instructions stored in the computer-readable storage medium 1203. The processor 1201 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0192] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, one or more instructions suitable for being loaded and executed by the processor 1201 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.

[0193] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor 1201 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above-mentioned model training method embodiment; in a specific implementation, the one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and execute the following steps:

[0194] Obtaining a text description to be processed and a first image; the first image contains a target object, and the text description is used to describe the scene content containing the target object and the target background;

[0195] Performing cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; the cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background;

[0196] The cross-modal attention features are edge-optimized according to the edge characteristics of the target object in the first image to obtain the second image.

[0197] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and, when performing cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature, specifically perform the following steps:

[0198] Performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object;

[0199] Performing an attention operation on the target object features based on the attention mechanism to generate a first feature attention map; the first feature attention map is used to indicate the object characteristics of the target object; and

[0200] Based on the attention mechanism, an attention operation is performed on the scene content described in the text description to generate a second feature attention map; the second feature attention map is used to indicate the background characteristics of the target background;

[0201] The first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0202] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and, when performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object, specifically perform the following steps:

[0203] Acquire an object background mask image corresponding to the first image, where the mask area in the object background mask image is the area where the target object is located in the first image;

[0204] Performing feature extraction processing on the target object in the object background mask image to obtain reference object features corresponding to the target object;

[0205] Format mapping processing is performed on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object.

[0206] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and, when performing an attention operation on the target object features based on the attention mechanism to generate a first feature attention map, specifically perform the following steps:

[0207] Perform attention calculation on the target object features based on the attention mechanism to obtain the intermediate feature attention map;

[0208] Obtaining an object mask image corresponding to the first image;

[0209] The object mask image is used to perform feature value extraction on the intermediate feature attention map to generate the first feature attention map.

[0210] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and, when performing the fusion processing of the first feature attention map and the second feature attention map to obtain the cross-modal attention feature, specifically perform the following steps:

[0211] Obtain a control coefficient, which is used to indicate the fusion weight of the first feature attention map during the fusion process;

[0212] According to the fusion weight indicated by the control coefficient, the first feature attention map and the second feature attention map are fused to obtain the cross-modal attention feature.

[0213] In one implementation, one or more instructions in a computer-readable storage medium are loaded by the processor 1201 and, when performing edge optimization processing on a cross-modal attention feature based on edge characteristics of a target object in a first image to obtain a second image, specifically perform the following steps:

[0214] Obtaining an edge feature map corresponding to the first image, where the edge feature map is used to characterize edge characteristics of a target object in the first image;

[0215] The boundary between the target object and the target background in the cross-modal attention feature is optimized according to the edge feature map to obtain a second image.

[0216] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 1201 and, when executing the process of obtaining an edge feature map corresponding to the first image, specifically perform the following steps:

[0217] Obtain the object background mask image, object mask image and object edge mask image corresponding to the first image; object edge mask image

[0218] performing splicing processing on the object background mask image, the object mask image and the object edge mask image to generate a spliced mask image;

[0219] The spliced mask image is edge enhanced to obtain an edge feature map; the edge feature map is used to indicate the edge characteristics of the target object.

[0220] In one implementation, the cross-modal fusion process and the edge optimization process are performed by calling a background replacement model; the background replacement model includes a concept injection module, an edge control module, and an attention module; the concept injection module includes an attention submodule, and the attention module includes an attention submodule;

[0221] Among them, the attention submodule in the concept injection module is located in the attention module, and the attention submodule in the concept injection module is created based on the original attention submodule in the attention module; the attention submodule belonging to the concept injection module in the attention module is used to: perform attention operations on the object features of the target object based on the attention mechanism, and the original attention submodule in the attention module is used to perform attention operations on the text description based on the attention mechanism.

[0222] In one implementation, the probability injection module further includes a feature extraction submodule and a format mapping submodule;

[0223] The feature extraction submodule is used to perform feature extraction processing on the target object in the object background mask image corresponding to the first image to obtain reference object features corresponding to the target object;

[0224] The format mapping submodule is used to perform format mapping processing on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object; the feature format of the target object features meets the feature input requirements of the attention module;

[0225] The attention submodule belonging to the concept injection module in the attention module is used to perform attention operation on the target object features based on the attention mechanism to obtain the intermediate attention map.

[0226] In one implementation, the edge control module is configured to perform edge enhancement processing on a stitched mask image; the stitched mask image is obtained by stitching an object background mask image, an object mask image, and an object edge mask image corresponding to the first image in an input channel dimension;

[0227] Among them, the object background mask image is used to enhance the learning of object characteristics of the target object in the first image; the object mask image is used to enhance the learning of the edge of the target object in the first image; and the object edge mask image is used to enhance the learning of the edge of the target object.

[0228] In one implementation, the attention module includes multiple network layers with different resolution levels; the edge characteristics of the target object are represented by an edge feature map; one or more instructions in a computer-readable storage medium are loaded by the processor 1201 and executed to perform edge optimization processing on the cross-modal attention features based on the edge characteristics of the target object in the first image to obtain a second image, specifically performing the following steps:

[0229] Embed the edge feature map output by the edge control module into the output position of each network layer in multiple network layers;

[0230] Perform splicing operation on the output result of each network layer and the edge feature map to obtain the edge splicing result corresponding to each network layer;

[0231] Based on the edge splicing results corresponding to each network layer, the boundary between the target object and the target background in the cross-modal attention feature is optimized to obtain the second image.

[0232] In an embodiment of the present application, the user is required to provide a first image containing a target object, and a text description is also required to be obtained, which is used to describe the scene content containing the target object and the target background. In this way, a cross-modal attention feature can be generated with the help of the scene content described by the text description and the target object in the first image. The cross-modal attention feature has both the background characteristics of the target background in the scene content and the object characteristics of the target object; by injecting the target object features of the target object extracted from the first image in the process of generating the target background and the target object based on the text description, the target object features of the target object can be combined to assist the text description in generating the target object, that is, ensuring that the generated target object can match the content (or object, such as the target object) described by the text description, thereby improving the matching degree between the generated target object and the target object in the text description. The embodiment of the present application also supports edge optimization processing of the cross-modal attention feature based on the edge characteristics of the target object in the first image to generate a second image; by optimizing the boundary edge between the foreground (i.e., the target object) and the background (i.e., the target background) in the cross-modal attention feature, it is ensured that the edge of the generated target object is not deformed, thereby ensuring the naturalness of the fusion between the foreground and the background, and improving the image quality of the generated second image.

[0233] The present application also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described image processing method.

[0234] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0235] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD) or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0236] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any technical object of a person skilled in the art that is within the technical scope disclosed in the present application and that can be easily conceived of by a person skilled in the art is within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.

Claims

1. An image processing method, characterized in that: include: Acquire a text description to be processed and a first image; the first image contains a target object, and the text description is used to describe a scene content containing the target object and a target background; Performing cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; The cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background; The cross-modal attention feature is edge optimized according to the edge characteristics of the target object in the first image to obtain a second image; the second image contains the target object and the target background.

2. The method according to claim 1, wherein The cross-modal fusion processing of the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature includes: performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object; Performing an attention operation on the target object features based on an attention mechanism to generate a first feature attention map; the first feature attention map is used to indicate the object characteristics of the target object; and Performing an attention operation on the scene content described by the text description based on an attention mechanism to generate a second feature attention map; the second feature attention map is used to indicate background characteristics of the target background; The first feature attention map and the second feature attention map are fused to obtain a cross-modal attention feature.

3. The method according to claim 2, wherein The performing feature extraction processing on the target object in the first image to obtain target object features corresponding to the target object includes: Acquire an object background mask image corresponding to the first image, where a mask area in the object background mask image is an area where the target object in the first image is located; Performing feature extraction processing on the target object in the object background mask image to obtain reference object features corresponding to the target object; Format mapping processing is performed on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object.

4. The method according to claim 2, wherein The performing an attention operation on the target object feature based on the attention mechanism to generate a first feature attention map includes: Performing attention calculation on the target object features based on the attention mechanism to obtain an intermediate feature attention map; Obtaining an object mask image corresponding to the first image; The object mask image is used to perform feature value extraction processing on the intermediate feature attention map to generate a first feature attention map.

5. The method according to claim 2, wherein The fusing the first feature attention map and the second feature attention map to obtain a cross-modal attention feature includes: Obtaining a control coefficient, where the control coefficient is used to indicate a fusion weight of the first feature attention map during the fusion process; The first feature attention map and the second feature attention map are fused according to the fusion weight indicated by the control coefficient to obtain a cross-modal attention feature.

6. The method according to claim 1, wherein The performing edge optimization processing on the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain a second image includes: Acquire an edge feature map corresponding to the first image, where the edge feature map is used to characterize edge characteristics of the target object in the first image; The boundary between the target object and the target background in the cross-modal attention feature is optimized according to the edge feature map to obtain a second image.

7. The method according to claim 6, wherein The obtaining of an edge feature map corresponding to the first image includes: Acquire an object background mask image, an object mask image, and an object edge mask image corresponding to the first image; an object edge mask image; performing a splicing process on the object background mask image, the object mask image and the object edge mask image to generate a spliced mask image; Edge enhancement processing is performed on the spliced mask image to obtain an edge feature map; the edge feature map is used to indicate edge characteristics of the target object.

8. The method according to claim 1, wherein The cross-modal fusion process and the edge optimization process are performed by calling a background replacement model; the background replacement model includes a concept injection module, an edge control module and an attention module; the concept injection module includes an attention submodule, and the attention module includes an attention submodule; Among them, the attention submodule in the concept injection module is located in the attention module, and the attention submodule in the concept injection module is created based on the original attention submodule in the attention module; the attention submodule in the attention module belonging to the concept injection module is used to: perform attention operation on the object features of the target object based on the attention mechanism, and the original attention submodule in the attention module is used to perform attention operation on the text description based on the attention mechanism.

9. The method according to claim 8, wherein The probability injection module also includes a feature extraction submodule and a format mapping submodule; The feature extraction submodule is used to perform feature extraction processing on the target object in the object background mask image corresponding to the first image to obtain reference object features corresponding to the target object; The format mapping submodule is used to perform format mapping processing on the reference object features corresponding to the target object to obtain the target object features corresponding to the target object; the feature format of the target object features meets the feature input requirements of the attention module; The attention submodule belonging to the concept injection module in the attention module is used to perform attention operation on the target object features based on the attention mechanism to obtain an intermediate attention map.

10. The method according to claim 8, wherein The edge control module is used to perform edge enhancement processing on the stitching mask image; the stitching mask image is obtained by stitching the object background mask image, the object mask image and the object edge mask image corresponding to the first image in the input channel dimension; Among them, the object background mask image is used to enhance the learning of object characteristics of the target object in the first image; the object mask image is used to enhance the learning of the edge of the target object in the first image; and the object edge mask image is used to enhance the learning of the edge of the target object.

11. The method according to claim 8, wherein The attention module includes multiple network layers, and the multiple network layers have different resolution levels; the edge characteristics of the target object are represented by an edge feature map; and the edge optimization processing of the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain the second image includes: Embed the edge feature map output by the edge control module into the output position of each network layer in the plurality of network layers; Performing a splicing operation on the output result of each network layer and the edge feature map to obtain an edge splicing result corresponding to each network layer; Based on the edge stitching results corresponding to each of the network layers, edge optimization processing is performed on the boundary between the target object and the target background in the cross-modal attention feature to obtain the second image.

12. An image processing device, characterized in that: include: An acquisition unit, configured to acquire a text description to be processed and a first image; the first image includes a target object, and the text description is used to describe a scene content including the target object and a target background; a processing unit, configured to perform cross-modal fusion processing on the scene content described by the text description and the target object in the first image to generate a cross-modal attention feature; The cross-modal attention feature is used to indicate the object characteristics of the target object and the background characteristics of the target background; The processing unit is further used to perform edge optimization processing on the cross-modal attention feature according to the edge characteristics of the target object in the first image to obtain a second image.

13. A computer device, characterized in that: include: a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein when the computer program is executed by the processor, the image processing method according to any one of claims 1 to 11 is implemented.

14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the image processing method according to any one of claims 1 to 11.

15. A computer program product, characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the image processing method according to any one of claims 1 to 11 is implemented.