Image generation method and device, computer equipment, storage medium and program product

By acquiring and processing image and text description information, and using multimodal models to extract and update image features, the repeated calculation and resource waste of image generation and conversion in the prior art are solved, and accurate update and efficient generation of images are achieved.

CN120235975APending Publication Date: 2025-07-01CHINA TELECOM CLOUD TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510332474.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing image generation and conversion technologies independently process text prompts and image generation tasks, resulting in repeated calculations and resource waste.

Method used

By obtaining the original image and text description information for the original image, text encoder in the multimodal model extracts text features, identify the target area in the image, and update the target area based on the updated text features to generate the target image.

Benefits of technology

Accurate updates to images are achieved, feature bias is reduced and time-saving. By introducing text description information and multimodal models, the accuracy and efficiency of the image generation process are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235975A_ABST
    Figure CN120235975A_ABST
Patent Text Reader

Abstract

The invention relates to an image generation method and device, computer equipment, a storage medium and a program product. The method comprises the following steps: acquiring an original image and text description information for the original image; wherein the text description information comprises to-be-replaced object information and updated object information of the original image; performing feature extraction on the text description information through a text encoder in the multi-modal model to obtain text features; wherein the text features comprise a first text feature corresponding to the to-be-replaced object information and a second text feature corresponding to the updated object information; performing object recognition on the original image based on the first text feature to obtain a target area in the original image; wherein the target area is an area, pointed by the to-be-replaced object information, in the original image; and updating the target area in the original image based on the second text feature to obtain a target image. By adopting the method, accurate updating of the image can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to an image generation method, apparatus, computer device, storage medium, and program product. Background Art

[0002] Image conversion technology is to fuse the apparent features of the target object in the reference image into the target image, so as to generate new content and style in the image or video.

[0003] Most of the existing image generation and conversion technologies independently process text prompts and image generation tasks. Although they each have their own advantages, there are problems of repeated calculation and resource waste. Summary of the Invention

[0004] Based on this, in view of the above technical problems, it is necessary to provide an image generation method, apparatus, computer device, storage medium, and program product that can achieve precise update of images.

[0005] In a first aspect, this application provides an image generation method, which includes:

[0006] Obtain an original image and text description information for the original image; wherein, the text description information includes information about the object to be replaced and updated object information of the original image;

[0007] Extract features from the text description information through a text encoder in a multimodal model to obtain text features; wherein, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the updated object information;

[0008] Based on the first text features, perform object recognition on the original image to obtain the target area in the original image; wherein, the target area is the area in the original image pointed to by the information about the object to be replaced;

[0009] Based on the second text features, update the target area in the original image to obtain a target image.

[0010] In one embodiment, based on the first text features, performing object recognition on the original image to obtain the target area in the original image includes:

[0011] Extract multi-scale image features of the original image through a target detection model, and match the first text features with the multi-scale image features to obtain the target area in the original image.

[0012] In one embodiment, based on the second text features, updating the target area in the original image to obtain a target image includes:

[0013] Mask the area outside the target area in the original image to obtain a mask image corresponding to the original image; update the mask image based on the second text feature to obtain the target image.

[0014] In one embodiment, masking the area outside the target area in the original image to obtain a mask image corresponding to the original image includes:

[0015] Adjust the pixel values of the target area in the original image to a first value, and adjust the pixel values of the area outside the target area in the original image to a second value to obtain a mask image corresponding to the original image.

[0016] In one embodiment, updating the mask image based on the second text feature to obtain the target image includes:

[0017] Update the mask image based on the second text feature to obtain an intermediate image; smooth the edge area corresponding to the target area in the intermediate image to obtain the target image.

[0018] In one embodiment, through the image encoder in the multimodal model, extract the features of the target image to obtain the target image features corresponding to the target image; compare the target image features with the text features to obtain a similarity value; in the case where the similarity value is less than a preset threshold, output the target image.

[0019] In a second aspect, the present application also provides an image generation device, which includes:

[0020] An acquisition module, configured to acquire the original image and the text description information for the original image; wherein, the text description information includes the information of the object to be replaced and the information of the object to be updated in the original image;

[0021] An extraction module, configured to extract the features of the text description information through the text encoder in the multimodal model to obtain text features; wherein, the text features include the first text feature corresponding to the information of the object to be replaced and the second text feature corresponding to the information of the object to be updated;

[0022] An identification module, configured to identify the object in the original image based on the first text feature to obtain the target area in the original image; wherein, the target area is the area in the original image pointed to by the information of the object to be replaced;

[0023] An update module, configured to update the target area in the original image based on the second text feature to obtain the target image.

[0024] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0025] Obtain an original image and text description information for the original image; wherein, the text description information includes information about the object to be replaced and updated object information of the original image;

[0026] Extract features from the text description information through a text encoder in a multimodal model to obtain text features; wherein, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the updated object information;

[0027] Based on the first text features, perform object recognition on the original image to obtain a target region in the original image; wherein, the target region is the region in the original image pointed to by the information about the object to be replaced;

[0028] Update the target region in the original image based on the second text features to obtain a target image.

[0029] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0030] Obtain an original image and text description information for the original image; wherein, the text description information includes information about the object to be replaced and updated object information of the original image;

[0031] Extract features from the text description information through a text encoder in a multimodal model to obtain text features; wherein, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the updated object information;

[0032] Based on the first text features, perform object recognition on the original image to obtain a target region in the original image; wherein, the target region is the region in the original image pointed to by the information about the object to be replaced;

[0033] Update the target region in the original image based on the second text features to obtain a target image.

[0034] In a fifth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the following steps are implemented:

[0035] Obtain an original image and text description information for the original image; wherein, the text description information includes information about the object to be replaced and updated object information of the original image;

[0036] Extract features from the text description information through the text encoder in the multimodal model to obtain text features; among them, the text features include the first text features corresponding to the object information to be replaced and the second text features corresponding to the updated object information;

[0037] Based on the first text features, perform object recognition on the original image to obtain the target area in the original image; among them, the target area is the area in the original image pointed to by the object information to be replaced;

[0038] Based on the second text features, update the target area in the original image to obtain the target image.

[0039] The above image generation method, device, computer device, storage medium and program product provide a basis for determining the updated content of the original image by obtaining the original image and the text description information for the original image; further, the text description information can be feature-extracted through the text encoder in the multimodal model to obtain text features including the first text features corresponding to the object information to be replaced and the second text features corresponding to the updated object information, providing a solution for realizing accurate image update; furthermore, based on the first text features, object recognition can be performed on the original image to obtain the target area in the original image, and based on the second text features, the target area in the original image can be updated to obtain the target image. This solution provides content for determining the updated content of the image by introducing the text description information; moreover, by introducing the text encoder in the multimodal model, the sameness of the text features in the image generation process is ensured, reducing feature deviation and saving time; finally, by introducing the first text features and the second text features, accurate update of the original image is realized. Description of the Drawings

[0040] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0041] Figure 1 It is an application environment diagram of an image generation method provided in an embodiment of the present application;

[0042] Figure 2 It is a flowchart of an image generation provided in an embodiment of the present application;

[0043] Figure 3 It is a flowchart of obtaining the target area provided in an embodiment of the present application;

[0044] Figure 4 It is a schematic flow chart of a process for obtaining a target image provided in an embodiment of the present application;

[0045] Figure 5 It is a schematic flow chart of another image generation method provided in an embodiment of the present application;

[0046] Figure 6 It is a structural block diagram of an image generation device provided in an embodiment of the present application;

[0047] Figure 7 It is a structural block diagram of another image generation device provided in an embodiment of the present application;

[0048] Figure 8 It is an internal structure diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners

[0049] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0050] The image generation method provided in the embodiments of the present application can be applied to an Figure 1 application environment as shown. Among them, the terminal 102 communicates with the server 104 through a network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The server 104 can obtain the original image and the text description information for the original image, and extract features from the text description information through the text encoder in the multimodal model to obtain text features; further, the original image can be object-recognized and updated according to the text features to obtain the target image, and finally the target image is displayed through the terminal 102. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0051] In an exemplary embodiment, as Figure 2 shown, an image generation method is provided. Taking the method applied to the Figure 1 server 104 in it as an example for description, it can include the following steps:

[0052] S201. Obtain the original image and the text description information for the original image.

[0053] Among them, the text description information includes the information of the object to be replaced and the information of the updated object in the original image.

[0054] Exemplarily, the original image and the text description information for the original image can be input through the terminal 102; furthermore, the original image and the text description information for the original image can be saved in the data storage system to realize the preservation of the original image and the text description information for the original image; further, the server 104 can directly obtain the original image and the text description information for the original image through the data storage system and perform image analysis and other processing.

[0055] S202. Through the text encoder in the multimodal model, extract features from the text description information to obtain text features.

[0056] Among them, the text features include the first text features corresponding to the information of the object to be replaced and the second text features corresponding to the information of the updated object.

[0057] Optionally, the feature extraction of the text description information can be realized through the text encoder of the CLIP (Contrastive Language–Image Pretraining) model. Among them, the CLIP model is a multimodal pre-training model that can understand images and texts at the same time. Its core idea is to embed images and texts into the same semantic space through contrastive learning, so as to realize cross-modal understanding and matching.

[0058] Exemplarily, the text description information can be input into the text encoder in the multimodal model, and the text encoder in the multimodal model can convert the text description information into a high-dimensional feature vector, that is, text features, to better represent the semantic information of the text description information.

[0059] It should be noted that the text encoder can extract the text features of the text description information based on the Transformer architecture through multiple layers of self-attention mechanisms (Self-Attention) and feed-forward neural networks (Feed-Forward Network).

[0060] S203. Based on the first text features, perform object recognition on the original image to obtain the target area in the original image.

[0061] Among them, the target area is the area in the original image pointed to by the information of the object to be replaced.

[0062] Exemplarily, the first text feature can be input into the GroundingDINO model. The Swin Transformer in the GroundingDINO model (a computer vision model based on the Transformer architecture) detects and identifies each object in the original image, and determines the area where the object corresponding to the first text feature in the original image is located as the target area.

[0063] It should be noted that the GroundingDINO model is an object detection model based on DINO (Distillation with NoLabels), which combines the ability of grounding (grounding refers to associating text descriptions with objects in images). It can directly detect targets in images through text descriptions and is a multi-modal object detection model. That is, GroundingDINO not only relies on visual information but also combines text information, enabling more flexible and semantic object detection.

[0064] S204, based on the second text feature, update the target area in the original image to obtain the target image.

[0065] Among them, the target image can represent the image generated after the original image is updated.

[0066] Exemplarily, the isolation of the target area in the original image can be achieved based on the Stable Diffusion model; further, based on the second text feature, the updated content of the target area in the original image can be determined; furthermore, the VAE (Variational Autoencoder) can be used to update the target area in the original image to obtain the target image.

[0067] It should be noted that the Stable Diffusion model is a generative model based on the Diffusion Model, which is used to generate high-quality images. By simulating the diffusion process (gradually generating images from noise) and combining deep learning techniques, it can generate realistic and diverse images.

[0068] The above image generation method provides a basis for determining the updated content of the original image by obtaining the original image and the text description information for the original image; further, the text encoder in the multimodal model can be used to extract features from the text description information to obtain text features including the first text feature corresponding to the information of the object to be replaced and the second text feature corresponding to the information of the updated object, providing a solution for accurate image update; furthermore, based on the first text feature, object recognition can be performed on the original image to obtain the target area in the original image, and based on the second text feature, the target area in the original image can be updated to obtain the target image. This solution provides content for determining the updated content of the image by introducing text description information; moreover, by introducing the text encoder in the multimodal model, the consistency of text features in the image generation process is ensured, reducing feature deviation and saving time; finally, by introducing the first text feature and the second text feature, accurate update of the original image is achieved.

[0069] Based on the above embodiments, the embodiments of the present application further explain the above embodiment S203 in detail. Specifically, the process of obtaining the target area involved in the embodiments of the present application is as Figure 3 shown, and specifically includes the following steps:

[0070] S301, through the target detection model, perform multi-scale feature extraction on the original image to obtain the multi-scale image features of the original image.

[0071] Among them, the target detection model can detect targets in the image; optionally, it can be the GroundingDINO model.

[0072] Exemplarily, the feature extraction of the original image can be performed through the GroundingDINO model (i.e., the target detection model). For example, through methods such as the feature pyramid network, multi-scale feature maps, dilated convolution, and adaptive pooling, the multi-scale image features of the original image can be finally obtained.

[0073] S302, match the first text feature and the multi-scale image features to obtain the target area in the original image.

[0074] Exemplarily, the multi-scale image features can be matched one by one with the first text feature, and the area in the original image corresponding to the successfully matched multi-scale image features is determined as the target area in the original image, realizing accurate target positioning to adapt to the accurate detection of specified objects in complex scenes.

[0075] In the embodiments of the present application, by introducing multi-scale image features and matching the multi-scale image features with the first text feature, accurate positioning of the target area is achieved, laying a foundation for accurately realizing image update.

[0076] Based on the above embodiments, the embodiments of the present application will explain the above embodiment S204 in detail. Specifically, the process of obtaining the target image involved in the embodiments of the present application is as follows Figure 4 shown, and specifically includes the following steps:

[0077] S401, mask the area outside the target area in the original image to obtain a mask image corresponding to the original image.

[0078] Exemplarily, through the masking technique, the masking of a specific area in the original image can be achieved to obtain a mask image.

[0079] One implementable way is to adjust the pixel values of the target area in the original image to a first value, and adjust the pixel values of the area outside the target area in the original image to a second value to obtain a mask image corresponding to the original image.

[0080] Exemplarily, the pixel values of the target area in the original image can be adjusted to 1 (or 255), and the pixel values of the area outside the target area in the original image can be adjusted to 0, so as to obtain a mask image corresponding to the original image.

[0081] S402, update the mask image based on the second text feature to obtain the target image.

[0082] Exemplarily, through the Stable Diffusion model, under the guidance of the second text feature, the mask image can be updated, that is, the target area in the mask image is updated according to the second text feature, and finally the target image is obtained.

[0083] One implementable way is to update the mask image based on the second text feature to obtain an intermediate image; smooth the edge area corresponding to the target area in the intermediate image to obtain the target image.

[0084] As shown in the above example, based on the second text feature (such as "apple"), the area with pixel value 1 (or 255) in the mask image can be updated. For example, the object in the area with pixel value 1 (or 255) in the mask image is updated to "apple", that is, the intermediate image is obtained; further, through smoothing processing methods such as Gaussian filtering, bilateral filtering, guided filtering, and anisotropic diffusion, the smoothing processing of the connection between the area with pixel value 1 (or 255) and the area with pixel value 0 can be realized to avoid the appearance of rigid transition lines, and finally the target image is obtained.

[0085] In the embodiments of the present application, by introducing the mask image, an implementation solution for realizing accurate image update is provided, laying a foundation for obtaining an accurate target image.

[0086] Based on the above embodiments, the embodiments of the present application relate to the process of comparing the similarity of a target image, specifically including:

[0087] Extract features of the target image through the image encoder in the multimodal model to obtain target image features corresponding to the target image; compare the target image features with text features to obtain a similarity value; when the similarity value is less than a preset threshold, output the target image.

[0088] Among them, the preset threshold can be set in advance and can be adaptively adjusted according to the size of the updated content of the image.

[0089] Exemplarily, the target image can be input into the multimodal model, and the image encoder in the multimodal model extracts features of the target image to obtain target image features corresponding to the target image; further, the target image features can be compared with text features, and a similarity value indicating the similarity degree between the target image features and the text features can be determined according to the comparison result; and when the similarity value is less than the preset threshold, it characterizes the accuracy and consistency of the target image, so that the target image can be output.

[0090] In the embodiments of the present application, by introducing similarity comparison, a judgment basis is provided for the accuracy of the target image.

[0091] Based on the above embodiments, an alternative example of an image generation method is provided in this embodiment. As Figure 5 shown, the specific implementation process is as follows:

[0092] S501, obtain the original image and text description information for the original image.

[0093] Among them, the text description information includes information about the object to be replaced and updated object information of the original image.

[0094] S502, extract features of the text description information through the text encoder in the multimodal model to obtain text features.

[0095] Among them, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the updated object information.

[0096] S503, perform multi-scale feature extraction on the original image through the target detection model to obtain multi-scale image features of the original image.

[0097] S504, match the first text features and the multi-scale image features to obtain the target area in the original image.

[0098] Among them, the target area is the area in the original image pointed to by the information of the object to be replaced.

[0099] S505, adjust the pixel values of the target area in the original image to a first value, and adjust the pixel values of the area outside the target area in the original image to a second value, so as to obtain a mask image corresponding to the original image.

[0100] S506, update the mask image based on the second text feature to obtain an intermediate image.

[0101] S507, smooth the edge area corresponding to the target area in the intermediate image to obtain a target image.

[0102] Furthermore, through the image encoder in the multimodal model, extract features from the target image to obtain target image features corresponding to the target image; compare the target image features with the text features to obtain a similarity value; in the case where the similarity value is less than a preset threshold, output the target image.

[0103] The specific processes of the above S501 - S507 can refer to the description of the method embodiments above, and their implementation principles and technical effects are similar, so they will not be elaborated here.

[0104] It should be understood that although each step in the flowcharts involved in the above - mentioned embodiments is shown in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above - mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0105] Based on the same inventive concept, an embodiment of the present application also provides an image generation device for implementing the above - mentioned image generation method. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the following image generation devices can refer to the limitations on the image generation method in the above text, and will not be elaborated here.

[0106] In an exemplary embodiment, as Figure 6 shown, an image generation device 1 is provided, including: an acquisition module 10, an extraction module 20, an identification module 30, and an update module 40, where:

[0107] An acquisition module 10, configured to acquire an original image and text description information for the original image; wherein, the text description information includes information about an object to be replaced and information about an updated object in the original image.

[0108] An extraction module 20, configured to extract features from the text description information through a text encoder in a multimodal model to obtain text features; wherein, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the information about the updated object.

[0109] An identification module 30, configured to perform object identification on the original image based on the first text features to obtain a target region in the original image; wherein, the target region is the region in the original image pointed to by the information about the object to be replaced.

[0110] An update module 40, configured to update the target region in the original image based on the second text features to obtain a target image.

[0111] In one embodiment, the identification module 30 is specifically configured to:

[0112] Extract multi-scale image features of the original image through a target detection model to obtain multi-scale image features of the original image; match the first text features with the multi-scale image features to obtain the target region in the original image.

[0113] In one embodiment, as Figure 7 shown, the update module 40 specifically further includes:

[0114] A masking unit 41, configured to mask the regions outside the target region in the original image to obtain a masked image corresponding to the original image.

[0115] An image update unit 42, configured to update the masked image based on the second text features to obtain a target image.

[0116] In one embodiment, the masking unit 41 is specifically configured to:

[0117] Adjust the pixel values of the target region in the original image to a first value, and adjust the pixel values of the regions outside the target region in the original image to a second value to obtain a masked image corresponding to the original image.

[0118] In one embodiment, the image update unit 42 is specifically configured to:

[0119] Update the masked image based on the second text features to obtain an intermediate image; perform smoothing processing on the edge region corresponding to the target region in the intermediate image to obtain a target image.

[0120] In one embodiment, the image generation device 1 further specifically includes:

[0121] A similarity comparison module, configured to extract features of a target image through an image encoder in a multimodal model to obtain target image features corresponding to the target image; compare the target image features with text features to obtain a similarity value; and output the target image when the similarity value is less than a preset threshold.

[0122] Each module in the above image generation device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute operations corresponding to each of the above modules.

[0123] In an exemplary embodiment, a computer device is provided. The computer device can be a server, and its internal structural diagram can be as Figure 8 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store image-related data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements an image generation method.

[0124] Those skilled in the art can understand that Figure 8 the structure shown in

[0125] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0126] Obtain the original image and the text description information for the original image; wherein, the text description information includes the information of the object to be replaced and the information of the updated object in the original image;

[0127] Extract features from the text description information through the text encoder in the multimodal model to obtain text features; wherein, the text features include the first text features corresponding to the information of the object to be replaced and the second text features corresponding to the information of the updated object;

[0128] Based on the first text features, perform object recognition on the original image to obtain the target region in the original image; wherein, the target region is the region in the original image pointed to by the information of the object to be replaced;

[0129] Based on the second text features, update the target region in the original image to obtain the target image.

[0130] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0131] Through the target detection model, perform multi-scale feature extraction on the original image to obtain the multi-scale image features of the original image; match the first text features and the multi-scale image features to obtain the target region in the original image.

[0132] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0133] Mask the region outside the target region in the original image to obtain the masked image corresponding to the original image; based on the second text features, update the masked image to obtain the target image.

[0134] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0135] Adjust the pixel values of the target region in the original image to the first value, and adjust the pixel values of the region outside the target region in the original image to the second value to obtain the masked image corresponding to the original image.

[0136] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0137] Based on the second text features, update the masked image to obtain an intermediate image; smooth the edge region corresponding to the target region in the intermediate image to obtain the target image.

[0138] In one embodiment, when the processor executes the computer program, the following steps are also implemented:

[0139] Extract features from the target image through the image encoder in the multimodal model to obtain the target image features corresponding to the target image; compare the target image features with the text features to obtain a similarity value; and output the target image when the similarity value is less than a preset threshold.

[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0141] Obtain the original image and the text description information for the original image; wherein, the text description information includes the object information to be replaced and the update object information of the original image;

[0142] Extract features from the text description information through the text encoder in the multimodal model to obtain text features; wherein, the text features include the first text features corresponding to the object information to be replaced and the second text features corresponding to the update object information;

[0143] Based on the first text features, perform object recognition on the original image to obtain the target area in the original image; wherein, the target area is the area in the original image pointed to by the object information to be replaced;

[0144] Update the target area in the original image based on the second text features to obtain the target image.

[0145] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0146] Extract multi-scale features from the original image through the target detection model to obtain the multi-scale image features of the original image; match the first text features with the multi-scale image features to obtain the target area in the original image.

[0147] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0148] Mask the area outside the target area in the original image to obtain the masked image corresponding to the original image; update the masked image based on the second text features to obtain the target image.

[0149] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0150] Adjust the pixel values of the target area in the original image to a first value, and adjust the pixel values of the area outside the target area in the original image to a second value to obtain the masked image corresponding to the original image.

[0151] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0152] Update the mask image based on the second text feature to obtain an intermediate image; smooth the edge region corresponding to the target region in the intermediate image to obtain a target image.

[0153] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0154] Extract features of the target image through the image encoder in the multimodal model to obtain target image features corresponding to the target image; compare the target image features with the text features to obtain a similarity value; output the target image when the similarity value is less than a preset threshold.

[0155] In one embodiment, a computer program product is provided, including a computer program, which when executed by a processor implements the following steps:

[0156] Obtain an original image and text description information for the original image; wherein, the text description information includes information about the object to be replaced and the update object in the original image;

[0157] Extract features of the text description information through the text encoder in the multimodal model to obtain text features; wherein, the text features include first text features corresponding to the information about the object to be replaced and second text features corresponding to the update object information;

[0158] Based on the first text feature, perform object recognition on the original image to obtain the target region in the original image; wherein, the target region is the region in the original image pointed to by the information about the object to be replaced;

[0159] Update the target region in the original image based on the second text feature to obtain a target image.

[0160] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0161] Extract multi-scale features of the original image through a target detection model to obtain multi-scale image features of the original image; match the first text feature and the multi-scale image features to obtain the target region in the original image.

[0162] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0163] Mask the region outside the target region in the original image to obtain a mask image corresponding to the original image; update the mask image based on the second text feature to obtain a target image.

[0164] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0165] Adjust the pixel values of the target region in the original image to a first value, and adjust the pixel values of the region outside the target region in the original image to a second value to obtain a mask image corresponding to the original image.

[0166] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0167] Update the mask image based on the second text feature to obtain an intermediate image; smooth the edge region corresponding to the target region in the intermediate image to obtain a target image.

[0168] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:

[0169] Extract features of the target image through the image encoder in the multimodal model to obtain target image features corresponding to the target image; compare the target image features with the text features to obtain a similarity value; and output the target image when the similarity value is less than a preset threshold.

[0170] It should be noted that the information (including but not limited to device information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0171] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0172] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0173] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An image generation method, characterized in that: The method comprises: Acquire an original image and text description information for the original image; wherein the text description information includes information about an object to be replaced and information about an updated object of the original image; By using a text encoder in a multimodal model, feature extraction is performed on the text description information to obtain text features; wherein the text features include a first text feature corresponding to the object information to be replaced and a second text feature corresponding to the update object information; Based on the first text feature, object recognition is performed on the original image to obtain a target area in the original image; wherein the target area is the area in the original image pointed to by the object information to be replaced; Based on the second text feature, the target area in the original image is updated to obtain a target image.

2. The method according to claim 1, characterized in that: The performing object recognition on the original image based on the first text feature to obtain a target area in the original image includes: Performing multi-scale feature extraction on the original image through the target detection model to obtain multi-scale image features of the original image; The first text feature and the multi-scale image feature are matched to obtain a target area in the original image.

3. The method according to claim 1, characterized in that: The updating of the target area in the original image based on the second text feature to obtain the target image includes: Masking an area outside the target area in the original image to obtain a mask image corresponding to the original image; Based on the second text feature, the mask image is updated to obtain a target image.

4. The method according to claim 3, characterized in that The step of masking the area outside the target area in the original image to obtain a mask image corresponding to the original image includes: The pixel values ​​of the target area in the original image are adjusted to a first value, and the pixel values ​​of an area outside the target area in the original image are adjusted to a second value, so as to obtain a mask image corresponding to the original image.

5. The method according to claim 3, characterized in that: The updating of the mask image based on the second text feature to obtain a target image includes: Based on the second text feature, updating the mask image to obtain an intermediate image; The edge area corresponding to the target area in the intermediate image is smoothed to obtain a target image.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: By using the image encoder in the multimodal model, feature extraction is performed on the target image to obtain target image features corresponding to the target image; Performing a similarity comparison between the target image feature and the text feature to obtain a similarity value; When the similarity value is less than a preset threshold, the target image is output.

7. An image generating device, characterized in that: The device comprises: An acquisition module, used to acquire an original image and text description information for the original image; wherein the text description information includes information about an object to be replaced and information about an updated object of the original image; An extraction module, used to extract features from the text description information through a text encoder in a multimodal model to obtain text features; wherein the text features include a first text feature corresponding to the object information to be replaced and a second text feature corresponding to the update object information; A recognition module, configured to perform object recognition on the original image based on the first text feature to obtain a target area in the original image; wherein the target area is the area in the original image pointed to by the object information to be replaced; An updating module is used to update the target area in the original image based on the second text feature to obtain a target image.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.