Image generation methods and apparatuses, and device and storage medium
By editing the original image using an image generation model and combining it with a text generation model to generate descriptive text, the problem of limited image generation methods is solved. This enables interactive generation of images and text, and the generated emoji images match the content, thus enriching the image generation methods.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-09-01
- Publication Date
- 2026-04-23
AI Technical Summary
Existing image generation methods are relatively simple and lack diversity and richness, especially in the interactive generation methods between images and text.
The original image elements are edited using an image generation model, and descriptive text matching the edited image is generated using a text generation model. This text is then added to the image to create an emoji image.
It enables automatic editing of image elements and automatic generation of text. The text content in the generated emoticon images matches the image content, enriching the image generation methods and improving generation efficiency and effectiveness.
Smart Images

Figure CN2025118294_23042026_PF_FP_ABST
Abstract
Description
Image generation methods, apparatus, devices and storage media
[0001] This application claims priority to Chinese Patent Application No. 202411452955.0, filed on October 16, 2024, entitled “Image Generation Method, Apparatus, Device and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technology, and in particular to an image generation method, apparatus, device, and storage medium. Background Technology
[0003] With the rapid development of computer technology, image generation has gradually become an important research direction. Image generation mainly includes two branches: text-to-image and image-to-image.
[0004] In related technologies, such as text-to-image generation, inputting text into a text-to-image model will produce an image that matches the input text. Similarly, in image-to-image generation, inputting an image into an image-to-image model will produce an image that is similar to the input image.
[0005] Therefore, the image generation methods in the aforementioned related technologies are relatively simple. Summary of the Invention
[0006] This application provides an image generation method, apparatus, device, and storage medium, which can enrich the ways of image generation. The technical solution provided by this application includes the following aspects.
[0007] According to one aspect of the embodiments of this application, an image generation method is provided, the method being executed by a computer device, the method comprising the following steps.
[0008] Obtain the original image, which includes at least two image elements;
[0009] The original image is input into an image generation model to obtain at least one edited image output by the image generation model, wherein the edited image has at least one changed image element compared to the original image;
[0010] The at least one edited image is input into a text generation model to obtain descriptive text corresponding to the at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image;
[0011] For the at least one edited image, add the corresponding descriptive text to the edited image to obtain at least one emoji image.
[0012] According to one aspect of the embodiments of this application, another image generation method is provided, the method comprising the following steps.
[0013] Display the original image, which includes at least two image elements;
[0014] In response to an editing operation on at least one of the at least two image elements, at least one edited image is displayed, wherein at least one image element has changed compared to the original image;
[0015] After selecting a first edited image from the at least one edited image, a descriptive text for the first edited image is added to the first edited image to obtain an emoji image corresponding to the first edited image. The descriptive text for the first edited image is automatically generated text that matches the image content of the first edited image.
[0016] According to one aspect of the embodiments of this application, an image generation apparatus is provided, the apparatus comprising the following modules.
[0017] An image acquisition module is used to acquire an original image, wherein the original image includes at least two image elements;
[0018] An image generation module is used to input the original image into an image generation model to obtain at least one edited image output by the image generation model, wherein the edited image has at least one image element changed compared to the original image;
[0019] A text generation module is used to input the at least one edited image into a text generation model to obtain descriptive text corresponding to the at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image;
[0020] The emoji generation module is used to add descriptive text corresponding to the edited image to the edited image to obtain at least one emoji image.
[0021] According to one aspect of the embodiments of this application, another image generation apparatus is provided, the apparatus comprising the following modules.
[0022] A display module is used to display an original image, which includes at least two image elements;
[0023] The display module is further configured to, in response to an editing operation on at least one of the at least two image elements, display at least one edited image, wherein the edited image has at least one image element changed compared to the original image;
[0024] The adding module is further configured to, after selecting a first edited image from the at least one edited image, add descriptive text of the first edited image to the first edited image to obtain an emoji image corresponding to the first edited image, wherein the descriptive text of the first edited image is automatically generated text that conforms to the image content of the first edited image.
[0025] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described image generation method.
[0026] According to one aspect of the present application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, the computer program being loaded and executed by a processor to implement the above-described image generation method.
[0027] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program, the computer program being loaded and executed by a processor to implement the above-described image generation method.
[0028] The technical solutions provided in this application can bring the following beneficial effects.
[0029] An image generation model is used to modify image elements in the original image to obtain an edited image. A text generation model is then used to generate descriptive text that matches the content of the edited image. Finally, the descriptive text is added to the edited image to obtain the emoji image. The image generation in this application is no longer a simple image-to-image or text-to-image conversion. This application achieves image element editing in the original image through an image generation model, and generates descriptive text that matches the edited image through a text generation model. The final emoji image is obtained by adding the descriptive text to the edited image. Therefore, this application combines automatic image editing and automatic text generation, enriching the image generation methods.
[0030] Furthermore, the image editing and text generation processes are performed sequentially. Image editing is performed first to obtain the edited image. Then, text generation is performed on the edited image to ensure that the descriptive text in the final emoji image matches the edited image. In other words, the text content and image content in the emoji images generated by this application match, resulting in better image generation quality. Attached Figure Description
[0031] Figure 1 is a schematic diagram of a computer system provided in an embodiment of this application;
[0032] Figure 2 is a block diagram of an image generation method provided in an embodiment of this application;
[0033] Figure 3 is a flowchart of an image generation method provided in an embodiment of this application;
[0034] Figure 4 is a flowchart of an image generation method provided in another embodiment of this application;
[0035] Figure 5 is a flowchart of an image generation method provided in another embodiment of this application;
[0036] Figure 6 is a block diagram of an image generation method provided in another embodiment of this application;
[0037] Figure 7 is a schematic diagram of an image generation method provided in an embodiment of this application;
[0038] Figure 8 is a schematic diagram of a method for selecting a second image element according to an embodiment of this application;
[0039] Figure 9 is a schematic diagram of a method for uploading and replacing images provided in an embodiment of this application;
[0040] Figure 10 is a schematic diagram of an emoticon image provided in an embodiment of this application;
[0041] Figure 11 is a schematic diagram of an image generation method provided in another embodiment of this application;
[0042] Figure 12 is a schematic diagram of an image generation method provided in another embodiment of this application;
[0043] Figure 13 is a schematic diagram of an image generation method provided in another embodiment of this application;
[0044] Figure 14 is a schematic diagram of a method for adjusting descriptive text provided in an embodiment of this application;
[0045] Figure 15 is a schematic diagram of an application scenario of the emoticon image provided in an embodiment of this application;
[0046] Figure 16 is a block diagram of an image generation method provided in another embodiment of this application;
[0047] Figure 17 is a block diagram of an image generation apparatus provided in an embodiment of this application;
[0048] Figure 18 is a block diagram of an image generation apparatus provided in another embodiment of this application;
[0049] Figure 19 is a structural block diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0051] Please refer to Figure 1, which shows a schematic diagram of a computer system provided in an exemplary embodiment of this application. The computer system may include: a terminal device 10 and a server 20.
[0052] Terminal device 10 includes, but is not limited to, mobile phones, tablets, smart voice interaction devices, game consoles, wearable devices, multimedia playback devices, PCs (Personal Computers), in-vehicle terminals, smart home appliances, and other electronic devices. A client application for the target application can be installed on terminal device 10. Optionally, the target application can be an application that requires downloading and installation, or it can be an application that can be used instantly; this embodiment of the application does not limit this.
[0053] In this embodiment, the target application is an application for image generation. For example, as shown in FIG2, the target application acquires an original image 210, which includes at least two image elements; it generates at least one edited image 220 based on the original image using an image generation model, where at least one image element has changed compared to the original image; it generates descriptive text corresponding to each of the at least one edited image 220 using a text generation model, the descriptive text conforming to the image content of the edited image 220; and it adds the descriptive text corresponding to each of the at least one edited image 220 to obtain at least one emoticon image 230. Of course, the specific type of the target application is not limited. The target application can be an image application, image processing application, image generation application, image search application, video application, animation application, social application, question-and-answer application, virtual reality application, augmented reality application, etc. This embodiment does not limit the specific category of the target application. In other embodiments, the target application can be considered a separate functional module, such as being implemented as one of the image generation functional modules within a social application. For example, a client running the aforementioned target application is located in terminal device 10.
[0054] Server 20 is used to provide backend services for the client of the target application in terminal device 10. For example, server 20 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms, but it is not limited to these.
[0055] Terminal device 10 and server 20 can communicate with each other via a network. This network can be a wired network or a wireless network.
[0056] The method provided in this application embodiment can be executed by a computer device. A computer device can be any electronic device capable of storing and processing data. For example, a computer device can be the terminal device 10 in Figure 1, or it can be a server 20.
[0057] Please refer to Figure 3, which shows a flowchart of an image generation method provided in an embodiment of this application. The executing entity of this method can be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, or the server 20 described above. In the following method embodiments, for ease of description, only the executing entity of each step is described as a "computer device". The method may include at least one of the following steps (310-340):
[0058] Step 310: Obtain the original image, which includes at least two image elements.
[0059] In some embodiments, the original image is the image to be edited. Exemplarily, the original image may be an image uploaded by the user or an image obtained from a database on a computer device. Exemplarily, the original image may be an image captured by a camera of a real-world object, or an image automatically generated by an artificial intelligence model. Of course, there are no limitations on the image size, image resolution, image category, etc., of the original image. Exemplarily, the original image may be a landscape photo, a portrait photo, a screenshot, etc.
[0060] In some embodiments, an image element is a directly visible element included in the original image. Exemplarily, the image element includes, but is not limited to, objects and the background in the original image. For example, when the original image is a portrait, the image element can be the individuals in the portrait, or it can be the background in which the individuals are situated. Exemplarily, each object included in the original image can be considered a separate image element, such as a person being considered an image element, or a tree being considered an image element. Exemplarily, the background in which an object in the original image is situated can be considered an image element.
[0061] In other embodiments, the image element can be not only the visible elements included in the original image, but also elements related to the original image. For example, the image element may be the brightness, sharpness, filters, etc. of the original image.
[0062] Step 320: Input the original image into the image generation model to obtain at least one edited image output by the image generation model. The edited image has at least one image element changed compared to the original image.
[0063] In some embodiments, the image generation model is a neural network model for editing at least one image element in an original image. Exemplarily, this image generation model may also be referred to as an image inpainting model. Exemplarily, this image generation model is a pre-training model (PTM). Exemplarily, in this application, the image generation model is directly invoked through its interface to generate at least one edited image based on the original image. Exemplarily, the image generation model uses a stable diffusion model image editing API (Application Programming Interface).
[0064] Pre-trained models, also known as foundational models or large models, refer to deep neural networks (DNNs) with large parameters, trained on massive amounts of unlabeled data. Leveraging the function approximation capabilities of large-parameter DNNs, pre-trained models (PTMs) extract common features from the data. Through fine-tuning, efficient parameter fine-tuning, and prompt-tuning techniques, they are suitable for downstream tasks. Therefore, pre-trained models can achieve ideal results in small-shot or zero-shot scenarios. PTMs can be categorized according to the data modality they process, including language models, vision models (swin-transformer, ViT (Vision Transformers), V-MOE (Vision Mixture of Experts)), speech models, and multimodal models. Multimodal models refer to models that establish feature representations for two or more data modalities. Pre-trained models are important tools for outputting AI-generated content and can also serve as a general interface connecting multiple specific task models. In the embodiments of this application, the image generation model, or at least one module or sub-model within the image generation model, is a pre-trained model.
[0065] In some embodiments, the input to the image generation model is an original image, and the output is at least one edited image. In some embodiments, the edited image is an image obtained by changing at least one image element based on the original image. For example, the image element that has changed in the original image is considered the first image element.
[0066] In some embodiments, the change is considered a replacement. In some embodiments, the first image element is replaced in the edited image, such as by replacing the first image element with a second image element, which is a different image element from the first image element. Exemplarily, the second image element and the first image element are completely different. Exemplarily, the input to the image generation model is the original image and the first image element in the original image, and the output is at least one edited image. Exemplarily, the input to the image generation model is the original image, the first image element in the original image, and the second image element used to replace the first image element, and the output is at least one edited image.
[0067] In some embodiments, the change is considered an addition. In some embodiments, the edited image is changed from a first image element to a third image element. This third image element includes the content of the first image element, but also includes other elements (the additional elements can also be called a fourth image element). For example, if the first image element is a child, the third image element is a mother holding the child's hand, then the additional fourth image element is also a mother. Exemplarily, the input to this image generation model is the original image, the first image element, and the fourth image element; the output is at least one edited image.
[0068] In some embodiments, the change is considered to be a modification based on the input descriptive text. In some embodiments, the edited image is obtained by editing a first image element in conjunction with the edited text. Exemplarily, the edited text is text input by the user. Exemplarily, the edited text is text used to edit the original image. Exemplarily, the edited text is "Change the child in the picture to an old person". Then, based on the edited text, the first image element is determined to be a child. Exemplarily, the input of the image generation model is the original image and the edited text, and the output is at least one edited image. Exemplarily, the input of the image generation model is the original image, the first image element, and the edited text, and the output is at least one edited image.
[0069] In some embodiments, when generating at least one edited image using an image generation model, the generation method includes at least one of the following: parallel generation, serial generation, or a combination of serial and parallel generation. For example, in parallel generation, at least one edited image is generated simultaneously by the image generation model. For example, in serial generation, at least one edited image is generated sequentially by the image generation model. For example, in a combination of parallel and serial generation, at least one edited image is generated batch by batch sequentially by the image generation model.
[0070] Step 330: Input at least one edited image into the text generation model to obtain descriptive text corresponding to at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image.
[0071] In some embodiments, the text generation model is a neural network model for generating descriptive text. Exemplarily, this text generation model may also be called a creative copywriting generation model. Exemplarily, this text generation model is obtained by further fine-tuning a hybrid multimodal base model. Exemplarily, the base model refers to a basic model trained on a large amount of pre-trained data, which typically has strong general capabilities but limited capabilities in vertical domains. Exemplarily, the hybrid multimodal base model is a pre-trained model. Exemplarily, the text generation model in this application is obtained by further fine-tuning the pre-trained hybrid multimodal base model using a first text database. Exemplarily, the first text database includes multiple tag descriptive texts. Exemplarily, the hybrid multimodal base model is further fine-tuned using these multiple tag descriptive texts as tags. Exemplarily, the first text database is also called a creative copywriting database, and the tag descriptive text included in the first text database may also be called creative copywriting. Exemplarily, the text in the first text database can be collected from the internet based on its popularity or can be manually written by experts. The text generation model in the embodiments of this application, or at least one module or sub-model in the text generation model, is a pre-trained model.
[0072] In some embodiments, the text generation model takes an edited image as input and outputs descriptive text corresponding to the edited image. For example, the descriptive text corresponding to the edited image conforms to the image content of the edited image. For example, since the text generation model is fine-tuned by a creative copywriting database, the descriptive text can also reflect a certain degree of creativity. This differs somewhat from text generated by image-to-text technology in related technologies. Of course, the descriptive text automatically generated by the model can be adjusted or modified by the user.
[0073] In some embodiments, when generating descriptive text corresponding to at least one edited image using a text generation model, the generation includes at least one of the following methods: parallel generation, serial generation, or a combination of serial and parallel generation. For example, in parallel generation, the descriptive text corresponding to at least one edited image is generated simultaneously by the text generation model. For example, in serial generation, the descriptive text corresponding to at least one edited image is generated sequentially by the text generation model. For example, in a combination of parallel and serial generation, the descriptive text corresponding to at least one edited image is generated batch by batch sequentially by the text generation model.
[0074] In some embodiments, the image generation model and text generation model in this application can be considered as two independent models, or as two sub-models integrated into the same larger model. For example, an emoji generation model generates at least one emoji image based on an original image. For example, the input to the emoji generation model is the original image, and the output is at least one emoji image. For example, the emoji generation model is a neural network model for generating emoji images. For example, the image generation model and text generation model are considered as two sub-models or sub-modules within the emoji generation model. Of course, other models mentioned below can be considered as independent models, or as sub-models or sub-modules within the emoji generation model.
[0075] Step 340: For at least one edited image, add the corresponding descriptive text to the edited image to obtain at least one emoji image.
[0076] In some embodiments, emoticon images are used to express user emotions through images. Optionally, the type of emoticon image is an image. Optionally, emoticon images include text content and image content. Exemplarily, in this application, emoticon images can be automatically generated based on the original image using an image generation model + a text generation model. This requires minimal manual operation from the user, thus resulting in relatively high image generation efficiency.
[0077] In some embodiments, descriptive text corresponding to the edited image is added to the edited image to obtain an emoji image. For example, the position of the descriptive text within the emoji image is not limited.
[0078] In some embodiments, an emoji synthesis model is used to add descriptive text corresponding to at least one edited image to at least one edited image, thereby obtaining at least one emoji image. Exemplarily, the emoji synthesis model is a neural network model for adding descriptive text to edited images to obtain emoji images. Exemplarily, the emoji synthesis model automatically determines a suitable position in the edited image for adding descriptive text and adds the descriptive text to the edited image to obtain the emoji image.
[0079] In other embodiments, the user may pre-specify the location where the descriptive text will be added. For example, a compositing module in a computer device adds the descriptive text to the edited image based on the descriptive text, the edited image, and the location where the descriptive text will be added. In other embodiments, the text generation model automatically determines the location of the descriptive text in the edited image while generating the descriptive text. Of course, the location of the descriptive text automatically determined by the model can be adjusted or modified by the user.
[0080] The technical solution provided in this application uses an image generation model to modify image elements in the acquired original image to obtain an edited image. A text generation model is then used to generate descriptive text that matches the content of the edited image. Finally, the descriptive text is added to the edited image to obtain an emoji image. The image generation in this application is no longer a simple image-to-image or text-to-image conversion. This application, on the one hand, uses an image generation model to edit image elements in the original image; on the other hand, it uses a text generation model to generate descriptive text that matches the edited image. The final emoji image is obtained by adding the descriptive text to the edited image. Therefore, this application combines automatic image editing and automatic text generation, enriching the image generation methods.
[0081] Furthermore, the image editing and text generation processes are performed sequentially. Image editing is performed first to obtain the edited image. Then, text generation is performed on the edited image to ensure that the descriptive text in the final emoji image matches the edited image. In other words, the text content and image content in the emoji images generated by this application match, resulting in better image generation quality.
[0082] Please refer to Figure 4, which shows a flowchart of an image generation method provided in another embodiment of this application. The execution subject of this method can be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, or the server 20 described above. In the following method embodiments, for ease of description, only the execution subject of each step is described as a "computer device". The method may include at least one of the following steps (410-460):
[0083] Step 410: Obtain the original image, which includes at least two image elements.
[0084] In some embodiments, the following steps are included after step 410.
[0085] In some embodiments, the original image is input into the content detection model to obtain the retained original image output by the content detection model. The content detection model is used to retain the original image if the original image meets a first condition, and to clean the original image if the original image does not meet the first condition. The first condition is the condition for the content detection model to filter the original image.
[0086] In some embodiments, the original image is filtered by a content detection model. The content detection model is used to retain the original image if it meets a first condition, and to clean the original image if it does not meet the first condition. The first condition is the condition for the content detection model to filter the original image.
[0087] Exemplarily, the content detection model is a neural network model used to detect content in an original image. Exemplarily, the content detection model is a pre-trained model, or at least one module or sub-model of the content detection model is a pre-trained model. Exemplarily, the content detection model may also be called a content inspection model. Exemplarily, the content detection model is a GPT-4v model. Exemplarily, the GPT-4v model is called through its interface. Exemplarily, the input of the content detection model is the original image, and the output is either pass or fail. When the output is pass, the original image is retained. When the output is fail, the original image is cleaned up or rejected. In the embodiments of this application, the content detection model, or at least one module or sub-model of the content detection model, is a pre-trained model.
[0088] For example, the first condition is the criterion used by the content detection model to filter the original image. For example, the first condition is that the original image cannot contain inappropriate information, such as pornography or violence. For example, when the content detection model detects inappropriate information in the original image, it cleans up or rejects the original image. When the content detection model detects that the original image does not contain inappropriate information, it retains the original image.
[0089] In some embodiments, the original image is retained and input into the object recognition model to obtain at least two image elements output by the object recognition model.
[0090] In some embodiments, the original image retained by the object recognition model is used to identify at least two image elements included in the original image. The object recognition model in this application embodiment, or at least one module or sub-model of the object recognition model, is a pre-trained model.
[0091] For example, the object recognition model is a neural network model for recognizing image elements in an image. For example, the object recognition model is a pre-trained model, or at least one module or sub-model in the object recognition model is a pre-trained model. For example, the object recognition model is a RAM++ model. For example, the RAM++ model is invoked through its interface.
[0092] For example, the input to the object recognition model is the original image that has been detected and retained by the object recognition model, and the output is at least two image elements present in the original image, such as all the objects present in the original image.
[0093] The technical solution provided in this application uses a content detection model to detect the original image, thereby filtering the original image, avoiding unwanted information, and improving the reliability of image generation. Furthermore, by using an object recognition model to obtain at least two image elements included in the original image, the accuracy of object and background recognition is improved, as well as recognition efficiency.
[0094] Step 420: Determine at least one first image element to be replaced from at least two image elements included in the original image.
[0095] In some embodiments, the user needs to manually determine the first image element to be replaced, or the at least one image element to be replaced can be determined automatically. For example, when the user needs to manually determine the first image element to be replaced, this can be done through a selection operation of at least one image element displayed on the screen of the terminal device. The image element selected by the user is then used as the image element to be replaced. For example, when the at least one image element to be replaced is determined automatically, the image generation model automatically identifies the original image and determines the incongruous image element in the original image as the at least one first image element to be replaced. For example, when the at least one image element to be replaced is determined automatically, the image generation model can also determine the at least one image element to be replaced based on the edited text or edited voice input by the user, without requiring manual selection by the user.
[0096] In some embodiments, at least two mask images are generated corresponding to image elements, which are used to distinguish image elements from other regions besides image elements in the original image.
[0097] For example, a mask image is used to distinguish image elements from other areas in the original image. For example, the mask image preserves the image elements in the original image while covering other areas in the original image, such as covering other areas in the original image with black or gray, so that the user can directly and clearly see the position of the selected first image element in the original image.
[0098] In some embodiments, the image segmentation model is invoked in parallel through at least two interfaces corresponding to the image segmentation model. Each of the at least two interfaces corresponding to the image segmentation model is used to invoke the image segmentation model to generate a mask image corresponding to one of the at least two image elements.
[0099] Exemplarily, an image segmentation model generates mask images corresponding to at least two image elements. Exemplarily, this image segmentation model is a neural network model for generating mask images corresponding to at least two image elements. Exemplarily, the input of this image segmentation model is the original image, and the output is the mask images corresponding to at least two image elements. Exemplarily, this image segmentation model is a SAM (Segment Anything Model). The image segmentation model in this embodiment, or at least one module or sub-model within the image segmentation model, is a pre-trained model.
[0100] In some embodiments, when the computer device performing the above method is a server, after generating mask images corresponding to at least two image elements respectively, the server sends the mask images corresponding to at least two image elements respectively to the terminal device, which then provides them to the user for selection.
[0101] In some embodiments, the image element corresponding to the selected mask image in the mask images corresponding to at least two image elements is determined as at least one first image element to be replaced.
[0102] For example, the image element corresponding to the selected mask image is determined as the first image element. When at least two mask images are selected, there are at least two first image elements.
[0103] The technical solution provided in this application embodiment can directly and clearly display the position of the first image element to be replaced in the original image by selecting a mask image, thereby reducing the occurrence of selection errors and improving the accuracy of image generation.
[0104] In some embodiments, the image segmentation model is invoked in parallel through at least two interfaces corresponding to the image segmentation model. Each of the at least two interfaces corresponding to the image segmentation model is used to invoke the image segmentation model to generate a mask image corresponding to one of the at least two image elements.
[0105] In some embodiments, the image segmentation model is invoked in parallel through at least two interfaces corresponding to the image segmentation model to generate mask images corresponding to at least two image elements respectively. Each of the at least two interfaces corresponding to the image segmentation model is used to invoke the image segmentation model to generate a mask image corresponding to one of the at least two image elements. In some embodiments, the interface is an API.
[0106] The technical solution provided in this application improves the efficiency of mask image generation by using a parallel generation method that calls the image segmentation model in parallel through at least two interfaces corresponding to the image segmentation model. Furthermore, the interface-based approach not only saves time but also reduces the hardware requirements for computer equipment.
[0107] Step 430: Determine the second image element to replace the first image element.
[0108] In some embodiments, the second image element used to replace the first image element is determined by the user or automatically generated by the image generation model. The second image element is an image element different from the first image element. For example, if the first image element is an image background, such as a seaside, then the second image element is also an image background, such as a forest.
[0109] In some embodiments, the second image element is an image element that is not present in the original image. Of course, the second image element can also be an image element included in the original image. For example, when two image elements out of at least two image elements included in the original image are selected sequentially, the image element selected first is used as the first image element, and the image element selected later is used as the second image element. The first image element is replaced with the second image element while keeping the second image element unchanged.
[0110] In some embodiments, at least two replacement elements for replacing the first image element are obtained from an image element library, and one or more replacement elements are selected from the at least two replacement elements as the second image element.
[0111] For example, a user can choose one or more replacement elements from at least two replacement elements as the second image element.
[0112] For example, when retrieving at least two replacement elements from the image element library to replace the first image element, the at least two replacement elements need to be determined based on the item category represented by the first image element. For example, if the first image element is a soccer ball, then the at least two replacement elements could also be ball-type objects, such as basketballs, table tennis balls, volleyballs, etc. For example, when retrieving at least two replacement elements from the image element library to replace the first image element, the at least two replacement elements need to be determined based on the image background represented by the first image element. For example, if the first image element is an image background (such as a seaside), then the at least two replacement elements could also be image backgrounds, such as forests, lakesides, tunnels, etc.
[0113] In some embodiments, a replacement image uploaded by a user is acquired, and image elements identified from the replacement image are used as second image elements. For example, a user may upload a replacement image, which is then acquired by a computer device, and image elements identified from the replacement image are used as second image elements.
[0114] In some embodiments, if the user does not determine the second image, the second image elements are automatically and randomly generated by the image generation model.
[0115] The technical solution provided in this application provides at least two replacement objects for the user to choose from, while also allowing the user to upload their own replacement images. This demonstrates the flexibility and diversity of image editing and provides users with a better image editing experience.
[0116] Step 440: Input the original image, the first image element, and the second image element into the image generation model to obtain at least one edited image output by the image generation model.
[0117] In some embodiments, where the second image element is determined by the user, the input to the image generation model is the original image, the first image element, and the second image element, and the output is at least one edited image. For example, the image generation model replaces the first image element with the second image element and performs adaptive processing on all image elements in the replaced image to make the individual image elements in the edited image appear more natural and harmonious. For example, when both the first image element and the second image are objects, the first image element is replaced with the second image element, and it is blended with surrounding image elements (such as the image background) to ensure that there are no visual discrepancies.
[0118] In some embodiments, where the second image element is generated automatically by the model, the input to the image generation model is the original image and the first image element, and the output is at least one edited image. The second image element is randomly generated automatically by the image generation model.
[0119] In some embodiments, when there are at least two edited images, the image generation model is invoked in parallel through at least two interfaces corresponding to the image generation model, wherein each of the at least two interfaces corresponding to the image generation model is used to invoke the image generation model to generate an edited image.
[0120] In some embodiments, when there are at least two edited images, the image generation model is invoked in parallel through at least two interfaces corresponding to the image generation model to generate at least two edited images based on the original image, a first image element, and a second image element. Each of the at least two interfaces corresponding to the image generation model is used to invoke the image generation model to generate one edited image. In some embodiments, the interface is an API.
[0121] The technical solution provided in this application improves the efficiency of generating edited images by using a parallel generation method that calls the image generation model in parallel through at least two interfaces corresponding to the image generation model. Furthermore, this interface-based approach not only saves time but also reduces the hardware requirements for computer equipment.
[0122] In some embodiments, if there is one first image element, there is also one second image element, and the number of edited images generated based on one first image element and one second image element is N.
[0123] In some embodiments, when there are at least two first image elements, there are also at least two second image elements. The number of edited images generated based on one first image element and one second image element is M, and the number of edited images generated based on at least two first image elements and at least two second image elements is N; where N is an integer greater than 1, and M is a positive integer less than N.
[0124] For example, a user can select one or at least two first image elements and one or at least two second image elements. For example, when one first image element and one second image element are selected, N edited images are generated. For example, N is 9.
[0125] For example, when at least two first image elements and at least two second image elements are selected, N edited images are generated. The final number of edited images generated by one first image element and the second image element replacing it is M. For example, when two first image elements and two corresponding second image elements are selected, nine edited images are generated. Specifically, the final number of edited images generated by one first image element and the second image element replacing it is five, and the final number of edited images generated by another first image element and the second image element replacing it is four. For example, when three first image elements and three corresponding second image elements are selected, nine edited images are generated. Specifically, the final number of edited images generated by each first image element and the second image element replacing it is three.
[0126] The technical solution provided in this application allows the selection of one or at least two first image elements and corresponding second image elements, and the total number of edited images generated remains unchanged, always being N. Therefore, without increasing the generation cost, it provides users with more options and space for image editing, not only improving the user experience of image editing but also further enriching the diversity of image generation.
[0127] Step 450: Input at least one edited image into the text generation model to obtain descriptive text corresponding to at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image.
[0128] In some embodiments, descriptive text corresponding to at least one edited image is generated by text generation models deployed on at least two GPUs (Graphics Processing Units). In some embodiments, the text generation models are deployed on at least two GPUs.
[0129] For example, text generation models deployed on eight GPUs respectively generate descriptive text corresponding to at least one edited image. For example, one text generation model deployed on each GPU generates descriptive text corresponding to one edited image at a time. For example, text generation models deployed on at least two GPUs respectively generate descriptive text corresponding to at least one edited image in parallel.
[0130] The technical solution provided in this application, through parallel generation using text generation models deployed on at least two GPUs respectively, also saves text generation time and improves the efficiency of generating descriptive text.
[0131] Step 460: For at least one edited image, add the corresponding descriptive text to the edited image to obtain at least one emoji image.
[0132] In some embodiments, for the i-th edited image among at least one edited image, adjustment information of the descriptive text for the i-th edited image is obtained, the adjustment information including at least one of the following: color, font, text content, display position, display size, where i is a positive integer.
[0133] For example, the i-th edited image can be considered as the image selected by the user to save from at least two edited images. For example, the adjustment information of the descriptive text for the i-th edited image is obtained based on the user's adjustment operation on the descriptive text of the i-th edited image.
[0134] For example, users can adjust various aspects of the descriptive text, such as color, font, text content, display position, and display size.
[0135] In some embodiments, the description text of the i-th edited image is adjusted based on the adjustment information to obtain the adjusted description text of the i-th edited image.
[0136] For example, based on the specific adjustment content indicated by the adjustment information, the descriptive text of the i-th edited image is adjusted to obtain the adjusted descriptive text of the i-th edited image. In some embodiments, at least one of the color, font, text content, display position, and display size of the descriptive text of the i-th edited image is adjusted based on the adjustment information to obtain the adjusted descriptive text of the i-th edited image. The adjusted descriptive text differs from the original descriptive text in at least one of the following aspects: color, font, text content, display position, and display size.
[0137] In some embodiments, the adjusted descriptive text of the i-th edited image is added to the i-th edited image to obtain the i-th emoji image.
[0138] Of course, users can also leave the description text unchanged and directly add the description text of the i-th edited image to the i-th edited image to obtain the i-th emoji image.
[0139] As shown in Figure 6, the user uploads an image, i.e., the original image. The user selects an object to replace, which includes selecting a first image element and a second image element. After selecting the object to replace, a 9-grid is automatically generated, i.e., at least two emoji images are generated (including the edited image and descriptive text). The user can choose the emoji images they want to download. Based on the user's selection, the system continuously refines and adjusts to better suit the user's preferences through background tagging.
[0140] The technical solution provided in this application allows for the replacement of image elements in the original image, offering an editing method for the original image and enriching image generation methods. Simultaneously, the generated descriptive text allows users to further adjust it to ensure the final generated emoji image meets user needs, thereby improving the hit rate of emoji images.
[0141] Please refer to Figure 5, which shows a flowchart of an image generation method provided in another embodiment of this application. The execution subject of this method can be the terminal device 10 described above, such as the client of the target application mentioned in the above embodiments, where the execution subject of each step is the client of the target application. In the following method embodiments, for ease of description, only the execution subject of each step will be described as a "computer device". The method may include at least one of the following steps (510-530):
[0142] Step 510: Display the original image, which includes at least two image elements.
[0143] For example, the original image is displayed on the screen of the terminal device.
[0144] Referring to Figure 7, interface 710 in Figure 7 prompts users to create emoticon images. Interface 720 in Figure 7 prompts users to log in to create emoticon images. Interface 730 in Figure 7 prompts users on how to create emoticon images. Interface 740 in Figure 7 displays the created emoticon images.
[0145] Referring to Figure 8, in response to a click operation on the upload control 810 in Figure 8, the original image 820 is uploaded.
[0146] Step 520, in response to an editing operation on at least one of at least two image elements, display at least one edited image, wherein at least one image element has changed compared to the original image.
[0147] In some embodiments, an editing operation on at least one of at least two image elements can be considered as an operation to edit that at least one image element. Such operations include, but are not limited to, clicks, long presses, double-clicks, etc.
[0148] For example, when a user clicks on at least one of at least two image elements, it is considered that an editing operation has been performed on at least one of the at least two image elements. For example, when a user long-presses on at least one of at least two image elements, it is considered that an editing operation has been performed on at least one of the at least two image elements.
[0149] For example, the image element pointed to by the editing operation is considered to be the image element to be edited. For example, the image generation model mentioned in the above embodiments generates at least one edited image based on the image element pointed to by the editing operation. Of course, the image generation model mentioned in this application embodiment can be deployed on a terminal device or on a server. When the image generation model is deployed on a server, the terminal device sends the image element pointed to by the editing operation and the original image to the server. The server, based on the image element pointed to by the editing operation and the original image, calls the image generation model to generate at least one edited image, sends the at least one edited image to the terminal device, and displays it on the terminal device for the user to select and save. Of course, when the image generation model is directly deployed on the terminal device, it can directly call the image generation model to generate at least one edited image based on the image element pointed to by the editing operation and the original image, and then display it.
[0150] In some embodiments, the editing operation is an operation for replacing at least one image element. Exemplarily, the editing operation is an operation of selecting a first image element to be replaced in the original image and an operation of selecting a second image element to replace the first image element.
[0151] In some embodiments, marking information corresponding to at least one image element is displayed, the marking information being used to indicate the image element. Referring to FIG8, marking information 830 corresponding to at least one image element is displayed.
[0152] For example, the labeling information corresponding to an image element is used to indicate that image element. For example, different image elements use different numbers, which are the labeling information of the image element. Of course, the labeling information can also be at least one of other graphic symbols, numbers, letters, etc.
[0153] In some embodiments, in response to a selection operation on the marker information corresponding to a first image element, a mask image corresponding to the first image element is displayed. The mask image is used to distinguish the first image element from other regions in the original image. The first image element is one or at least two of at least one image element to be replaced. Referring to FIG11, in response to a selection operation on the marker information corresponding to the first image element in the original image 1110, a mask image 1120 corresponding to the first image element is displayed, and finally at least two generated emoji images 1130 are displayed.
[0154] For example, the selection operation of the marker information corresponding to the first image element can be considered as the operation of selecting the marker information corresponding to the first image element. The operation type of this operation includes, but is not limited to, click operation, long press operation, double click operation, etc. Referring to Figure 8, the selection operation of the marker information corresponding to the first image element is a click operation on marker 1.
[0155] For example, in response to a selection operation of the tag information corresponding to the first image element, the terminal device obtains a mask image corresponding to an image element from the local device or from the server.
[0156] In some embodiments, in response to the selection of a second image element, at least one edited image is displayed, the second image element being used to replace the first image element.
[0157] For example, the operation of selecting a second image element is an operation for selecting a second image element.
[0158] In some embodiments, at least two replacement elements are displayed, which are elements in an image element library used to replace the first image element; in response to a selection operation on one or both of the at least two replacement elements, the selected replacement element is used as the selected second image element. Referring to FIG8, at least two replacement elements 840 are displayed. In response to a selection operation on one or both of the at least two replacement elements 840, the selected replacement element 850 is used as the selected second image element. In response to an operation on control 860, at least one edited image is displayed.
[0159] In some embodiments, in response to an operation of uploading a replacement image in the input field, image elements included in the replacement image are selected as second image elements. For example, a user can upload a replacement image in the input field. Referring to FIG9, in response to an operation on custom control 910, input field 920 is displayed. In response to an operation of uploading a replacement image in input field 920, image elements included in the replacement image are selected as second image elements.
[0160] Step 530: After selecting the first edited image from at least one edited image, add descriptive text to the first edited image to obtain the emoticon image corresponding to the first edited image. The descriptive text of the first edited image is automatically generated text that matches the image content of the first edited image.
[0161] Referring to Figure 10, at least two emoji images 1000 are displayed. Users can choose their favorite emoji image to download. Referring to Figure 12, the background of the original image 1210 is replaced to obtain the edited image 1220. Further, descriptive text is added to the edited image to obtain emoji image 1230. Referring to Figure 13, the background of the original image 1310 is replaced to obtain the edited image 1320. Further, descriptive text is added to the edited image to obtain emoji image 1330.
[0162] For example, at least one edited image is displayed on the terminal device, allowing the user to select the edited image they want to save. The edited image selected by the user is considered the first edited image here. For example, the user can select at least two edited images at once, that is, at least two first edited images.
[0163] For example, for the edited image selected by the user, the descriptive text corresponding to the edited image is displayed. For example, the text generation model mentioned in the above embodiments generates the descriptive text corresponding to the first edited image based on the first edited image. Of course, the text generation model mentioned in this application embodiment can be deployed on a terminal device or a server. When the text generation model is deployed on a server, the terminal device sends the first edited image selected by the user to the server. The server, based on the first edited image, calls the text generation model to generate the descriptive text corresponding to the first edited image, sends the descriptive text corresponding to the first edited image to the terminal device, and displays it on the terminal device for further adjustment or modification by the user. Of course, when the text generation model is directly deployed on the terminal device, it can directly call the text generation model based on the first edited image to generate the descriptive text corresponding to the first edited image and display it.
[0164] In some embodiments, descriptive text of the first edited image to be edited is displayed on the first edited image.
[0165] In some embodiments, in response to an adjustment operation on the descriptive text of the first edited image, the adjusted descriptive text obtained after adjusting the descriptive text of the first edited image is displayed. Referring to FIG14, the first descriptive text 1420, the second descriptive text 1430, and the third descriptive text 1440 are adjusted using different controls on the toolbar 1410.
[0166] For example, the adjustment operation for the descriptive text of the first edited image is an operation for adjusting at least one of the descriptive text's color, font, text content, display position, and display size. The operation type of this adjustment operation includes, but is not limited to, click operations, voice input operations, drag operations, etc.
[0167] In some embodiments, in response to a confirmation operation for the adjustment operation, adjusted descriptive text is added to the first edited image to obtain an emoji image corresponding to the first edited image.
[0168] For example, a confirmation operation for an adjustment operation is an operation used to confirm the aforementioned adjustment operation. Confirmation operations for adjustment operations include, but are not limited to, operations on a confirmation control. A confirmation control is a control on the user interface. Referring to Figure 14, a confirmation operation for an adjustment operation includes a click operation on the confirmation control 1450.
[0169] In some embodiments, adjustment information is obtained in response to a confirmation operation for the adjustment operation. The adjustment information includes at least one of the following: color, font, text content, display position, and display size, where i is a positive integer.
[0170] In some embodiments, the descriptive text of the first edited image is adjusted based on the adjustment information to obtain the adjusted descriptive text of the first edited image.
[0171] Referring to Figure 15, after a user saves an emoji image, they can send the emoji image 1520 in a social chat and simultaneously receive emoji images 1510 sent by friends.
[0172] Referring to Figure 16, content detection model 1610 performs content detection on the original image (i.e., the original image) provided by the user. Item recognition model 1620 identifies the various image elements included in the original image. Further, image elements are segmented and a mask image (i.e., a masked image) is generated using image segmentation model 1630 and a grounding model (given an input image and an item label, the model generates the bounding box of the item). For example, the user selects a mask image 1650 (corresponding to the first image element to be replaced) on the front end 1693 (i.e., the screen of the terminal device), and a replacement 1660 (i.e., the second image element to replace the first image element) selected by the user on the front end 1693. A replacement library 1640 (i.e., an image element library) is used to determine at least two replacement elements. Further, image editing model 1670 determines a creative image 1680 (i.e., the edited image) based on the selected replacement 1660 and the selected mask image 1650. Creative text 1691 (i.e., descriptive text) is generated using the creative text generation model 1690 (also known as the text-to-image model). Users can adjust the content of the creative text, including font size, color, etc. The creative text 1691 is then added to the creative image 1680 using the compositing module to obtain emoticon 1692 (i.e., emoticon image).
[0173] For technical details not mentioned in the embodiments of this application, please refer to the explanations of other embodiments in the context, and they will not be repeated here.
[0174] The technical solution provided in this application allows for image editing through editing operations on image elements, resulting in an edited image. Furthermore, descriptive text matching the content of the edited image is automatically generated. Finally, the descriptive text is added to the edited image to obtain an emoji image. Therefore, this application combines manual image editing with automatic text generation to generate emoji images, enriching the image generation methods.
[0175] Furthermore, by displaying marker information for image elements, users can select the first image element to be replaced with a single click, demonstrating the flexibility of image editing. Of course, displaying a mask image also helps inform users of the location of the image element they want to replace, reducing the possibility of accidental touches.
[0176] Of course, when replacing the first image element, users are also given multiple ways to choose the second image element, reflecting the diversity and flexibility of image editing methods.
[0177] Finally, users can also adjust the automatically generated descriptive text to ensure that the text content in the final emoji image meets their needs. This not only enriches the ways emoji images are generated but also improves the image generation quality to some extent. This encourages users to send the emoji image, thus promoting human-computer interaction.
[0178] The following are embodiments of the apparatus of this application, which can be used to execute the embodiments of the method of this application. For details not disclosed in the embodiments of the apparatus of this application, please refer to the embodiments of the method of this application.
[0179] Please refer to Figure 17, which shows a block diagram of an image generation apparatus provided in one embodiment of this application. This apparatus has the function of implementing the image generation method described above. This function can be implemented in hardware or by hardware executing corresponding software. The apparatus can be the computer device described above, or it can be installed within a computer device. As shown in Figure 17, the apparatus 1700 may include: an image acquisition module 1710, an image generation module 1720, a text generation module 1730, and an emoticon generation module 1740.
[0180] The image acquisition module 1710 is used to acquire an original image, which includes at least two image elements.
[0181] The image generation module 1720 is used to input the original image into the image generation model to obtain at least one edited image output by the image generation model, wherein the edited image has at least one image element changed compared to the original image.
[0182] The text generation module 1730 is used to input the at least one edited image into the text generation model to obtain the descriptive text corresponding to the at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image.
[0183] The emoji generation module 1740 is used to add descriptive text corresponding to the edited image to the edited image to obtain at least one emoji image.
[0184] In some embodiments, the image generation module 1720 is configured to determine at least one first image element to be replaced from the at least two image elements included in the original image; determine a second image element to replace the first image element; and input the original image, the first image element, and the second image element into the image generation model to obtain the at least one edited image output by the image generation model.
[0185] In some embodiments, the image generation module 1720 is configured to, when there are at least two edited images, call the image generation model in parallel through at least two interfaces corresponding to the image generation model, wherein each of the at least two interfaces corresponding to the image generation model is used to call the image generation model to generate an edited image.
[0186] In some embodiments, when there is one first image element, there is also one second image element, and the number of edited images generated based on one first image element and one second image element is N; when there are at least two first image elements, there are also at least two second image elements, and the number of edited images generated based on one first image element and one second image element is M, and the number of edited images generated based on at least two first image elements and at least two second image elements is N; where N is an integer greater than 1, and M is a positive integer less than N.
[0187] In some embodiments, the image generation module 1720 is configured to generate mask images corresponding to the at least two image elements respectively, the mask images being used to distinguish the image elements and other regions besides the image elements in the original image; and to determine the image element corresponding to the selected mask image in the mask images corresponding to the at least two image elements as the at least one first image element to be replaced.
[0188] In some embodiments, the image generation module 1720 is used to call the image segmentation model in parallel through at least two interfaces corresponding to the image segmentation model, wherein each of the at least two interfaces corresponding to the image segmentation model is used to call the image segmentation model to generate a mask image corresponding to one of the at least two image elements.
[0189] In some embodiments, the image generation module 1720 is configured to obtain at least two replacement elements from an image element library for replacing the first image element, and select one or more replacement elements from the at least two replacement elements as the second image element; or, obtain a replacement image uploaded by a user, and use image elements identified from the replacement image as the second image element.
[0190] In some embodiments, the text generation model is deployed on at least two GPUs.
[0191] In some embodiments, the image generation module 1720 is used to input the original image into a content detection model to obtain the retained original image output by the content detection model. The content detection model is used to retain the original image if the original image meets a first condition, and to clean the original image if the original image does not meet the first condition. The first condition is the condition for the content detection model to filter the original image. The retained original image is then input into an item recognition model to obtain the at least two image elements output by the item recognition model.
[0192] In some embodiments, the emoji generation module 1740 is configured to, for the i-th edited image among the at least one edited images, obtain adjustment information for the descriptive text of the i-th edited image, the adjustment information including at least one of the following: color, font, text content, display position, display size, where i is a positive integer; adjust the descriptive text of the i-th edited image based on the adjustment information to obtain the adjusted descriptive text of the i-th edited image; and add the adjusted descriptive text of the i-th edited image to the i-th edited image to obtain the i-th emoji image.
[0193] The technical solution provided in this application uses an image generation model to modify image elements in the acquired original image to obtain an edited image. A text generation model is then used to generate descriptive text that matches the content of the edited image. Finally, the descriptive text is added to the edited image to obtain an emoji image. The image generation in this application is no longer a simple image-to-image or text-to-image conversion. This application, on the one hand, uses an image generation model to edit image elements in the original image; on the other hand, it uses a text generation model to generate descriptive text that matches the edited image. The final emoji image is obtained by adding the descriptive text to the edited image. Therefore, this application combines automatic image editing and automatic text generation, enriching the image generation methods.
[0194] Furthermore, the image editing and text generation processes are performed sequentially. Image editing is performed first to obtain the edited image. Then, text generation is performed on the edited image to ensure that the descriptive text in the final emoji image matches the edited image. In other words, the text content and image content in the emoji images generated by this application match, resulting in better image generation quality.
[0195] Please refer to Figure 18, which shows a block diagram of an image generation apparatus according to another embodiment of this application. This apparatus has the function of implementing the image generation method described above; the function can be implemented in hardware or by hardware executing corresponding software. This apparatus can be the terminal device described above, or it can be installed within a terminal device. As shown in Figure 18, the apparatus 1800 may include a display module 1810 and an adding module 1820.
[0196] Display module 1810 is used to display an original image, which includes at least two image elements.
[0197] The display module 1810 is further configured to display at least one edited image in response to an editing operation on at least one of the at least two image elements, wherein at least one image element has changed compared to the original image.
[0198] The adding module 1820 is used to add descriptive text of the first edited image to the first edited image after selecting the first edited image from the at least one edited image, so as to obtain an emoticon image corresponding to the first edited image. The descriptive text of the first edited image is automatically generated text that conforms to the image content of the first edited image.
[0199] In some embodiments, the display module 1810 is configured to display marker information corresponding to the at least one image element, the marker information being used to indicate the image element; in response to a selection operation of the marker information corresponding to a first image element, a mask image corresponding to the first image element is displayed, the mask image being used to distinguish the first image element and other regions besides the first image element in the original image, the first image element being one or at least two of the at least one image element to be replaced; in response to a selection operation of a second image element, the at least one edited image is displayed, the second image element being used to replace the first image element.
[0200] In some embodiments, the display module 1810 is configured to display at least two replacement elements, the at least two replacement elements being elements in an image element library used to replace the first image element; in response to a selection operation for one or at least two of the at least two replacement elements, the selected replacement element is used as the selected second image element; or, in response to an operation of uploading a replacement image in an input field, the image elements included in the replacement image are used as the selected second image element.
[0201] In some embodiments, the adding module 1820 is configured to display descriptive text of the first edited image to be edited on the first edited image; in response to an adjustment operation on the descriptive text of the first edited image, display the adjusted descriptive text obtained after adjusting the descriptive text of the first edited image; and in response to a confirmation operation on the adjustment operation, add the adjusted descriptive text on the first edited image to obtain an emoji image corresponding to the first edited image.
[0202] The technical solution provided in this application allows for image editing through editing operations on image elements, resulting in an edited image. Furthermore, descriptive text matching the content of the edited image is automatically generated. Finally, the descriptive text is added to the edited image to obtain an emoji image. Therefore, this application combines manual image editing with automatic text generation to generate emoji images, enriching the image generation methods.
[0203] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0204] Please refer to Figure 19, which shows a structural block diagram of a computer device 1900 provided in one embodiment of this application. The computer device 1900 can be any electronic device with data computing, processing, and storage functions. The computer device 1900 can be used to implement the image generation method provided in the above embodiments.
[0205] Typically, computer device 1900 includes a processor 1901 and a memory 1902.
[0206] Processor 1901 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1901 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1901 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1901 may integrate a GPU, which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0207] The memory 1902 may include one or more computer-readable storage media, which may be non-transitory. The memory 1902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1902 is used to store a computer program configured to be executed by one or more processors to implement the image generation method described above.
[0208] Those skilled in the art will understand that the structure shown in FIG19 does not constitute a limitation on the computer device 1900, and may include more or fewer components than shown, or combine certain components, or employ different component arrangements.
[0209] In an exemplary embodiment, a computer-readable storage medium is also provided, wherein a computer program is stored in the storage medium, and the computer program, when executed by a processor, implements the above-described image generation method. Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0210] In an exemplary embodiment, a computer program product is also provided, the computer program product including a computer program stored in a computer-readable storage medium. A processor of a terminal device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the terminal device to perform the image generation method described above.
[0211] It should be noted that the collection and processing of relevant data (including original images, editing operations, etc.) in this application should strictly comply with the requirements of relevant national laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of laws and regulations and the authorization of the personal information subject.
[0212] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0213] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. An image generation method, the method being executed by a computer device, the method comprising: Obtain the original image, which includes at least two image elements; The original image is input into an image generation model to obtain at least one edited image output by the image generation model, wherein the edited image has at least one changed image element compared to the original image; The at least one edited image is input into a text generation model to obtain descriptive text corresponding to the at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image; For the at least one edited image, add the corresponding descriptive text to the edited image to obtain at least one emoji image.
2. The method of claim 1, wherein, The step of inputting the original image into the image generation model to obtain at least one edited image output by the image generation model includes: From the at least two image elements included in the original image, determine at least one first image element to be replaced; Determine a second image element to replace the first image element; The original image, the first image element, and the second image element are input into the image generation model to obtain the at least one edited image output by the image generation model.
3. The method of claim 2, wherein, The method further includes: When there are at least two edited images, the image generation model is called in parallel through at least two interfaces corresponding to the image generation model, wherein each of the at least two interfaces corresponding to the image generation model is used to call the image generation model to generate an edited image.
4. The method according to claim 2 or 3, wherein, When there is only one first image element, there is also only one second image element. The number of edited images generated based on one first image element and one second image element is N. When there are at least two first image elements, there are also at least two second image elements. The number of edited images generated based on one first image element and one second image element is M, and the number of edited images generated based on at least two first image elements and at least two second image elements is N. Where N is an integer greater than 1, and M is a positive integer less than N.
5. The method according to any one of claims 2 to 4, wherein, Determining at least one first image element to be replaced from the at least two image elements included in the original image includes: Generate mask images corresponding to the at least two image elements respectively, the mask images being used to distinguish the image elements and other regions besides the image elements in the original image; In the mask images corresponding to the at least two image elements respectively, the image element corresponding to the selected mask image is determined as the at least one first image element to be replaced.
6. The method of claim 5, wherein, The method further includes: The image segmentation model is invoked in parallel through at least two interfaces corresponding to the image segmentation model. Each of the at least two interfaces corresponding to the image segmentation model is used to invoke the image segmentation model to generate a mask image corresponding to one of the at least two image elements.
7. The method according to any one of claims 2 to 6, wherein, Determining the second image element to replace the first image element includes: Obtain at least two replacement elements from the image element library to replace the first image element, and select one or at least two replacement elements from the at least two replacement elements as the second image element; or, Obtain a replacement image uploaded by the user, and use the image elements identified from the replacement image as the second image element.
8. The method according to any one of claims 1 to 7, wherein, The text generation model is deployed on at least two GPUs.
9. The method according to any one of claims 1 to 8, wherein, After acquiring the original image, the process also includes: The original image is input into the content detection model to obtain the retained original image output by the content detection model. The content detection model is used to retain the original image if the original image meets a first condition, and to clean the original image if the original image does not meet the first condition. The first condition is the condition for the content detection model to filter the original image. The original image that is retained is input into the object recognition model to obtain the at least two image elements output by the object recognition model.
10. The method according to any one of claims 1 to 9, wherein, For the at least one edited image, adding the corresponding descriptive text to the edited image to obtain at least one emoji image includes: For the i-th edited image among the at least one edited images, obtain adjustment information for the descriptive text of the i-th edited image, the adjustment information including at least one of the following: color, font, text content, display position, display size, where i is a positive integer; The description text of the i-th edited image is adjusted based on the adjustment information to obtain the adjusted description text of the i-th edited image; Add the adjusted description text of the i-th edited image to the i-th edited image to obtain the i-th emoji image.
11. An image generation method, the method being performed by a computer device, the method comprising: Display the original image, which includes at least two image elements; In response to an editing operation on at least one of the at least two image elements, at least one edited image is displayed, wherein at least one image element has changed compared to the original image; After selecting a first edited image from the at least one edited image, a descriptive text for the first edited image is added to the first edited image to obtain an emoji image corresponding to the first edited image. The descriptive text for the first edited image is automatically generated text that matches the image content of the first edited image.
12. The method of claim 11, wherein, The response to an editing operation on at least one of the at least two image elements, displaying at least one edited image, includes: Displaying the marking information corresponding to each of the at least one image element, the marking information being used to indicate the image element; In response to a selection operation for the marker information corresponding to a first image element, a mask image corresponding to the first image element is displayed. The mask image is used to distinguish the first image element from other regions besides the first image element in the original image. The first image element is one or at least two of the at least one image element to be replaced. In response to the selection of a second image element, the at least one edited image is displayed, wherein the second image element is used to replace the first image element.
13. The method of claim 12, wherein, Before displaying the at least one edited image, the method further includes: Display at least two replacement elements, which are elements from an image element library used to replace the first image element; in response to a selection operation on one or both of the at least two replacement elements, the selected replacement element is used as the selected second image element; or... In response to the operation of uploading a replacement image in the input field, the image elements included in the replacement image are selected as the second image element.
14. The method according to any one of claims 11 to 13, wherein, The step of adding descriptive text to the first edited image to obtain an emoji image corresponding to the first edited image includes: Display descriptive text of the first edited image to be edited on the first edited image; In response to the adjustment operation of the descriptive text for the first edited image, the adjusted descriptive text obtained after adjusting the descriptive text for the first edited image is displayed; In response to a confirmation operation for the adjustment operation, the adjusted descriptive text is added to the first edited image to obtain an emoji image corresponding to the first edited image.
15. An image generation apparatus, the apparatus comprising: An image acquisition module is used to acquire an original image, wherein the original image includes at least two image elements; An image generation module is used to input the original image into an image generation model to obtain at least one edited image output by the image generation model, wherein the edited image has at least one image element changed compared to the original image; A text generation module is used to input the at least one edited image into a text generation model to obtain descriptive text corresponding to the at least one edited image output by the text generation model, wherein the descriptive text conforms to the image content of the edited image; An emoji generation module is used to add descriptive text corresponding to the edited image to the edited image to obtain at least one emoji image.
16. An image generation apparatus, the apparatus comprising: A display module is used to display an original image, which includes at least two image elements; The display module is further configured to, in response to an editing operation on at least one of the at least two image elements, display at least one edited image, wherein the edited image has at least one image element changed compared to the original image; An adding module is used to add descriptive text to the first edited image after selecting a first edited image from the at least one edited image, thereby obtaining an emoji image corresponding to the first edited image. The descriptive text of the first edited image is automatically generated text that conforms to the image content of the first edited image.
17. A computer device comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the image generation method as claimed in any one of claims 1 to 10, or to implement the image generation method as claimed in any one of claims 11 to 14.
18. A computer-readable storage medium storing a computer program, the computer program being loaded and executed by a processor to implement the image generation method as claimed in any one of claims 1 to 10, or to implement the image generation method as claimed in any one of claims 11 to 14.
19. A computer program product comprising a computer program loaded and executed by a processor to implement the image generation method as claimed in any one of claims 1 to 10, or to implement the image generation method as claimed in any one of claims 11 to 14.
Citation Information
Patent Citations
Image content regeneration method and device, equipment and storage medium
CN116740210A
Image editing method and device, equipment, storage medium and program product
CN117611709A
Image generation method and related device
CN118537447A
Airbag cushion Protector
KR1020240157415A
Meme generation method and apparatus, and device and medium
WO2021169134A1