Image generation method and device, equipment and medium
By receiving the text information input by the user and generating corresponding images, the poor image quality problem caused by object obfuscation in the prior art is solved, and higher quality image generation is achieved.
Patent Information
- Application Number
- CN202510147009.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art may easily lead to object confusion and poor image quality when generating an image including at least two objects.
By receiving the text information input by the user, a background image is generated, and a first image is generated based on the user's input to the background image, and finally a second image corresponding to the object information is generated based on the first image and the background image.
It effectively avoids object confusion and improves image quality, so that users can generate corresponding different objects when inputting in different areas on the background image.
Smart Images

Figure CN119991854A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of image processing technology, and specifically relates to an image generation method, device, equipment and medium. Background Art
[0002] With the development of image technology, large model generation technology has gradually penetrated into people's daily lives. In related technologies, for a specific object, users can use multiple pre-collected pictures of the object to train an object generation model for generating the object; using the object generation model, combined with some style templates or text templates, images with different styles and including the object can be generated.
[0003] However, in the related art, when an image including at least two objects is generated, there may be a situation where at least two objects are confused, resulting in poor image quality. Summary of the invention
[0004] The purpose of the embodiments of the present application is to provide an image generation method, device, equipment and medium that can solve the problem of poor image quality.
[0005] In a first aspect, an embodiment of the present application provides an image generation method, comprising:
[0006] Receiving text information input by a user, wherein the text information includes object information and scene description information;
[0007] Generate a background image based on the scene description information;
[0008] Receive a first input of a background image from a user;
[0009] In response to a first input, generating a first image;
[0010] Based on the first image and the background image, a second image corresponding to the object information is generated.
[0011] In a second aspect, an embodiment of the present application provides an image generating device, including:
[0012] A first receiving module, configured to receive text information input by a user, wherein the text information includes object information and scene description information;
[0013] A first generating module, used for generating a background image according to the scene description information;
[0014] A second receiving module, used for receiving a first input of a background image by a user;
[0015] A second generating module, configured to generate a first image in response to the first input;
[0016] The third generating module is used to generate a second image corresponding to the object information based on the first image and the background image.
[0017] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the program or instructions are executed by the processor, the steps of the image generation method provided in the embodiment of the present application are implemented.
[0018] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the image generation method provided in the embodiment of the present application are implemented.
[0019] In a fifth aspect, an embodiment of the present application provides a chip, comprising a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run a program or instruction to implement the steps of the image generation method provided in the embodiment of the present application.
[0020] In a sixth aspect, an embodiment of the present application provides a computer program product, which is stored in a storage medium and is executed by at least one processor to implement the steps of the image generation method provided in the embodiment of the present application.
[0021] In an embodiment of the present application, by receiving text information input by a user, wherein the text information includes object information and scene description information; generating a background image according to the scene description information; receiving a first input of the background image by the user; generating a first image in response to the first input; and generating a second image corresponding to the object information based on the first image and the background image. In this way, an image corresponding to the object information included in the text information input by the user can be generated based on the image generated by the user's input of the background image and the background image, so that when the user inputs in different areas on the background image, different objects can be generated in different areas of the background image, thereby avoiding confusion between different objects and improving image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flowchart of the image generation method provided in the embodiment of the present application;
[0023] Figure 2 is a comparative schematic diagram of the images generated provided in the embodiments of the present application;
[0024] Figure 3 is a schematic diagram of the structure of an image generating device provided in an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0026] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0028] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0029] The following, in conjunction with the accompanying drawings, describes in detail the image generation method, device, equipment and medium provided in the embodiments of the present application through specific embodiments and their application scenarios.
[0030] Figure 1 : is a flow chart of an image generation method provided in an embodiment of the present application. The image generation method may include:
[0031] Step 101: receiving text information input by a user, wherein the text information includes object information and scene description information;
[0032] Exemplarily, the text information input by the user is: a cat and a dog are surfing; then the objects included in the text information are the cat and the dog, and the scene description information is surfing.
[0033] Step 102: Generate a background image according to the scene description information;
[0034] The embodiment of the present application does not limit the method used to generate the background image according to the scene description information, and any available method can be applied to the embodiment of the present application.
[0035] Step 103: receiving a first input of a background image from a user;
[0036] In some possible implementations of the embodiment of the present application, the first input in the embodiment of the present application is used to trigger the generation of an image. The first input in the embodiment of the present application can be a circle selection input of a certain area in the background image, or a click input of a certain point in the background image.
[0037] Step 104: generating a first image in response to a first input;
[0038] In some possible implementations of the embodiments of the present application, in step 104, the area circled by the first input may be filled with a first color to obtain a first image, and the area within a preset size range centered on the point clicked by the user may be filled with the first color to obtain a first image. The first color may be set according to actual needs.
[0039] In some possible implementations of the embodiments of the present application, step 104 may include: filling the minimum circumscribed rectangle of the first input circled area with a first color to obtain a first image.
[0040] Exemplarily, assuming that the minimum horizontal coordinate of the area selected by the user is Iregion_x1, the maximum horizontal coordinate is Iregion_x2, the minimum vertical coordinate is Iregion_y1, and the maximum vertical coordinate is Iregion_y2, then the minimum circumscribed rectangle of the area selected by the first input is a rectangular area formed by four points (Iregion_x1, Iregion_y2), (Iregion_x1, Iregion_y1), (Iregion_x2, Iregion_y1) and (Iregion_x2, Iregion_y2).
[0041] For example, the minimum bounding rectangle of the area selected by the first input can be determined by the following formula (1):
[0042]
[0043] Among them, in formula (1), Irect=IregionpYnew1:Ynew2,Xnew1:Xnew2] represents a rectangular area composed of four points (Xnew1, Ynew1), (Xnew1, Ynew2), (Xnew2, Ynew1) and (Xnew2, Ynew2), Iregion_x is the horizontal coordinate list of the area Iregion selected by the user, and Iregion_y is the vertical coordinate list of the area Iregion selected by the user.
[0044] In an embodiment of the present application, the first image is obtained by filling the minimum circumscribed rectangle of the area circled by the first input with the first color, which can avoid the generated image having a serious sense of pasting due to the user's selection of too small an area or a different size ratio between the selected area and the object.
[0045] In some possible implementations of the embodiments of the present application, step 104 may include: filling the minimum bounding rectangle of the first input selected area with a first color, and blurring the edge of the minimum bounding rectangle filled with the first color to obtain a first image.
[0046] In some possible implementations of the present application, the width, height and center point of the minimum bounding rectangle can be obtained, and a certain range can be expanded with the center point as the center. The response value within this range is set to the highest value 1 to ensure that the object is generated intact in this area. The specific process is shown in the following formula (2):
[0047]
[0048] In formula (2), Ycenter is the vertical coordinate of the center point of the minimum circumscribed rectangle, Xcenter is the horizontal coordinate of the center point, a and b are expansion coefficients determined by the width and height of the minimum circumscribed rectangle obtained from experiments, H is the height of the minimum circumscribed rectangle, and W is the width of the minimum circumscribed rectangle.
[0049] Irect[Ycenter-a*H:Ycenter+a*H,Xcenter-a*W:Xcenter+a*W]=1 is used to set the response value in the rectangular area composed of four points (Xcenter-b*W, Ycenter-a*H), (Xcenter-b*W, Ycenter+a*H), (Xcenter+a*W, Ycenter-a*H) and (Xcenter+a*W, Ycenter+a*H) to 1. Blur function is the Gaussian blur operation in image processing. Irect_outer is the result of Gaussian blur operation on rectangle Irect.
[0050] The first object in the embodiment of the present application may be referred to as a mask image.
[0051] In the embodiment of the present application, by blurring the edge of the minimum circumscribed rectangle of the area circled by the first input, it is possible to ensure that the generated image of the object is naturally connected with the background, thereby improving the image quality.
[0052] Step 105: Generate a second image corresponding to the object information based on the first image and the background image.
[0053] In some possible implementations of the embodiments of the present application, step 105 may include: generating an image of the object corresponding to the object information in a first area corresponding to the first image on the background image using the object generation model and the first image.
[0054] In some possible implementations of the embodiments of the present application, the object generation model in the embodiments of the present application can be a low-rank adaptation (LoRA) model, wherein LoRA is a fine-tuning technology for a large language model (LLM), which aims to reduce model parameters by introducing a low-rank matrix and reduce fine-tuning costs while maintaining the performance of the model.
[0055] In some possible implementations of the embodiments of the present application, for each of the multiple objects, multiple images of the object can be used to train a LoRA model for generating an image of the object. When the image of the object is generated using the object generation model, the object generation model corresponding to the object is used for generation. The embodiments of the present application do not elaborate on the process of training the LoRA model for generating the object image, and the specific process can be referred to the description in the related art.
[0056] In some possible implementations of the embodiments of the present application, step 105 may include: in the first area, using the attention map of the object in the object generation model and the first image, generating an image of the object corresponding to the object information.
[0057] In some possible implementations of the embodiments of the present application, a corresponding object generation model can be trained for a specific object, and the attention mechanism in the object generation model builds a bridge between the image samples and the text labels in the training data; the attention mechanism is a technology that allows the model to dynamically focus on different parts when processing input data, and by calculating the weights of different parts of the input data, the model can process the input information more flexibly. The attention mechanism calculates the attention, that is, the response value, between the image sample and the text label, thereby ensuring that the generated image is consistent with the text label. Among them, the attention mechanism is shown in the following formula (3):
[0058]
[0059] In formula (3), Attention(Q,K,V) is the response value, Q is the hidden layer image feature of a layer in the diffusion model backbone network (Unet) weight, K and V are the extracted text features, d is the feature dimension, QK T Characterize the similarity between image features and text features, and the softmax function is used to convert the result into a probability distribution, indicating the correlation between K and Q.
[0060] It can be seen from the above formula that, given an image sample and its corresponding text label, the attention of each text element in the text label to each image area of the image sample during the training process can be different. By visualizing the above response values, an attention map is obtained. By visualizing the attention map, the response value of each word in each input label of the original image in the target area of the original image during the training process can be obtained. If the area on the visualized attention map is black, the response value is 0, indicating that the word is not activated on the target image; if the area on the visualized attention map is white, the response value is 1, indicating that the word is activated on the target image.
[0061] In some possible implementations of the embodiments of the present application, when generating an image of an object, the mask image of the area corresponding to the object is weighted with the attention map in the object generation model corresponding to the object to increase the attention weight of the object in the area corresponding to the object, while reducing the attention weights of other areas. Specifically, for each of the multiple objects, the attention map corresponding to the object is calculated, and the attention map reflects the activation degree of the text label in the image semantics during the image generation process.
[0062] For object i, by improving the mask image M of its corresponding area i The attention weight of the area is increased, and the attention weight outside the area is reduced to highlight the response probability of the object. When the image of object i is generated in the area corresponding to object i, the image of object i is generated in the area corresponding to object i through iterative calculation.
[0063] In some possible implementations of the embodiments of the present application, step 105 may include: in the first region, iteratively calculating the attention map of the object in the object generation model corresponding to the object and the first image in the first region, and when the number of iterations reaches a set number, generating an image of the object in the first region. The iterative calculation process is shown in the following formula (4):
[0064]
[0065] In formula (4), Attn i_k+1 is the attention map of object i obtained after calculation at the kth iteration; Attn i_k Calculate the attention map of the previous object i for the kth iteration; M i is the mask image corresponding to object i, that is, in the embodiment of the present application, the response value of the area corresponding to object i is 1, and the response values of other areas are 0; Attn j≠i is the attention map of other objects; a is the empirical coefficient corresponding to the background fusion degree; M j is the mask image corresponding to other objects.
[0066] Molecule Attn i_k *M i There is an activation value only in the area corresponding to object i, and the response values of other areas are forced to be set to 0; Attn j≠i *(1-M j ) represents the attention value of other objects in the background area.
[0067] When the number of iterations reaches the set number, the attention map of the object obtained at this time is the image of the object.
[0068] In the embodiment of the present application, the area can be well connected with the background without affecting the generation of the object in its corresponding area, avoiding direct truncation when the object exceeds the area selected by the user, so that the generated object can smoothly transition in the image.
[0069] Figure 2 is a comparative schematic diagram of the generated images provided in the embodiment of the present application. Figure 2 The image on the left is an image of a cat and a dog generated using related technology. Figure 2 The image on the right side of the figure is an image of a cat and a dog generated by the image generation method provided by the embodiment of the present application. Figure 2 It can be seen that there is confusion between cats and dogs in the images including cats and dogs generated using the relevant technology, while there is no confusion between cats and dogs in the images including cats and dogs generated using the image generation method provided in the embodiment of the present application, and cats and dogs can be clearly reflected.
[0070] In an embodiment of the present application, by receiving text information input by a user, wherein the text information includes object information and scene description information; generating a background image according to the scene description information; receiving a first input of the background image by the user; generating a first image in response to the first input; and generating a second image corresponding to the object information based on the first image and the background image. In this way, an image corresponding to the object information included in the text information input by the user can be generated based on the image generated by the user's input of the background image and the background image, so that when the user inputs in different areas on the background image, different objects can be generated in different areas of the background image, thereby avoiding confusion between different objects and improving image quality.
[0071] The image generation method provided in the embodiment of the present application can be executed by an image generation device. In the embodiment of the present application, the image generation device provided in the embodiment of the present application is described by taking the image generation method executed by the image generation device as an example.
[0072] Figure 3 300 is a schematic diagram of the structure of an image generating device provided in an embodiment of the present application. The image generating device 300 may include:
[0073] A first receiving module 301 is used to receive text information input by a user, wherein the text information includes object information and scene description information;
[0074] A first generating module 302, used to generate a background image according to the scene description information;
[0075] The second receiving module 303 is used to receive a first input of a background image by a user;
[0076] A second generating module 304, configured to generate a first image in response to the first input;
[0077] The third generating module 305 is used to generate a second image corresponding to the object information based on the first image and the background image.
[0078] In an embodiment of the present application, by receiving text information input by a user, wherein the text information includes object information and scene description information; generating a background image according to the scene description information; receiving a first input of the background image by the user; generating a first image in response to the first input; and generating a second image corresponding to the object information based on the first image and the background image. In this way, an image corresponding to the object information included in the text information input by the user can be generated based on the image generated by the user's input of the background image and the background image, so that when the user inputs in different areas on the background image, images of different objects can be generated in different areas of the background image, thereby avoiding confusion between different objects and improving image quality.
[0079] In some possible implementations of the embodiment of the present application, the second generating module 304 is specifically used for:
[0080] The minimum circumscribed rectangle of the first input selected area is filled with a first color to obtain a first image.
[0081] In the embodiment of the present application, the first image is obtained by filling the color of the minimum circumscribed rectangle of the area circled by the user, which can avoid the serious sense of pasting in the generated image caused by the user's selected area being too small or the selected area being in a different size ratio with the object.
[0082] In some possible implementations of the embodiment of the present application, the second generating module 304 is specifically used for:
[0083] The minimum circumscribed rectangle of the first input circled area is filled with a first color, and the edge of the minimum circumscribed rectangle filled with the first color is blurred to obtain a first image.
[0084] In the embodiment of the present application, by blurring the edge of the minimum circumscribed rectangle of the area circled by the user, it is possible to ensure that the generated image of the object is naturally connected with the background, thereby improving the image quality.
[0085] In some possible implementations of the embodiment of the present application, the third generating module 305 is specifically used for:
[0086] In a first area corresponding to the first image on the background image, an image of the object corresponding to the object information is generated using the object generation model and the first image.
[0087] In some possible implementations of the embodiment of the present application, the third generating module 305 is specifically used for:
[0088] In the first region, an image of the object corresponding to the object information is generated using an attention map of the object in the object generation model and the first image.
[0089] In the embodiment of the present application, the area can be well connected with the background without affecting the generation of the object in its corresponding area, avoiding direct truncation when the object exceeds the area selected by the user, so that the generated object can smoothly transition in the image.
[0090] The image generating device in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or other devices other than a terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.
[0091] The image generation device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0092] The image generation device provided in the embodiment of the present application can achieve Figure 1 to Figure 2 To avoid repetition, the various processes implemented in the image generation method embodiment are not described here.
[0093] Alternatively, if Figure 4 As shown, an embodiment of the present application further provides an electronic device 400, including a processor 401 and a memory 402, wherein the memory 402 stores programs or instructions that can be executed on the processor 401, and when the program or instructions are executed by the processor 401, the various steps of the image generation method embodiment provided in the embodiment of the present application are implemented, and the same technical effect can be achieved. To avoid repetition, they are not described here.
[0094] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application.
[0095] The electronic device 500 includes but is not limited to: a radio frequency unit 501, a network module 502, an audio output unit 503, an input unit 504, a sensor 505, a display unit 506, a user input unit 507, an interface unit 508, a memory 509, and a processor 510.
[0096] Those skilled in the art will appreciate that the electronic device 500 may also include a power source (such as a battery) for supplying power to various components, and the power source may be logically connected to the processor 510 through a power management system, thereby implementing functions such as managing charging, discharging, and power consumption management through the power management system. Figure 5 The electronic device structure shown in the figure does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently, which will not be described in detail here.
[0097] The processor 510 is used to: receive text information input by a user, wherein the text information includes object information and scene description information; generate a background image according to the scene description information;
[0098] A user input unit 507, used to receive a first input of a user on a background image;
[0099] The processor 510 is further configured to: generate a first image in response to a first input; and generate a second image corresponding to the object information based on the first image and the background image.
[0100] In an embodiment of the present application, by receiving text information input by a user, wherein the text information includes object information and scene description information; generating a background image according to the scene description information; receiving a first input of the background image by the user; generating a first image in response to the first input; and generating a second image corresponding to the object information based on the first image and the background image. In this way, an image corresponding to the object information included in the text information input by the user can be generated based on the image generated by the user's input of the background image and the background image, so that when the user inputs in different areas on the background image, different objects can be generated in different areas of the background image, thereby avoiding confusion between different objects and improving image quality.
[0101] In some possible implementations of the embodiments of the present application, the processor 510 is specifically configured to:
[0102] The minimum circumscribed rectangle of the first input selected area is filled with a first color to obtain a first image.
[0103] In the embodiment of the present application, the first image is obtained by filling the color of the minimum circumscribed rectangle of the area circled by the user, which can avoid the serious sense of pasting in the generated image caused by the user's selected area being too small or the selected area being in a different size ratio with the object.
[0104] In some possible implementations of the embodiments of the present application, the processor 510 is specifically configured to:
[0105] The minimum circumscribed rectangle of the first input circled area is filled with a first color, and the edge of the minimum circumscribed rectangle filled with the first color is blurred to obtain a first image.
[0106] In the embodiment of the present application, by blurring the edge of the minimum circumscribed rectangle of the area circled by the user, it is possible to ensure that the generated image of the object is naturally connected with the background, thereby improving the image quality.
[0107] In some possible implementations of the embodiments of the present application, the processor 510 is specifically configured to:
[0108] In a first area corresponding to the first image on the background image, an image of the object corresponding to the object information is generated using the object generation model and the first image.
[0109] In some possible implementations of the embodiments of the present application, the processor 510 is specifically configured to:
[0110] In the first region, an image of the object corresponding to the object information is generated using an attention map of the object in the object generation model and the first image.
[0111] In the embodiment of the present application, the area can be well connected with the background without affecting the generation of the object in its corresponding area, avoiding direct truncation when the object exceeds the area selected by the user, so that the generated object can smoothly transition in the image.
[0112] It should be understood that in the embodiment of the present application, the input unit 504 may include a graphics processor (GPU) 5041 and a microphone 5042, and the graphics processor 5041 processes the image data of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 506 may include a display panel 5061, and the display panel 5061 may be configured in the form of a liquid crystal display, an organic light emitting diode, etc. The user input unit 507 includes a touch panel 5071 and at least one of other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include two parts: a touch detection device and a touch controller. Other input devices 5072 may include, but are not limited to, a physical keyboard, function keys (such as a volume control button, a switch button, etc.), a trackball, a mouse, and a joystick, which will not be repeated here.
[0113] The memory 509 can be used to store software programs and various data. The memory 509 may mainly include a first storage area for storing programs or instructions and a second storage area for storing data, wherein the first storage area may store an operating system, an application program or instructions required for at least one function (such as a sound playback function, an image playback function, etc.), etc. In addition, the memory 509 may include a volatile memory or a non-volatile memory, or the memory 509 may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDRSDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synchronous link dynamic random access memory (SLDRAM) and a direct memory bus random access memory (DRRAM). The memory 509 in the embodiment of the present application includes but is not limited to these and any other suitable types of memory.
[0114] The processor 510 may include one or more processing units; optionally, the processor 510 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communication signals, such as a baseband processor. It is understandable that the modem processor may not be integrated into the processor 510.
[0115] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the various processes of the image generation method embodiment provided in the embodiment of the present application are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0116] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, and examples of computer-readable storage media include non-transitory computer-readable media, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.
[0117] An embodiment of the present application also provides a chip, which includes a processor and a communication interface, wherein the communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the image generation method embodiment provided in the embodiment of the present application, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0118] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0119] The embodiment of the present application also provides a computer program product, which is stored in a storage medium. The program product is executed by at least one processor to implement the various processes of the image generation method embodiment provided in the embodiment of the present application, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0120] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0121] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0122] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. An image generation method, characterized in that: The method comprises: Receiving text information input by a user, wherein the text information includes object information and scene description information; Generate a background image according to the scene description information; receiving a first input of the background image by a user; In response to the first input, generating a first image; Based on the first image and the background image, a second image corresponding to the object information is generated.
2. The method according to claim 1, characterized in that The generating of the first image comprises: The minimum circumscribed rectangle of the area selected by the first input is filled with a first color to obtain the first image.
3. The method according to claim 1, characterized in that: The generating of the first image comprises: The first image is obtained by filling the minimum circumscribed rectangle of the area circled by the first input with a first color and blurring the edge of the minimum circumscribed rectangle filled with the first color.
4. The method according to claim 1, characterized in that: The step of generating a second image corresponding to the object information based on the first image and the background image includes: In a first area on the background image corresponding to the first image, an image of the object corresponding to the object information is generated using an object generation model and the first image.
5. The method according to claim 1, characterized in that The first area corresponding to the first image on the background image generates an image of the object corresponding to the object information by using an object generation model and the first image, comprising: In the first area, an image of the object corresponding to the object information is generated using the attention map of the object in the object generation model and the first image.
6. An image generating device, characterized in that: The device comprises: A first receiving module, configured to receive text information input by a user, wherein the text information includes object information and scene description information; A first generating module, used for generating a background image according to the scene description information; A second receiving module, used for receiving a first input of the background image by a user; A second generating module, configured to generate a first image in response to the first input; The third generating module is used to generate a second image corresponding to the object information based on the first image and the background image.
7. The device according to claim 6, characterized in that The second generating module is specifically used for: The minimum circumscribed rectangle of the area selected by the first input is filled with a first color to obtain the first image.
8. The device according to claim 6, characterized in that The second generating module is specifically used for: The first image is obtained by filling the minimum circumscribed rectangle of the area circled by the first input with a first color and blurring the edge of the minimum circumscribed rectangle filled with the first color.
9. The device according to claim 6, characterized in that The third generation module is specifically configured to generate an image of the object corresponding to the object information by using an object generation model and the first image in a first area corresponding to the first image on the background image.
10. The device according to claim 9, characterized in that The third generation module is specifically used for: In the first area, an image of the object corresponding to the object information is generated using the attention map of the object in the object generation model and the first image.
11. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores a program or instruction that can be run on the processor, and when the program or instruction is executed by the processor, the steps of the image generation method according to any one of claims 1 to 5 are implemented.
12. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by a processor, the steps of the image generation method according to any one of claims 1 to 5 are implemented.