Picture generation method and device, storage medium and electronic equipment
By generating and adjusting pictures in the graffiti painting scene, using the picture generation model and combining user input, the problem of low creation efficiency in traditional graffiti methods is solved, and an efficient and convenient graffiti creation experience is achieved.
Patent Information
- Application Number
- CN202510565859.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional graffiti methods rely on physical media, limiting the size, color selection and modification possibility of graffiti works, resulting in low creative efficiency and may destroy the beauty of the original works.
In the graffiti painting scene, the user's picture generation description is determined, the initial preview picture is generated, and the reference graffiti sketch input by the user is monitored, the picture generation model is used to adjust the content, and the target picture is automatically generated, and the preview update display is achieved in combination with the user input.
It greatly breaks through the limitations of the traditional graffiti creation model, provides graffiti users with a convenient and efficient creative experience, and improves the efficiency and flexibility of graffiti creation.
Smart Images

Figure CN120495446A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a method, device, storage medium and electronic device for generating an image. Background Art
[0002] In the realm of artistic creation, graffiti, as a free, spontaneous, and creative form of expression, has long been beloved by people of all ages. It's not just a visual art form, but also a direct expression of personal emotions and thoughts. With the advancement of technology and the growing demand for diverse artistic creation, traditional graffiti has gradually revealed some limitations, particularly in terms of flexibility and creative expansion.
[0003] Traditional graffiti often relies on physical surfaces, such as walls, graffiti areas, or paper. The physical properties of these media limit the size, color choices, and potential for modification. For example, once paint is applied to a physical medium, modification is extremely difficult, often requiring a complex overlay or repainting process that is not only time-consuming but can also ruin the overall aesthetic of the original work. This makes traditional graffiti methods particularly inconvenient, significantly limiting the artist's efficiency. Summary of the Invention
[0004] The embodiments of the present application provide a picture generation method, device, storage medium, and electronic device, which can improve the efficiency of graffiti creation by graffiti artists.
[0005] In a first aspect, an embodiment of the present application provides a method for generating an image, comprising:
[0006] In a graffiti drawing scene, determining a drawing generation description of the user, and generating an initial preview drawing in a drawing generation preview area based on the drawing generation description;
[0007] monitoring a reference doodle sketch input by a user in a doodle canvas area;
[0008] Based on the reference graffiti sketch and the picture generation description, a picture generation model is used to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and a preview update display process is performed on the picture generation preview area based on the target generated picture.
[0009] In some embodiments, adjusting the content of the initial preview generated picture using a picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture includes:
[0010] Obtaining, through a picture generation model, a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture;
[0011] Determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch through a picture generation macro model, and performing sketch semantic encoding on the sketch line information, the sketch shape, the sketch outline information, and the content space layout information to obtain a sketch semantic vector;
[0012] The picture generation model performs picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector and the sketch semantic vector to obtain picture adjustment semantic information, and uses the picture adjustment semantic information to adjust the picture content of the initial preview generated picture to obtain the target generated picture.
[0013] In some embodiments, performing image adjustment inference processing based on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain image adjustment semantic information includes:
[0014] Mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space;
[0015] In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information;
[0016] Picture adjustment information is parsed based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initially previewed picture.
[0017] In some embodiments, the step of adjusting the content of the initial preview generated picture using the picture adjustment semantic information to obtain a target generated picture includes:
[0018] Acquire multiple layer vectors for the initial preview generated picture based on a layered generation architecture, and determine layer adjustment information for each of the layer vectors based on the picture adjustment semantic information;
[0019] Adjusting the layer vectors based on the layer reference parameters of the layer vectors and the layer adjustment information to obtain adjusted layer vectors;
[0020] The initial preview generated picture is encoded with picture content based on each of the adjustment layer vectors to obtain a target generated picture.
[0021] In some embodiments, generating an initial preview image in the image generation preview area based on the image generation description includes:
[0022] The picture generation description is converted into a text semantic vector using a picture generation model, and text-image processing is performed based on the text semantic vector to obtain an initial preview generated picture, which is displayed in a picture generation preview area.
[0023] In some embodiments, receiving a user input of a picture adjustment description;
[0024] Adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using a picture generation macro model to generate an adjusted generated picture;
[0025] Based on the adjusted generated picture, a preview update display process is performed on the picture generation preview area.
[0026] In some embodiments, the step of adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by the picture generation macro model to generate the adjusted generated picture includes:
[0027] Obtaining, through a picture generation model, an adjustment text semantic vector corresponding to the picture adjustment description and a preview image feature vector corresponding to the preview generated picture;
[0028] The picture generation model performs picture adjustment inference processing based on the adjustment text semantic vector and the preview image feature vector to obtain picture adjustment description semantic information, and uses the picture adjustment description semantic information to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
[0029] In a second aspect, an embodiment of the present application further provides a picture generating device, comprising:
[0030] A generating module, configured to determine a drawing generation description of the user in a graffiti drawing scene, and generate an initial preview drawing in a drawing generation preview area based on the drawing generation description;
[0031] a detection module, configured to detect a reference graffiti sketch input by a user in a graffiti canvas area;
[0032] An updating module is configured to adjust the content of the initial preview generated picture using a picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture, and perform preview update display processing on the picture generation preview area based on the target generated picture.
[0033] In some embodiments, the update module is specifically configured to:
[0034] Obtaining, through a picture generation model, a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture;
[0035] Determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch through a picture generation macro model, and performing sketch semantic encoding on the sketch line information, the sketch shape, the sketch outline information, and the content space layout information to obtain a sketch semantic vector;
[0036] The picture generation model performs picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector and the sketch semantic vector to obtain picture adjustment semantic information, and uses the picture adjustment semantic information to adjust the picture content of the initial preview generated picture to obtain the target generated picture.
[0037] In some embodiments, the update module is specifically configured to:
[0038] Mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space;
[0039] In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information;
[0040] Picture adjustment information is parsed based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initially previewed picture.
[0041] In some embodiments, the update module is specifically configured to:
[0042] Acquire multiple layer vectors for the initial preview generated picture based on a layered generation architecture, and determine layer adjustment information for each of the layer vectors based on the picture adjustment semantic information;
[0043] Adjusting the layer vectors based on the layer reference parameters of the layer vectors and the layer adjustment information to obtain adjusted layer vectors;
[0044] The initial preview generated picture is encoded with picture content based on each of the adjustment layer vectors to obtain a target generated picture.
[0045] In some embodiments, the generating module is specifically configured to:
[0046] The picture generation description is converted into a text semantic vector using a picture generation model, and text-image processing is performed based on the text semantic vector to obtain an initial preview generated picture, which is displayed in a picture generation preview area.
[0047] In some embodiments, the update module is specifically configured to:
[0048] receiving a picture adjustment description input by a user;
[0049] Adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using a picture generation macro model to generate an adjusted generated picture;
[0050] Based on the adjusted generated picture, a preview update display process is performed on the picture generation preview area.
[0051] In some embodiments, the update module is specifically configured to:
[0052] Obtaining, through a picture generation model, an adjustment text semantic vector corresponding to the picture adjustment description and a preview image feature vector corresponding to the preview generated picture;
[0053] The picture generation model performs picture adjustment inference processing based on the adjustment text semantic vector and the preview image feature vector to obtain picture adjustment description semantic information, and uses the picture adjustment description semantic information to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
[0054] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is run on a computer, the computer is enabled to execute the picture generation method provided in any embodiment of the present application.
[0055] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising a processor and a memory, wherein the memory has a computer program, and the processor is configured to execute the picture generation method provided in any embodiment of the present application by calling the computer program.
[0056] The technical solution provided by the embodiments of the present application determines a user's picture generation description in a graffiti drawing scene, generates an initial preview picture in a picture generation preview area based on the picture generation description, monitors a reference graffiti sketch input by the user in the graffiti canvas area, and uses a picture generation macro model to adjust the picture content of the initial preview picture based on the reference graffiti sketch and the picture generation description to obtain a target picture. The preview update display processing of the picture generation preview area is performed based on the target picture. The present application can automatically generate related pictures based on the picture generation description and reference graffiti sketch input by the user, greatly breaking through the limitations of traditional graffiti creation models and providing unprecedented convenience and efficiency for graffiti artists, thereby greatly improving the efficiency of graffiti artists' graffiti creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0058] Figure 1 Schematic diagram of an application scenario of the picture generation method provided in an embodiment of the present application.
[0059] Figure 2 A schematic diagram of the first operating interface provided in an embodiment of the present application.
[0060] Figure 3 A schematic diagram of the second operation interface provided in an embodiment of the present application.
[0061] Figure 4 Schematic diagram of the third operation interface provided in the embodiment of the present application
[0062] Figure 5 A schematic diagram of the fourth operating interface provided in an embodiment of the present application.
[0063] Figure 6 A schematic diagram of the fifth operating interface provided in an embodiment of the present application.
[0064] Figure 7 A schematic diagram of the structure of the image generation device provided in an embodiment of the present application.
[0065] Figure 8 This is a schematic diagram of the first structure of the electronic device provided in an embodiment of the present application.
[0066] Figure 9 A second structural diagram of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0068] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0069] In the realm of artistic creation, graffiti, as a free, spontaneous, and creative form of expression, has long been beloved by people of all ages. It's not just a visual art form, but also a direct expression of personal emotions and thoughts. With the advancement of technology and the growing demand for diverse artistic creation, traditional graffiti has gradually revealed some limitations, particularly in terms of flexibility and creative expansion.
[0070] Traditional graffiti often relies on physical surfaces, such as walls, graffiti areas, or paper. The physical properties of these media limit the size, color choices, and potential for modification. For example, once paint is applied to a physical medium, modification is extremely difficult, often requiring a complex overlay or repainting process that is not only time-consuming but can also ruin the overall aesthetic of the original work. This makes traditional graffiti methods particularly inconvenient, significantly limiting the artist's efficiency.
[0071] To improve the efficiency of graffiti artists, embodiments of the present application provide a method for generating an image. The method can be performed by the image generation device provided in embodiments of the present application, or an electronic device incorporating the image generation device. The image generation device can be implemented in hardware or software. The electronic device can be a smartphone, tablet computer, PDA, laptop computer, or desktop computer.
[0072] See also Figure 1 , Figure 1 This is a schematic diagram of the first process of the picture generation method provided in the embodiment of the present application. The specific process of the picture generation method provided in the embodiment of the present application can be as follows:
[0073] S1. In a graffiti drawing scene, determine a user's drawing generation description, and generate an initial preview drawing in a drawing generation preview area based on the drawing generation description.
[0074] The image generation description refers to the specific information or instructions provided by the user to generate a picture. This information may include elements such as the picture's content, style, color, and layout, and is usually expressed in the form of text, voice, gestures, or selections on a graphical interface. The specificity and detail of the description will affect the quality and accuracy of the generated picture. For example, when the image generation description is a text description, the text description may be: "A little girl in a red dress is flying a kite at the beach, with a blue sky, white clouds, and golden sand in the background." For another example, when the image generation description is a voice description, the voice description may be: "Draw a cool boy wearing sunglasses, driving a cool sports car." In this embodiment, a picture generation description input box can be provided to facilitate the user to enter text or voice through this input box.
[0075] The image generation preview area refers to the screen or interface area used to display the initial preview image generated based on the user's image generation description. This area is usually located in a prominent position on the user interface so that the user can clearly see the generated image and make adjustments or modifications as needed.
[0076] The initial preview image is a preliminary preview image generated by the system based on the user's image description. This image may not be the final version, but it clearly captures the main content and style of the user's description. Users can view this image in the image preview area and make further adjustments or modifications as needed.
[0077] In this embodiment, in the graffiti drawing scene, the user's drawing generation description is determined, and an initial preview drawing is generated in the drawing generation preview area based on the drawing generation description, so that the user can intuitively see the generated drawing effect.
[0078] For example, please refer to Figure 2 , Figure 2 This is a schematic diagram of the first operation interface provided in an embodiment of the present application. The schematic diagram illustrates the graffiti canvas area and the picture generation preview area, and provides a picture generation description input box below the above areas. The text "a lovely angel with an angel halo on her head" is entered in the input box. The initial preview generated based on the picture generation description can be as follows Figure 2 The picture in the generated preview area is shown.
[0079] S2. Monitoring a reference graffiti sketch input by the user in the graffiti canvas area.
[0080] The graffiti canvas area refers to the virtual creative space that appears on the screen of electronic devices (such as tablets, smartphones, computer monitors, etc.). It uses digital technology and software tools to provide users with creative elements such as brushes, colors, and textures, allowing users to freely draw graffiti sketches on the screen through touch, click, or external input devices (such as digital pens, mice, etc.).
[0081] The reference doodle sketch may be a hand-drawn, relatively rough sketch used to express the general outline or shape of the picture that the user wants to create.
[0082] In this embodiment, the electronic device monitors the reference graffiti sketch input by the user in the graffiti canvas area.
[0083] For example, please refer to Figure 3 , Figure 3 This is a schematic diagram of a second operation interface provided in an embodiment of the present application. In this schematic diagram, the graffiti canvas area displays a reference graffiti sketch input by the user.
[0084] S3. Based on the reference graffiti sketch and the picture generation description, the picture generation model is used to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and the picture generation preview area is previewed and updated based on the target generated picture.
[0085] The picture generation model is configured to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and perform preview update display processing on the picture generation preview area based on the target generated picture.
[0086] In this embodiment, based on the reference graffiti sketch and the picture generation description, the picture generation model is used to adjust the picture content of the initial preview generated picture to obtain the target generated picture, and the picture generation preview area is previewed and updated based on the target generated picture. In this embodiment, the reference graffiti sketch input by the user and the picture generation description are combined to adjust the picture content to automatically generate the target picture, providing the user with a physical and digital picture requirement input entrance, thereby providing the user with a broad creative space, and based on the input provided by the user, it can automatically generate a picture associated with the user input, which can greatly improve the graffiti creation efficiency and graffiti creation flexibility of the graffiti artist.
[0087] For example, please continue to refer to Figure 3 , Figure 3 In the graffiti canvas area, the reference graffiti sketch input by the user is displayed, and the text description "A cute angel with an angel halo on her head" is entered in the picture generation description input box. Based on the reference graffiti sketch and the picture generation description, a winged angel picture as shown in the picture generation preview area can be automatically generated.
[0088] For further information, please refer to Figure 4 ,exist Figure 3 Based on this, users can continue to draw in the graffiti canvas area. At this time, the reference graffiti sketch displayed in the graffiti canvas area is Figure 4 As shown, based on the reference doodle sketch and the picture generation description, a picture of an angel with spread wings as shown in the picture generation preview area can be automatically generated.
[0089] Furthermore, if the user is not satisfied with the preview generated picture displayed in the current picture generation preview area, he can click "Effect Refresh" to update the preview generated picture displayed in the picture preview generation area. Please refer to Figure 5 , Figure 5 The icon clicked by the middle finger pointer is the icon corresponding to "effect refresh", and the preview generated picture displayed in the picture preview generation area is updated. The effect refresh can be achieved by the electronic device re-performing step S3. When the user is satisfied with the preview generated picture displayed in the picture preview generation area, please refer to Figure 6 , you can click "Save as New Layer" to replace the reference sketch displayed in the doodle canvas area with the preview generated image. It is understood that users can continue to update the preview generated image through the effect refresh function and continue to save the image. I will not go into details here.
[0090] It should be understood that in this embodiment, the user can provide two key inputs: a reference doodle sketch, which may be a preliminary design drawn by the user to express their creative intent; and a generated image description, which is a more detailed and specific textual description that further defines the user's desired image style, colors, element layout, and other aspects. The doodle generation system utilizes a pre-trained large-scale computer model, the generated image model, to process these two inputs. Based on the requirements in the generated image description, the system generates a preliminary image preview, also known as the initial generated image. This initial generated image preview may not fully meet the user's expectations. Therefore, the system further adjusts and optimizes the image content based on the requirements in the generated image description and / or the image content in the reference doodle sketch. This may include changing colors, adjusting shapes, adding or removing elements, and so on, to ensure that the generated image better meets the user's creative intent. After these adjustments and optimizations, the system ultimately generates an image that meets the user's requirements, known as the target generated image.
[0091] In some embodiments, the training process of the image generation model can be as follows:
[0092] (1) obtaining a basic large language model, an image encoding and decoding model, and a multimodal fusion model, and creating an initial large picture generation model for a target generated picture generation prompt word generation scenario based on the basic large language model, the image encoding and decoding model, and the multimodal fusion model;
[0093] Among them, basic large language models include but are not limited to large models in the field of natural language processing (NLP), such as GPT-3, GPT-4, ChatGPT, BERT, and RoBERTa models.
[0094] The basic large language model is used to understand and parse the natural language instructions in the image generation description, extract key information, and convert it into a vector representation that the model can understand. Examples of basic large language models include BERT, the GPT series (such as GPT-3), and T5. These models are trained on large amounts of text data and possess powerful natural language understanding and generation capabilities.
[0095] Among them, the image encoding and decoding model includes two model parts according to the function. The image encoding model part is used to convert the reference graffiti sketch and the initial preview generated picture into a feature vector in a high-dimensional space so that the model can perform subsequent processing. The image decoding model part is used to convert the processed feature vector back to the image format to generate the target generated picture. For example, convolutional neural network (CNN): used for tasks such as image feature extraction and image classification. Variational autoencoder (VAE): learns data distribution by reconstructing data, which can be used for image generation and image restoration. Generative adversarial networks (GANs): generate realistic images through competitive learning, including variants such as conditional generative adversarial networks (cGANs).
[0096] The multimodal fusion model is used to fuse text vectors and image feature vectors to generate image content that matches the text description. The multimodal fusion model involves complex attention mechanisms and fusion strategies to ensure the effective interaction and integration of text and image information in the model.
[0097] (2) obtaining a sample reference graffiti sketch and a sample picture generation description, and annotating the sample reference graffiti sketch and the sample picture generation description to generate a picture label;
[0098] (3) using sample reference graffiti sketches and sample picture generation descriptions to perform at least one round of model training on the initial picture generation model;
[0099] (4) During the forward propagation training of the model, the initial picture is used to generate a large model to determine the predicted target generated picture based on the sample reference graffiti sketch and the sample picture generation description;
[0100] (5) During the model backpropagation training process, the image generation loss is determined based on the predicted target generation image and the target generation image label, and the image generation loss is used to adjust the model parameters of the initial image generation model until the initial image generation model completes model training and obtains the image generation model.
[0101] Specifically, a preset model loss calculation formula is used to calculate the image generation loss based on the predicted target generation image and the target generation image label. The preset model loss calculation formula can be a fit of one or more of the contrast loss calculation formula, cross entropy loss calculation formula, hinge loss calculation formula, and the like in related technologies.
[0102] Optionally, the model training termination conditions for generating the large model from the initial image may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset threshold, etc. Specific model training termination conditions can be determined based on actual conditions and are not specifically limited here.
[0103] During specific implementation, the present application is not limited by the execution order of the various steps described. If no conflict occurs, some steps can be performed in other orders or simultaneously.
[0104] As can be seen from the above, the picture generation method provided in the embodiment of the present application determines the user's picture generation description in a graffiti drawing scene, generates an initial preview generated picture in the picture generation preview area based on the picture generation description, monitors the reference graffiti sketch input by the user in the graffiti canvas area, and uses the picture generation macro model to adjust the picture content of the initial preview generated picture based on the reference graffiti sketch and the picture generation description to obtain a target generated picture, and then performs preview update and display processing on the picture generation preview area based on the target generated picture. In this application, relevant pictures can be automatically generated based on the picture generation description and reference graffiti sketch input by the user, which greatly breaks through the limitations of traditional graffiti creation models and brings unprecedented convenience and efficiency to graffiti artists, thereby greatly improving the efficiency of graffiti artists in graffiti creation.
[0105] In some embodiments, step S3 of "adjusting the content of the initial preview generated picture using the picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture" may include the following steps S31 to S33:
[0106] S31, obtaining a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture through the picture generation large model;
[0107] A text semantic vector is a mathematical representation of text (such as a sentence, a description, or an article) converted through an algorithm or model. It is typically a vector in a high-dimensional space. This vector captures the semantic information in the text, making similar texts closer in the vector space and dissimilar texts farther apart. This vector captures key information such as context, sentiment, and style.
[0108] For example, in this embodiment, a pre-trained text encoder (such as a CLIP text encoder or a Transformer model) can be used to convert the text description input by the user into a high-dimensional semantic vector, that is, the above-mentioned text semantic vector.
[0109] Among them, the initial preview image feature vector is a mathematical representation of an image (such as a painting, a photo of a scene or an object) converted through a certain image processing algorithm or deep learning model. It is also a vector in a high-dimensional space. This vector can capture key features in the image, such as color, texture, shape, edges, etc., as well as more advanced semantic information (such as the type of object or scene). It has a wide range of applications in computer vision, image recognition, image generation and other fields. For example, an image encoder (such as the encoding part of U-Net or ResNet) can be used to extract features from the initial preview generated picture to capture global style and local details.
[0110] In this embodiment, the picture generation model is used to obtain the text semantic vector corresponding to the picture generation description and the initial preview feature map vector corresponding to the initial preview generated picture. In this embodiment, the text description and the generated picture are associated in a certain mathematical representation for subsequent processing, analysis or comparison.
[0111] S32, determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch by generating a large model through the picture, and performing sketch semantic encoding on the sketch line information, sketch shape, sketch outline information, and content space layout information to obtain a sketch semantic vector;
[0112] Sketch line information refers to the direction, length, thickness, curvature, and other attributes of each line in a graffiti sketch, as well as the connections between the lines. For example, the line information of a hand-drawn circle may include the smoothness of the line, whether the thickness is uniform, and the starting and ending points of the line.
[0113] Sketch shapes refer to basic geometric forms or object outlines presented in a doodle sketch, such as circles, squares, triangles, etc., or more complex object shapes. For example, a hand-drawn cup may be an oval with a handle.
[0114] Sketch outline information refers to the boundary formed by the outermost lines in a doodle sketch, which defines the overall shape and scope of the sketch. For example, the outline information of a hand-drawn tree sketch may be composed of the outer edge lines of the trunk and crown.
[0115] Content spatial layout information refers to the spatial distribution and arrangement of elements (such as shapes, lines, and text) within a graffiti sketch, including their position, size, orientation, and relationships. For example, a graffiti sketch containing multiple objects, such as a house, a tree, and a person, would have their relative positions, sizes, and orientations constituting the content spatial layout information.
[0116] A sketch semantic vector is a mathematical representation of sketch features, such as line information, shape, outline information, and content spatial layout information, converted through an algorithm or model. It is typically a vector in a high-dimensional space. This vector captures the semantic information of the sketch, allowing similar sketches to be closer in vector space and dissimilar sketches to be farther apart. For example, a sketch semantic vector representing a "house" might be closer in vector space to another sketch that also represents a "house" but with slightly different details, but farther away from a sketch semantic vector representing a "car."
[0117] In this embodiment, a large-scale model is used to determine the sketch line information, sketch shape, sketch outline information, and content spatial layout information of a reference graffiti sketch. Sketch semantic encoding is then performed on these information to generate a sketch semantic vector. In this embodiment, the reference graffiti sketch is converted into a mathematical representation that is easy for a computer to process and compare, facilitating subsequent tasks such as sketch recognition, retrieval, generation, or editing. This sketch semantic vector allows computers to more easily understand graffiti sketches, enabling more efficient sketch processing and application.
[0118] S33. Perform image adjustment reasoning based on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector through the image generation large model to obtain image adjustment semantic information, and use the image adjustment semantic information to adjust the image content of the initial preview generated image to obtain the target generated image.
[0119] In this embodiment, the text semantic vector, the initial preview image feature vector, and the sketch semantic vector are fused through the picture generation model to form a comprehensive semantic representation, that is, the picture adjustment semantic information, and the picture content of the initial preview generated picture is adjusted according to the picture adjustment semantic information to obtain the target generated picture.
[0120] This embodiment combines text semantic vectors, initial preview image feature vectors, and sketch semantic vectors to implement a flexible and efficient image generation and adjustment system. This approach not only improves the quality of image generation but also enhances user engagement and satisfaction.
[0121] In some embodiments, step S33 of "performing picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain picture adjustment semantic information" may include the following steps S331 to S333:
[0122] S331, mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space;
[0123] In this embodiment, the text semantic vector can be converted into a vector space compatible with the image feature vector using embedding technology in deep learning (such as Word2Vec, BERT or image embedding model). For the initial preview image feature vector, if it is already in a suitable feature space, no additional conversion is required; otherwise, it is converted to the target feature space through a convolutional neural network (CNN) or other image feature extraction method. For the sketch semantic vector, an appropriate image or graphics processing method is also used to ensure that its vector representation is in the same space as the text and initial preview image feature vector. This embodiment can make the above three vectors now located in the same feature vector space, providing a basis for subsequent feature alignment.
[0124] S332. In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information;
[0125] In this embodiment, a deep learning model including a cross-attention mechanism can be designed, which can process feature vectors from text, initial preview images, and sketches. The cross-attention mechanism allows the model to focus on relevant information in other modalities when processing the features of each modality. For example, when processing sketch features, the model can focus on the details described in the text, or when processing text features, focus on specific shapes or layouts in the sketch. Through training, the model learns how to effectively integrate information from different modalities to generate multimodal feature mapping information. This embodiment can generate multimodal feature mapping information that contains key information of text semantic vectors, initial preview image feature vectors, and sketch semantic vectors.
[0126] For example, a cross-attention module can be built: the preview image features are used as the query, and the text and sketch features are used as the key and value, respectively. Through multi-layer cross-attention, the preview image can dynamically "pay attention" to the semantic cues in the text and the structural information in the sketch during the generation process, thereby determining how to make local or overall adjustments in the image.
[0127] S333: Analyze the picture adjustment information based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initial preview generated picture.
[0128] In this embodiment, a parser can be designed that can receive multimodal feature mapping information and output image adjustment semantic information. The parser may include a series of fully connected layers, convolutional layers, or recurrent neural networks (RNNs), etc., for extracting and adjusting relevant image features from the multimodal feature mapping information. Through training, the parser learns how to generate accurate image adjustment semantic information based on the multimodal feature mapping information. The image adjustment semantic information may include instructions such as color adjustment, shape modification, and object position change, which are used to guide the subsequent image adjustment process. This embodiment can obtain image adjustment semantic information for the image generated for the initial preview, and this information will be used to guide further adjustment and optimization of the image.
[0129] In some embodiments, step S33 of "adjusting the image content of the initial preview generated image using the image adjustment semantic information to obtain a target generated image" may include the following steps:
[0130] S334: Acquire multiple layer vectors for the initial preview generated image based on the layered generation architecture, and determine layer adjustment information for each layer vector based on the image adjustment semantic information;
[0131] The layered generative architecture refers to a generative approach that breaks down an image into multiple layers, each containing different parts or features of the image. This approach makes the image generation process more modular and controllable. For example, one layer might contain the background, another might contain foreground objects, and yet another might contain color or texture information.
[0132] Layer vectors are mathematical representations of a layer, containing information or characteristics about all pixels in the layer. These vectors are the basis for image generation and editing.
[0133] The layer adjustment information refers to the adjustment instructions for the layer vectors, which are used to guide the further modification and optimization of the image.
[0134] In this embodiment, the system first generates multiple layer vectors for the initial preview image using a layered generation architecture. These vectors are mathematical representations of the layer, containing information or features about all pixels in the layer. Next, the system determines layer adjustment information for each layer vector based on the image adjustment semantics (which were parsed in the previous step and contain adjustment instructions for the image). This adjustment information may include color changes, shape modifications, object position adjustments, and more.
[0135] For example, the initial preview generated image can be decomposed into multiple layers, including a background layer, a foreground layer, and a local detail layer. Features are extracted from each layer, and multimodal fusion and semantic adjustment are performed on each layer. Each layer is weighted and synthesized according to the layer reference parameters set for each layer to generate the target generated image. S335: The layer vectors are adjusted based on the layer reference parameters and layer adjustment information of each layer vector to obtain an adjusted layer vector.
[0136] The layer reference parameters adjust the contribution of different layers to the final image during the generation process. These parameters reflect the importance and influence of a layer within the image. These parameters guide the adjustment of layer vectors, ensuring overall consistency and harmony in the image.
[0137] In this embodiment, before adjusting layer vectors, the system first determines the reference parameters for each layer. These parameters reflect the importance and influence of the layer in the image. For example, a background layer might have a low reference value because the background is generally not the focus of the user's attention; whereas a foreground object layer might have a high reference value because it is the main content of the image. Next, the system adjusts each layer vector based on the layer reference parameters and the layer adjustment information. The adjustment process may include changing the color, shape, texture, or position of the layer. The system balances the adjustment strength between different layers based on the layer reference parameters to ensure the overall consistency and harmony of the image.
[0138] S336 , encoding the initial preview generated image based on each adjustment layer vector to obtain a target generated image.
[0139] Image content encoding refers to the process of recombining the adjusted layer vectors to generate the adjusted image. This process ensures the integrity and consistency of the image.
[0140] The target generated image refers to the image after adjustment and optimization. It combines the advantages of the initial preview generated image and the user's adjustment requirements and is the final output result.
[0141] In this embodiment, after obtaining the adjusted layer vectors, the system reassembles them to generate the adjusted image. This process is called image content encoding. The system reassembles the vectors of each layer according to their original or new positions to form a complete image. Finally, the system outputs the adjusted image as the target generated image. This image combines the advantages of the initial preview generated image with the user's adjustment requirements, better meeting user expectations and needs.
[0142] This embodiment successfully applies a layered generation architecture to the image generation and editing process. Multiple layer vectors are acquired, and adjustment information for each layer vector is determined based on image adjustment semantics. The layer vectors are then adjusted based on layer reference parameters and recombined to generate the target generated image. This approach not only improves image generation quality but also makes the image editing process more flexible and controllable.
[0143] In some embodiments, the step S1 of “generating an initial preview generation picture in a picture generation preview area based on the picture generation description” may include the following steps:
[0144] S12. Use the picture generation model to convert the picture generation description into a text semantic vector, perform text-image processing based on the text semantic vector to obtain an initial preview generated picture, and display the initial preview generated picture in the picture generation preview area.
[0145] In this embodiment, the user can enter a picture generation description in the picture generation description input box. The picture generation model can first pre-process the input picture generation description, including removing irrelevant characters, word segmentation, part-of-speech tagging, etc., to improve the accuracy of subsequent processing. Then, the text encoder part of the picture generation model is used to convert the pre-processed text description into a text semantic vector. These vectors are mathematical representations of the text description in a high-dimensional space and contain key information in the description. The picture generation model is then used to perform text-graph processing on the semantic information included in the text semantic vector to obtain an initial preview generated picture corresponding to the text semantic vector, and the initial preview generated picture is displayed in the picture generation preview area. The user can zoom in, out, drag or rotate the picture in this picture generation preview area to more fully understand the generated picture effect.
[0146] In some embodiments, the picture generation method provided by the present application can also receive a picture adjustment description input by the user, adjust the preview generated picture in the picture generation preview area based on the picture adjustment description through the picture generation large model, generate an adjusted generated picture, and perform preview update display processing on the picture generation preview area based on the adjusted generated picture.
[0147] In this embodiment, after the initial preview generated picture is generated and the preview generated picture is generated in the picture generation preview area, the preview generated picture in the picture generation preview area can be adjusted according to the picture adjustment description input by the user to generate an adjusted generated picture, and the preview update display processing of the picture generation preview area is performed based on the adjusted generated picture.
[0148] Specifically, the preview generated picture in the picture generation preview area is adjusted based on the picture adjustment description through the picture generation big model to generate the adjusted generated picture. The adjustment text semantic vector corresponding to the picture adjustment description and the preview image feature vector corresponding to the preview generated picture can be obtained through the picture generation big model. The picture adjustment inference processing is performed based on the adjustment text semantic vector and the preview image feature vector through the picture generation big model to obtain the picture adjustment description semantic information. The picture adjustment description semantic information is used to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
[0149] In this embodiment, the large image generation model can utilize a deep learning algorithm, taking the adjustment text semantic vector and the preview image feature vector as input and performing complex reasoning. This process aims to understand the requirements in the adjustment description and determine how to implement these requirements on the preview image. Through this reasoning process, the model generates "image adjustment description semantic information," which contains specific instructions for adjusting the preview image.
[0150] Finally, the model uses the semantic information from the image adjustment description to modify the preview image. This process involves image editing, rendering, and compositing techniques to ensure that the adjusted image meets the user's description requirements while maintaining overall aesthetics and consistency. The adjusted image is called the "adjusted generated image" and reflects the user or system's modification requirements for the preview image. This implementation allows the user or system to modify and optimize the image content through simple text descriptions.
[0151] In one embodiment, a picture generating device is also provided. Figure 7 , Figure 7 This is a schematic diagram of the structure of a picture generation device 400 provided in an embodiment of the present application. The picture generation device 400 is applied to an electronic device and includes a generation module 401, a detection module 402, and an update module 403, as follows:
[0152] A generating module 401 is configured to determine a drawing generation description of the user in a graffiti drawing scene, and generate an initial preview drawing in a drawing generation preview area based on the drawing generation description;
[0153] A detection module 402 is configured to monitor a reference graffiti sketch input by a user in the graffiti canvas area;
[0154] The updating module 403 is configured to adjust the content of the initial preview generated picture using a picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture, and perform preview update display processing on the picture generation preview area based on the target generated picture.
[0155] In some embodiments, the updating module 403 is specifically configured to:
[0156] Obtaining, through a picture generation model, a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture;
[0157] Determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch through a picture generation macro model, and performing sketch semantic encoding on the sketch line information, the sketch shape, the sketch outline information, and the content space layout information to obtain a sketch semantic vector;
[0158] The picture generation model performs picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector and the sketch semantic vector to obtain picture adjustment semantic information, and uses the picture adjustment semantic information to adjust the picture content of the initial preview generated picture to obtain the target generated picture.
[0159] In some embodiments, the updating module 403 is specifically configured to:
[0160] Mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space;
[0161] In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information;
[0162] Picture adjustment information is parsed based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initially previewed picture.
[0163] In some embodiments, the updating module 403 is specifically configured to:
[0164] Acquire multiple layer vectors for the initial preview generated picture based on a layered generation architecture, and determine layer adjustment information for each of the layer vectors based on the picture adjustment semantic information;
[0165] Adjusting the layer vectors based on the layer reference parameters of the layer vectors and the layer adjustment information to obtain adjusted layer vectors;
[0166] The initial preview generated picture is encoded with picture content based on each of the adjustment layer vectors to obtain a target generated picture.
[0167] In some embodiments, the generating module 401 is specifically configured to:
[0168] The picture generation description is converted into a text semantic vector using a picture generation model, and text-image processing is performed based on the text semantic vector to obtain an initial preview generated picture, which is displayed in a picture generation preview area.
[0169] In some embodiments, the updating module 403 is specifically configured to:
[0170] receiving a picture adjustment description input by a user;
[0171] Adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using a picture generation macro model to generate an adjusted generated picture;
[0172] Based on the adjusted generated picture, a preview update display process is performed on the picture generation preview area.
[0173] In some embodiments, the updating module 403 is specifically configured to:
[0174] Obtaining, through a picture generation model, an adjustment text semantic vector corresponding to the picture adjustment description and a preview image feature vector corresponding to the preview generated picture;
[0175] The picture generation model performs picture adjustment inference processing based on the adjustment text semantic vector and the preview image feature vector to obtain picture adjustment description semantic information, and uses the picture adjustment description semantic information to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
[0176] It should be noted that the picture generation device provided in the embodiment of the present application and the picture generation method in the above embodiment have the same concept. Any method provided in the picture generation method embodiment can be implemented through the picture generation device. The specific implementation process is detailed in the picture generation method embodiment and will not be repeated here.
[0177] In addition, in order to better implement the picture generation method in the embodiment of the present application, based on the picture generation method, the present application also provides an electronic device, which can be a smart phone, tablet computer, PDA, laptop computer, or desktop computer. Figure 8 , Figure 8This is a schematic diagram of a first structure of an electronic device provided in an embodiment of the present application. The electronic device 500 includes a processor 501 and a memory 502. The processor 501 is electrically connected to the memory 502.
[0178] The processor 501 is the control center of the electronic device 500. It connects the various parts of the entire electronic device using various interfaces and lines. By running or calling computer programs stored in the memory 502, as well as calling data stored in the memory 502, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device. The processor 501 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0179] Memory 502 can be used to store computer programs and data. The computer programs stored in memory 502 contain instructions that can be executed by the processor. Computer programs can be composed of various functional modules. Processor 401 executes various functional applications and data processing by calling the computer programs stored in memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the electronic device 500 (such as audio data, video data, etc.). In addition, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0180] In this embodiment, the processor 501 in the electronic device 500 loads instructions corresponding to one or more computer program processes into the memory 502 according to the following steps, and the processor 501 runs the computer program stored in the memory 502 to implement various functions:
[0181] In a graffiti drawing scene, determining a drawing generation description of the user, and generating an initial preview drawing in a drawing generation preview area based on the drawing generation description;
[0182] monitoring a reference doodle sketch input by a user in a doodle canvas area;
[0183] Based on the reference graffiti sketch and the picture generation description, a picture generation model is used to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and a preview update display process is performed on the picture generation preview area based on the target generated picture.
[0184] In some embodiments, see Figure 9 , Figure 9 This is a schematic diagram of a second structure of an electronic device provided in an embodiment of the present application. The electronic device 500 further includes a radio frequency circuit 503, a display screen 504, a control circuit 505, an input unit 506, an audio circuit 507, a sensor 508, and a power supply 509. The processor 501 is electrically connected to the radio frequency circuit 503, the display screen 504, the control circuit 505, the input unit 506, the audio circuit 507, the sensor 508, and the power supply 509, respectively.
[0185] The radio frequency circuit 503 is used to transmit and receive radio frequency signals to communicate with network devices or other electronic devices through wireless communication.
[0186] The display screen 504 may be used to display information input by a user or information provided to a user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces may be composed of images, texts, icons, videos, and any combination thereof.
[0187] The control circuit 505 is electrically connected to the display screen 504 and is used to control the display screen 504 to display information.
[0188] The input unit 506 may be configured to receive input numbers, characters, or user characteristics (e.g., fingerprints), and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. The input unit 506 may include a fingerprint recognition module.
[0189] The audio circuit 507 can provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 507 includes a microphone. The microphone is electrically connected to the processor 501. The microphone is used to receive voice information input by the user.
[0190] The sensor 508 is used to collect external environment information. The sensor 508 may include one or more sensors such as an ambient brightness sensor, an acceleration sensor, and a gyroscope.
[0191] The power supply 509 is used to supply power to various components of the electronic device 500. In some embodiments, the power supply 509 can be logically connected to the processor 501 through a power management system, so that the power management system can manage charging, discharging, and power consumption.
[0192] Although not shown in the figure, the electronic device 500 may also include a camera, a Bluetooth module, etc., which will not be described in detail here.
[0193] In this embodiment, the processor 501 in the electronic device 500 loads instructions corresponding to one or more computer program processes into the memory 502 according to the following steps, and the processor 501 runs the computer program stored in the memory 502 to implement various functions:
[0194] In a graffiti drawing scene, determining a drawing generation description of the user, and generating an initial preview drawing in a drawing generation preview area based on the drawing generation description;
[0195] monitoring a reference doodle sketch input by a user in a doodle canvas area;
[0196] Based on the reference graffiti sketch and the picture generation description, a picture generation model is used to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and a preview update display process is performed on the picture generation preview area based on the target generated picture.
[0197] In some embodiments, when the processor 501 performs the step of adjusting the content of the initial preview generated picture using the picture generation model based on the reference graffiti sketch and the picture generation description to obtain the target generated picture, the processor 501 may perform the following steps:
[0198] Obtaining, through a picture generation model, a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture;
[0199] Determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch through a picture generation macro model, and performing sketch semantic encoding on the sketch line information, the sketch shape, the sketch outline information, and the content space layout information to obtain a sketch semantic vector;
[0200] The picture generation model performs picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector and the sketch semantic vector to obtain picture adjustment semantic information, and uses the picture adjustment semantic information to adjust the picture content of the initial preview generated picture to obtain the target generated picture.
[0201] In some implementations, when the processor 501 performs the image adjustment inference processing based on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain the image adjustment semantic information, the processor 501 may execute:
[0202] Mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space;
[0203] In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information;
[0204] Picture adjustment information is parsed based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initially previewed picture.
[0205] In some implementations, when the processor 501 performs the step of adjusting the image content of the initial preview generated image using the image adjustment semantic information to obtain the target generated image, the processor 501 may perform:
[0206] Acquire multiple layer vectors for the initial preview generated picture based on a layered generation architecture, and determine layer adjustment information for each of the layer vectors based on the picture adjustment semantic information;
[0207] Adjusting the layer vectors based on the layer reference parameters of the layer vectors and the layer adjustment information to obtain adjusted layer vectors;
[0208] The initial preview generated picture is encoded with picture content based on each of the adjustment layer vectors to obtain a target generated picture.
[0209] In some embodiments, when the processor 501 generates an initial preview image in the image generation preview area based on the image generation description, it may execute:
[0210] The picture generation description is converted into a text semantic vector using a picture generation model, and text-image processing is performed based on the text semantic vector to obtain an initial preview generated picture, which is displayed in a picture generation preview area.
[0211] In some implementations, the processor 501 may further execute:
[0212] receiving a picture adjustment description input by a user;
[0213] Adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using a picture generation macro model to generate an adjusted generated picture;
[0214] In some embodiments, the processor 501 performs the following operations when performing the adjustment process on the preview image generated in the preview area of the picture generation based on the picture adjustment description using the picture generation macro model:
[0215] Obtaining, through a picture generation model, an adjustment text semantic vector corresponding to the picture adjustment description and a preview image feature vector corresponding to the preview generated picture;
[0216] The picture generation model performs picture adjustment inference processing based on the adjustment text semantic vector and the preview image feature vector to obtain picture adjustment description semantic information, and uses the picture adjustment description semantic information to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
[0217] An embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is run on a computer, the computer executes the picture generation method described in any of the above embodiments.
[0218] It should be noted that, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the storage medium may include but is not limited to: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0219] Furthermore, the terms "first," "second," and "third," etc., in this application are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but rather some embodiments may include steps or modules not listed, or other steps or modules that are inherent to such process, method, product, or apparatus.
[0220] The above describes in detail the image generation method, device, storage medium, and electronic device provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is intended only to facilitate understanding of the method and core concept of the present application. Furthermore, those skilled in the art will appreciate that variations in the specific implementation methods and scope of application may occur based on the concepts of the present application. In summary, the contents of this specification should not be construed as limiting the present application.
Claims
1. A method for generating a picture, characterized in that: include: In a graffiti drawing scene, determining a drawing generation description of the user, and generating an initial preview drawing in a drawing generation preview area based on the drawing generation description; monitoring a reference doodle sketch input by a user in a doodle canvas area; Based on the reference graffiti sketch and the picture generation description, a picture generation model is used to adjust the picture content of the initial preview generated picture to obtain a target generated picture, and a preview update display process is performed on the picture generation preview area based on the target generated picture.
2. The method according to claim 1, characterized in that The step of adjusting the content of the initial preview generated picture using a large picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture includes: Obtaining, through a picture generation model, a text semantic vector corresponding to the picture generation description and an initial preview image feature vector corresponding to the initial preview generation picture; Determining sketch line information, sketch shape, sketch outline information, and content space layout information of the reference graffiti sketch through a picture generation macro model, and performing sketch semantic encoding on the sketch line information, the sketch shape, the sketch outline information, and the content space layout information to obtain a sketch semantic vector; The picture generation model performs picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector and the sketch semantic vector to obtain picture adjustment semantic information, and uses the picture adjustment semantic information to adjust the picture content of the initial preview generated picture to obtain the target generated picture.
3. The method according to claim 2, characterized in that The performing picture adjustment inference processing based on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain picture adjustment semantic information includes: Mapping the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to the same feature vector space; In the same feature vector space, a cross-attention mechanism is used to perform feature mapping alignment processing on the text semantic vector, the initial preview image feature vector, and the sketch semantic vector to obtain multimodal feature mapping information; Picture adjustment information is parsed based on the multimodal feature mapping information to obtain picture adjustment semantic information for the initially previewed picture.
4. The method according to claim 2, characterized in that The step of adjusting the content of the initial preview generated picture by using the picture adjustment semantic information to obtain a target generated picture includes: Acquire multiple layer vectors for the initial preview generated picture based on a layered generation architecture, and determine layer adjustment information for each of the layer vectors based on the picture adjustment semantic information; Adjusting the layer vectors based on the layer reference parameters of the layer vectors and the layer adjustment information to obtain adjusted layer vectors; The initial preview generated picture is encoded with picture content based on each of the adjustment layer vectors to obtain a target generated picture.
5. The method according to claim 1, characterized in that The step of generating an initial preview picture in a picture generation preview area based on the picture generation description includes: The picture generation description is converted into a text semantic vector using a picture generation model, and text-image processing is performed based on the text semantic vector to obtain an initial preview generated picture, which is displayed in a picture generation preview area.
6. The method according to claim 1, characterized in that The method further comprises: receiving a picture adjustment description input by a user; Adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using a picture generation macro model to generate an adjusted generated picture; A preview update display process is performed on the picture generation preview area based on the adjusted generated picture.
7. The method according to claim 6, characterized in that The step of adjusting the preview generated picture in the picture generation preview area based on the picture adjustment description by using the picture generation macro model to generate the adjusted generated picture includes: Obtaining, through a picture generation model, an adjustment text semantic vector corresponding to the picture adjustment description and a preview image feature vector corresponding to the preview generated picture; The picture generation model performs picture adjustment inference processing based on the adjustment text semantic vector and the preview image feature vector to obtain picture adjustment description semantic information, and uses the picture adjustment description semantic information to adjust the picture content of the preview generated picture to obtain the adjusted generated picture.
8. A picture generating device, characterized in that: include: A generating module, configured to determine a drawing generation description of the user in a graffiti drawing scene, and generate an initial preview drawing in a drawing generation preview area based on the drawing generation description; a detection module, configured to detect a reference graffiti sketch input by a user in a graffiti canvas area; An updating module is configured to adjust the content of the initial preview generated picture using a picture generation model based on the reference graffiti sketch and the picture generation description to obtain a target generated picture, and perform preview update display processing on the picture generation preview area based on the target generated picture.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed on a computer, the computer is caused to execute the picture generating method according to any one of claims 1 to 7.
10. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, wherein: The processor is configured to execute the picture generating method according to any one of claims 1 to 7 by calling the computer program.