Drawing adjustment processing method and device, storage medium and electronic equipment

By receiving user instructions in the literary and biographical scene and using the painting processing model to generate picture adjustment description text, the problem of time-consuming and labor-consuming traditional painting adjustment is solved, and efficient picture adjustment is achieved.

CN120495472APending Publication Date: 2025-08-15SHENZHEN QIHOO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510570050.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional painting methods are difficult to make large adjustments, which consumes time and effort and may destroy the original paintings, limiting the efficiency of painting adjustment.

Method used

By receiving user picture adjustment instructions in the literary picture scene, using the painting processing model to analyze instructions and adjust description text generation, intelligent adjustment of reference generated pictures is achieved.

Benefits of technology

Improve the efficiency of painting adjustment, and automatically adjust the drawing according to user needs to generate target adjustments that meet expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495472A_ABST
    Figure CN120495472A_ABST
Patent Text Reader

Abstract

The invention discloses a drawing adjustment processing method and device, a storage medium and electronic equipment, and the method comprises the steps: receiving a drawing adjustment instruction of a user for a reference generation drawing in a text drawing scene; and performing instruction content analysis on the picture adjustment instruction through the drawing processing large model to obtain a picture adjustment description text for the reference generation picture, performing picture adjustment processing on the reference generation picture based on the picture adjustment description text, generating a target adjustment picture for the reference generation picture, and displaying the target adjustment picture to a user. Therefore, the reference generated picture can be automatically adjusted directly through the picture adjusting instruction input by the user, a painter can be assisted in drawing adjustment, and the drawing adjustment efficiency of the painter is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a painting adjustment processing method, device, storage medium and electronic device. Background Art

[0002] Painting, as a creative and expressive art form, has long been beloved by art enthusiasts. However, traditional painting methods have limitations in terms of flexibility. For example, once a painting is completed, it is difficult to make significant adjustments. The artist often needs to repaint or overwrite the original parts, which is not only time-consuming and labor-intensive, but can also damage the overall feel and details of the original painting. Traditional methods are particularly inconvenient for paintings that require multiple adjustments or refinements, significantly limiting the artist's efficiency in making adjustments. Summary of the Invention

[0003] The embodiments of the present application provide a painting adjustment processing method, device, storage medium, and electronic device, which can assist painters in painting adjustments and improve the painters' painting adjustment efficiency.

[0004] In a first aspect, an embodiment of the present application provides a method for adjusting a painting, including:

[0005] In the Wensheng picture scenario, receiving a user's picture adjustment instruction for a reference generated picture;

[0006] parsing the picture adjustment instruction using a large painting processing model to obtain a picture adjustment description text for the reference generated picture, performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture;

[0007] The target adjustment picture is presented to the user.

[0008] In some embodiments, performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture includes:

[0009] Acquire a first picture description text for the reference generated picture through the painting processing large model, wherein the first picture description text is a picture description after content description processing is performed on the reference generated picture in advance;

[0010] Performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode;

[0011] Based on the graphic adjustment instruction set, the target adjusted picture is obtained by performing picture adjustment through the painting processing model.

[0012] In some embodiments, performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode includes:

[0013] determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for the reference generated picture;

[0014] Using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector;

[0015] The fused adjustment text vector is subjected to a graphic-text adjustment decoding process to obtain a graphic-text adjustment mode, and a description instruction conversion process is performed based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

[0016] In some embodiments, performing a graphic-text adjustment decoding process on the fused adjustment text vector to obtain a graphic-text adjustment mode, and performing a description instruction conversion process based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set includes:

[0017] performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map;

[0018] A graphic and text adjustment mode is determined based on the distribution type of the adjustment attention map, and a graphic and text adjustment instruction set is obtained by performing instruction generation processing using a fused adjustment text vector based on the graphic and text adjustment mode.

[0019] In some embodiments, determining a graphic-text adjustment mode based on the distribution type of the attention map adjustment, and performing instruction generation processing based on the graphic-text adjustment mode using a fused adjustment text vector to obtain a graphic-text adjustment instruction set include:

[0020] If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment area based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment area based on the fused adjustment text vector to obtain a first picture and text adjustment instruction set;

[0021] If the attention map is of a global distribution type, the global editing mode is determined by the painting processing model, and based on the global editing mode, the fused adjustment text vector is used to generate a picture reconstruction instruction to obtain a second picture and text adjustment instruction set.

[0022] In some embodiments, presenting the target adjustment picture to the user includes:

[0023] Obtaining scene context information of the cultural image scene;

[0024] Determining a user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention;

[0025] The target adjustment picture and the picture expansion adjustment solution are presented to the user.

[0026] In some embodiments, determining the user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating multiple picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention, includes:

[0027] Determining context scene association features and a predicted painting style using the painting processing large model based on the scene context information, and performing potential painting intention inference based on the context scene association features and the predicted painting style to obtain a user adjustment intention;

[0028] An object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item are determined based on the user adjustment intention, and a plurality of picture extension adjustment schemes are obtained by combining the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item.

[0029] In a second aspect, an embodiment of the present application further provides a painting adjustment processing device, comprising:

[0030] A receiving module, configured to receive, in a text-generated image scene, an image adjustment instruction from a user for a reference generated image;

[0031] a generation module configured to parse the picture adjustment instruction using a painting processing macromodel to obtain a picture adjustment description text for the reference generated picture, and perform picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture;

[0032] A display module is used to display the target adjustment picture to the user.

[0033] In some embodiments, the generating module is specifically configured to:

[0034] Acquire a first picture description text for the reference generated picture through the painting processing large model, wherein the first picture description text is a picture description after content description processing is performed on the reference generated picture in advance;

[0035] Performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode;

[0036] Based on the graphic adjustment instruction set, the target adjusted picture is obtained by performing picture adjustment through the painting processing model.

[0037] In some embodiments, the generating module is specifically configured to:

[0038] determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for the reference generated picture;

[0039] Using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector;

[0040] The fused adjustment text vector is subjected to a graphic-text adjustment decoding process to obtain a graphic-text adjustment mode, and a description instruction conversion process is performed based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

[0041] In some embodiments, the generating module is specifically configured to:

[0042] performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map;

[0043] A graphic and text adjustment mode is determined based on the distribution type of the adjustment attention map, and a graphic and text adjustment instruction set is obtained by performing instruction generation processing using a fused adjustment text vector based on the graphic and text adjustment mode.

[0044] In some embodiments, the generating module is specifically configured to:

[0045] If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment area based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment area based on the fused adjustment text vector to obtain a first picture and text adjustment instruction set;

[0046] If the attention map is of a global distribution type, the global editing mode is determined by the painting processing model, and based on the global editing mode, the fused adjustment text vector is used to generate a picture reconstruction instruction to obtain a second picture and text adjustment instruction set.

[0047] In some embodiments, the display module is specifically used to:

[0048] Obtaining scene context information of the cultural image scene;

[0049] Determining a user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention;

[0050] The target adjustment picture and the picture expansion adjustment solution are presented to the user.

[0051] In some embodiments, the display module is specifically used to:

[0052] Determining context scene association features and a predicted painting style using the painting processing large model based on the scene context information, and performing potential painting intention inference based on the context scene association features and the predicted painting style to obtain a user adjustment intention;

[0053] An object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item are determined based on the user adjustment intention, and a plurality of picture extension adjustment schemes are obtained by combining the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item.

[0054] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is run on a computer, the computer executes the painting adjustment processing method provided in any embodiment of the present application.

[0055] In a fourth aspect, an embodiment of the present application further provides an electronic device, comprising a processor and a memory, wherein the memory has a computer program, and the processor is configured to execute a painting adjustment processing method as provided in any embodiment of the present application by calling the computer program.

[0056] The technical solution provided by the embodiments of the present application receives a user's picture adjustment instruction for a reference generated picture in a Wensheng picture scenario, parses the picture adjustment instruction content using a large painting processing model to obtain a picture adjustment description text for the reference generated picture, performs picture adjustment processing on the reference generated picture based on the picture adjustment description text, generates a target adjustment picture for the reference generated picture, and displays the target adjustment picture to the user. In this way, the present application can automatically adjust the reference generated picture directly based on the picture adjustment instruction input by the user, assisting the painter in making painting adjustments and improving the painter's painting adjustment efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0058] Figure 1 This is a schematic diagram of the first flow chart of the painting adjustment processing method provided in an embodiment of the present application.

[0059] Figure 2 This is a schematic diagram of a first application scenario example of the painting adjustment processing method provided in an embodiment of the present application.

[0060] Figure 3 This is a schematic diagram of a second application scenario example of the painting adjustment processing method provided in an embodiment of the present application.

[0061] Figure 4 This is a schematic diagram of a third application scenario example of the painting adjustment processing method provided in an embodiment of the present application.

[0062] Figure 5 This is a structural diagram of the painting adjustment processing device provided in an embodiment of the present application.

[0063] Figure 6 This is a schematic diagram of the first structure of the electronic device provided in an embodiment of the present application.

[0064] Figure 7 A second structural diagram of the electronic device provided in an embodiment of the present application. Specific embodiments

[0065] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.

[0066] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0067] Painting, as a creative and expressive art form, has long been beloved by art enthusiasts. However, traditional painting methods have limitations in terms of flexibility. For example, once a painting is completed, it is difficult to make significant adjustments. The artist often needs to repaint or overwrite the original parts, which is not only time-consuming and labor-intensive, but can also damage the overall feel and details of the original painting. Traditional methods are particularly inconvenient for paintings that require multiple adjustments or refinements, significantly limiting the artist's efficiency in making adjustments.

[0068] To improve the efficiency of adjusting a painting, embodiments of the present application provide a painting adjustment processing method. The method may be performed by a painting adjustment processing device provided in embodiments of the present application, or an electronic device incorporating the painting adjustment processing device. The painting adjustment processing device may be implemented in hardware or software. The electronic device may be a smartphone, tablet computer, PDA, laptop computer, or desktop computer.

[0069] See also Figure 1 , Figure 1 This is a schematic diagram of the first process of the painting adjustment processing method provided in the embodiment of the present application. The specific process of the painting adjustment processing method provided in the embodiment of the present application can be as follows:

[0070] S1. In a scene of a generated image, receiving a user's instruction to adjust an image generated by a reference.

[0071] The "text-to-image" scenario refers to a technology application scenario in which a textual description (such as text, sentences, or paragraphs) is converted into an image or picture through a specific algorithm or technology. This technology is commonly used in fields such as artistic creation, design automation, and advertising creative generation. For example, a user enters a text describing a natural landscape, and the system generates a corresponding landscape painting based on the text.

[0072] A reference image is a preliminary image generated in a text-based image scenario based on a user's text description or other input (such as sketches, keywords, etc.). This image may not fully meet the user's final expectations and may require further adjustment. For example, a user enters the text description "A couple strolling by the beach on a sunny afternoon." The system generates a picture of a couple strolling by the beach based on this description as the reference image, but the user may feel that the couple's expressions or postures in the reference image need to be adjusted.

[0073] Image adjustment instructions are specific modification requests made by users to the initially generated image in the context of the text image scene. For example, object elements, foreground elements, and background elements in the reference generated image can be adjusted.

[0074] The object element is the most central and crucial element in an image, often representing the primary message or theme. In portraits, it's often the person; in landscapes, it might be a beautiful natural landscape or a specific landmark. Objects are often the most eye-catching part of an image and the primary message it conveys.

[0075] Foreground elements are those positioned between the main subject and the viewer. They can add depth, layering, and visual guidance to an image. Foreground elements can be anything that contrasts or echoes the main subject, such as tree branches, flowers, or the corner of a building. By skillfully utilizing foreground elements, you can make an image more vivid and interesting, and guide the viewer's gaze.

[0076] Background elements are used to highlight the subject and foreground elements in an image, providing the context or setting for the subject. Backgrounds can range from simple, solid colors to complex natural environments or man-made structures. Background design plays an important role in highlighting the subject, creating atmosphere, and guiding the eye. A suitable background can enhance the overall effect of an image and make the subject stand out more clearly.

[0077] Image adjustment instructions may involve adjustments to the color, shape, layout, and details of the reference image, aiming to make the generated image more consistent with the user's expectations. For example, after viewing the initial generated image, a user may request an adjustment such as "I want the couple's expressions to be more affectionate, and the girl's skirt to be a lighter blue."

[0078] In some embodiments, the image adjustment instruction can be a textual adjustment instruction input by the user for the reference generated image, or can be a template image input by the user, and the user wishes the reference generated image to be adjusted to mimic the image content elements in the template image. For example, the image content elements can include object elements, foreground elements, and background elements in the image.

[0079] S2. parse the picture adjustment instruction through the painting processing large model to obtain a picture adjustment description text for the reference generated picture, perform picture adjustment processing on the reference generated picture based on the picture adjustment description text, and generate a target adjusted picture for the reference generated picture.

[0080] The painting processing model is configured to parse the content of the picture adjustment instructions to obtain a picture adjustment description text for the reference generated picture, and then perform picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture. The painting processing model is specifically designed to process and analyze painting-related data. It can understand painting styles, techniques, color application, etc., and can also intelligently adjust or generate images based on input instructions or requirements. The painting processing model is based on deep learning technology and is trained with a large number of paintings and related knowledge to gain a deep understanding and processing ability of the art of painting.

[0081] The image adjustment description is a textual description generated by the painting processing model based on the input image adjustment instructions. It details the adjustments required to the reference generated image, such as changing color, adjusting composition, adding or removing elements, and so on. This description is a critical bridge between user intent and model operation, ensuring that the model accurately understands and implements user adjustment requests.

[0082] The target adjustment image is the final image obtained by adjusting the reference generated image according to the image adjustment description text. It reflects the painting effect or style change that the user hopes to achieve by painting the large model.

[0083] In this embodiment, the user provides an image adjustment instruction, which the painting processing model interprets and generates a detailed description of the image adjustment. The model then adjusts the reference image based on this description, ultimately generating a target image that meets the user's requirements.

[0084] In one example, suppose a user wants to change the sky color of a landscape painting (i.e., a reference-generated image) from blue to orange, while also adding some layering to the clouds. They input the landscape painting and the image adjustment instructions into the painting processing model. After receiving the image adjustment instructions, the painting processing model parses them to understand the user's desire to change the sky color and add layering to the clouds. The model-generated description of the image adjustment might read: "Change the sky color from blue to orange, while adding layering and detail to the cloud area." The target image might be an orange sky with richly layered clouds.

[0085] In some embodiments, the painting processing model is trained according to the following steps:

[0086] (1) Obtain a basic large language model and a basic text-image model, and create an initial painting processing large model for text-image scenes.

[0087] Among them, basic large language models include but are not limited to large models in the field of natural language processing (NLP), such as GPT-3, GPT-4, ChatGPT, BERT, and RoBERTa models.

[0088] Among them, the basic text-to-image model is the basic model for text-to-image generation. This type of model can generate corresponding images based on the input text description, such as the Stable Diffusion model, GAN (generative adversarial model), DALL-E model, and CogView model.

[0089] (2) Obtain a sample image adjustment instruction, and label the sample image adjustment instruction with an image adjustment description text label and a target adjustment image label.

[0090] (3) Using sample picture adjustment instructions to perform at least one round of model training on the initial painting processing model.

[0091] (4) During the forward propagation training of the model, the initial painting processing model determines the predicted picture adjustment description text based on the sample picture adjustment instruction, and performs picture generation processing based on the predicted picture adjustment description text to obtain the predicted target adjustment picture.

[0092] (5) During the model backpropagation training process, the text prediction loss is determined based on the predicted picture adjustment description text and the picture adjustment description text label, the picture generation loss is determined based on the predicted target adjustment picture and the target adjustment picture label, and the model comprehensive loss is determined based on the text prediction loss and the picture generation loss. The model comprehensive loss is used to adjust the model parameters of the initial painting processing large model until the initial painting processing large model completes the model training and obtains the painting processing large model.

[0093] Specifically:

[0094] The first model loss calculation formula is used to calculate the text prediction loss based on the predicted picture adjustment description text and the picture adjustment description text label. The second model loss calculation formula is used to calculate the image generation loss based on the predicted target adjustment picture and the target adjustment picture label. The third model loss calculation formula is used to calculate the model comprehensive loss based on the text prediction loss and the picture generation loss.

[0095] Among them, the model loss calculation formulas such as the first model loss calculation formula, the second model loss calculation formula and the third model loss calculation formula can be fitted by one or more of the contrast loss calculation formula, cross entropy loss calculation formula, hinge loss calculation formula and the like in the relevant technology.

[0096] Optionally, the model training termination conditions for the initial painting processing large model may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. Specific model training termination conditions can be determined based on actual conditions and are not specifically limited here.

[0097] During specific implementation, the present application is not limited by the execution order of the various steps described. If no conflict occurs, some steps can be performed in other orders or simultaneously.

[0098] S3. Show the target adjustment picture to the user.

[0099] In this embodiment, after the target adjustment picture is generated based on the reference generated picture input by the user, the target adjustment picture is displayed to the user for reference.

[0100] In an example, see Figure 2 , Figure 2The first picture from top to bottom is a scene of a cultural image. The user inputs a reference generated picture generated by "drawing a baby elephant" into the cultural image system. Based on the reference generated picture, the user inputs the reference generated picture into the painting processing model and inputs a picture adjustment instruction for the reference generated picture, "add a red hat to the baby elephant". The painting processing model parses the picture adjustment instruction to obtain a picture adjustment description text for the reference generated picture. The picture adjustment description text can be "Please cleverly add a fashionable and eye-catching red hat to the baby elephant's head in this picture depicting the baby elephant. The hat should be designed to fit the outline of the baby elephant's head and look playful and cute. The red color should be bright but not dazzling, forming a sharp and harmonious contrast with the baby elephant's skin color." The painting processing model performs picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a Figure 2 The second picture viewed from top to bottom in the image is used as the target adjustment picture for the reference generated picture. Next, the user wants to continue adjusting the picture and continues to input the picture adjustment instruction "Change the picture background to blue sky and city". The painting processing model parses the instruction content of the new picture adjustment instruction to obtain a new picture adjustment description text, which may be "Please replace the background elements in the current picture, and replace the original background with a vast blue sky and a modern urban building complex in the distance". The painting processing model adjusts the picture based on the new picture adjustment description text. Figure 2 The second picture in the image is adjusted to generate Figure 2 The third picture from top to bottom is used as the target adjustment picture for the second picture. It is understandable that when the user wants to continue adjusting the picture, the Figure 2 The second image in is used as the reference image for processing.

[0101] During specific implementation, the present application is not limited by the execution order of the various steps described. If no conflict occurs, some steps can be performed in other orders or simultaneously.

[0102] As can be seen from the above, the painting adjustment processing method provided in the embodiments of the present application receives a user's picture adjustment instruction for a reference generated picture in a Wensheng picture scenario, parses the picture adjustment instruction content using the painting processing macro model to obtain a picture adjustment description text for the reference generated picture, performs picture adjustment processing on the reference generated picture based on the picture adjustment description text, generates a target adjustment picture for the reference generated picture, and displays the target adjustment picture to the user. In this way, the present application can automatically adjust the reference generated picture directly based on the picture adjustment instruction input by the user, assisting the painter in making painting adjustments and improving the painter's painting adjustment efficiency.

[0103] In some embodiments, step S2 of "performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture" may include the following steps S21 to S23:

[0104] S21. Obtaining a first picture description text for the reference generated picture through the painting processing macromodel, where the first picture description text is a picture description obtained by pre-processing the content description of the reference generated picture;

[0105] The first picture description text refers to a detailed, structured text description generated after in-depth analysis and understanding of the reference generated picture through the painting processing large model. This description text not only contains the basic content of the picture (such as theme, object, scene, etc.), but also covers advanced features such as the style, color matching, composition details, etc. It is understandable that the purpose of the first picture description text provided in this embodiment is to provide an accurate and comprehensive guide or reference so that the picture can be regenerated, modified, or optimized based on this description.

[0106] For example, consider a reference generated image, a landscape depicting an ancient castle. After inputting this reference generated image into the painting processing model, a description of the image is pre-acquired. An example of a possible "first description of the image" for this landscape might be: "In the center of the picture is an ancient and majestic castle, standing amidst a lush forest. The castle walls are built of thick stone blocks, their surface covered with traces of time. The sky in the picture is a light blue, with a few white clouds drifting leisurely. The mountains in the distance are stacked one upon another, forming a harmonious picture with the castle and forest nearby."

[0107] S22, performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing large model to obtain an image-text adjustment pattern, and generating an image-text adjustment instruction set based on the image-text adjustment pattern;

[0108] After image and text alignment, the painting processing model generates an image and text adjustment pattern. This pattern represents image and text adjustment requirements in a more understandable and operational manner. This pattern is a structured representation that includes multiple adjustment elements and parameters, each corresponding to a specific aspect or feature of the image. This pattern describes the correspondence between images and text, the adjustment strategy, and other aspects, providing a basis for the subsequent generation of image and text adjustment instruction sets.

[0109] The image adjustment instruction set contains a series of specific operational instructions that guide the image editing or generation process. These instructions may include color adjustment instructions, shape modification instructions, layout adjustment instructions, and more. Each instruction corresponds to an adjustment element and adjustment parameter in the image adjustment mode. The image adjustment instruction set can be directly used by image processing software or algorithms to achieve the goal of image adjustment.

[0110] In this embodiment, the painting processing large model performs image and text adjustment alignment processing based on the picture adjustment description text, the first picture description text and the reference generated picture, that is, it combines the picture adjustment description text, the first picture description text and the reference generated picture to analyze the adjustments that need to be made to the target adjustment picture to be generated based on the reference generated picture, and determines the image and text adjustment mode according to the adjustments that need to be made, and generates the image and text adjustment instruction set based on the image and text adjustment mode, so as to adjust and modify the picture through the image and text adjustment instruction set.

[0111] S23. Based on the image and text adjustment instruction set, the image is adjusted by using the painting processing model to obtain a target adjusted image.

[0112] In this embodiment, after the image and text adjustment instruction set is determined, the image adjustment can be performed by painting the large model to obtain the target adjustment image.

[0113] In some embodiments, step S22 of "performing an image-text adjustment alignment process based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode" may include the following steps S221-S223:

[0114] S221, determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for generating a reference picture;

[0115] The image adjustment description vector is a numerical representation of the image adjustment description text converted through a natural language processing model. This vector captures key information about the image adjustment instructions or features described in the text, such as color changes, shape adjustments, and layout changes. This conversion enables computers to understand and execute the adjustment instructions in the text and is commonly used in tasks such as image editing and style transfer.

[0116] The first image description vector is the result of converting the first image description text into a numerical representation. This vector reflects the description of the first image in the text, which may include features such as the image's theme, style, color, and object. This conversion enables computers to recognize and understand the image information described in the text. It is commonly used in tasks such as image retrieval and image generation, serving as a reference for generating or retrieving images.

[0117] The reference image encoding vector is a numerical representation of a specific image (in this context, the "reference generated image") converted through image encoding methods (such as deep learning models). This vector captures key visual features of the image, such as color distribution, texture patterns, and shape structure. This conversion enables the image to be processed and analyzed by computers in numerical form. It is often used as input or reference information in tasks such as image recognition, image classification, and image generation.

[0118] In this embodiment, first, the image adjustment description text is converted into an image adjustment description vector so that a computer can understand and execute the image adjustment requirements described in the text. Second, the first image description text is converted into a first image description vector. The first image description vector serves as a numerical description of the reference generated image, guiding or constraining the generation of a new image and ensuring that the subsequently generated target adjusted image is consistent or similar in certain features to the reference generated image. Finally, a reference image encoding vector is determined for the reference generated image. This vector serves as a reference point or benchmark for the subsequently generated target adjusted image, ensuring that the subsequently generated target adjusted image visually meets specific requirements or styles.

[0119] S222, using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector;

[0120] The image-text attention mechanism is a deep learning technique that integrates text and image information. It allows the model to dynamically focus on image regions related to text descriptions when processing images, or to focus on text segments related to the image content when processing text. This mechanism achieves cross-modal information fusion and interaction by calculating the correlation, or attention weight, between text and images.

[0121] In this embodiment, an image-text attention mechanism is used to fuse and adjust the image adjustment description vector, the first image description vector, and the reference image encoding vector to generate a fused adjusted text vector. This fused adjusted text vector combines the numerical representation of image and text information and is used to guide subsequent image processing or generation tasks. This vector contains information about the adjusted image features, style, or content, as well as feature information related to the reference generated image. This process involves modifying image content, transferring style, and enhancing or suppressing features to ensure that the generated new image not only meets the adjustment requirements of the image adjustment description text, but also maintains consistency or similarity with the reference generated image in certain features.

[0122] It is understandable that the fused adjustment text vector fuses the picture adjustment description, the first picture description and the reference image coding information and is expressed in the form of a numerical vector. This vector contains all the key information required for picture and text adjustment.

[0123] S223. Performing a graphic-text adjustment decoding process on the fused adjustment text vector to obtain a graphic-text adjustment mode, and performing a description instruction conversion process based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

[0124] Image and text adjustment decoding is the process of converting the fused adjusted text vector back into an image and text adjustment model that is easier to understand and operate. This process involves steps such as vector decoding, feature extraction, and information reorganization, aiming to convert the numerical vector information into a more specific and interpretable image and text adjustment model. The image and text adjustment model may include information such as image content modification, style transfer, and layout adjustment.

[0125] The description instruction conversion process is the process of converting the image and text adjustment mode into an image and text adjustment instruction set. This process converts the adjustment elements and adjustment parameters in the image and text adjustment mode into specific operational instructions, which can be directly used to guide the image editing or generation process. The description instruction conversion process may involve steps such as instruction formatting, parameter adjustment, and instruction sorting to ensure that the generated instruction set can accurately and efficiently achieve the goal of image and text adjustment. It can be understood that the image and text adjustment instruction set is the result of the description instruction conversion process.

[0126] In this embodiment, the fused adjustment text vector is subjected to graphic and text adjustment decoding processing to obtain a graphic and text adjustment pattern, and a description instruction conversion processing is performed based on the graphic and text adjustment pattern to obtain a graphic and text adjustment instruction set, which aims to convert the digitized vector information into a specific and operational graphic and text adjustment instruction set to guide the image editing or generation process.

[0127] In some embodiments, step S223 of "performing a graphic-text adjustment decoding process on the fused adjustment text vector to obtain a graphic-text adjustment mode, and performing a description instruction conversion process based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set" may include the following steps S2231-S2232:

[0128] S2231, performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map;

[0129] In this embodiment, attention mapping processing is performed on the first picture description vector based on the fused adjustment text vector to obtain an adjustment attention map. The adjustment attention map can intuitively indicate the specific position of the modified content described in the adjustment text in the original image.

[0130] S2232. Determine the image-text adjustment mode based on the distribution type of the attention map adjustment, and use the fused adjustment text vector to perform instruction generation processing based on the image-text adjustment mode to obtain an image-text adjustment instruction set.

[0131] In this embodiment, if the attention mapping in the distribution type of the attention map is concentrated in a local area, it means that the adjustment demand belongs to the local editing mode; if the attention mapping in the distribution type of the attention map is adjusted to show a global distribution, the adjustment demand may require overall reconstruction, which means that the adjustment demand is a global editing mode.

[0132] In some embodiments, step S2232 "determining a graphic-text adjustment mode based on the distribution type of the attention map, and performing instruction generation processing using the fused adjustment text vector based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set" may include the following steps S22321-S22322:

[0133] S22321. If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment region based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment region based on the fused adjustment text vector to obtain a first picture-text adjustment instruction set;

[0134] In this embodiment, after determining the local editing mode, the painting processing model will combine the fused adjustment text vector and the first picture description vector to identify the target adjustment area that needs to be adjusted, and generate specific adjustment instructions for the target adjustment area to obtain the first picture and text adjustment instruction set.

[0135] S22322. If the attention map is of a global distribution type, the global editing mode is determined by the large painting processing model, and based on the global editing mode, the fusion adjustment text vector is used to generate and process the picture reconstruction instructions to obtain a second picture and text adjustment instruction set.

[0136] In this embodiment, after determining the global editing mode, the large-scale drawing processing model generates image reconstruction instructions based on the fused and adjusted text vectors to obtain a second set of image and text adjustment instructions. This instruction set contains all the operations and information required to implement global editing, guiding image editing software or algorithms to reconstruct and adjust the original image. Compared to local editing, global editing involves more extensive changes, including redesigning the entire image style, tone, layout, or content.

[0137] In some embodiments, step S3 “showing the target adjustment picture to the user” may include the following steps S31 to S33:

[0138] S31, obtaining scene context information of the Wensheng picture scene;

[0139] Contextual scene information refers to all information involved in the image adjustment process, including user input and system output. This includes, but is not limited to, the user's input reference image, image adjustment instructions, and target image, as well as the intermediate results and suggestions generated by the system when processing these inputs. This information collectively constitutes the context of the user's drawing task and provides a foundation for understanding and predicting user intent.

[0140] S32: using a large painting processing model to determine the user's adjustment intention corresponding to the scene context information, and generating multiple picture extension adjustment schemes for the target adjustment picture based on the user's adjustment intention;

[0141] This embodiment uses a large painting processing model to infer image features that the user may wish to adjust, such as color saturation, brightness, composition, and level of detail, based on scene context information. This model then outputs one or more representations of the user's adjustment intent, which can be in the form of vectors, labels, or text, to guide subsequent image generation or adjustment processes. Based on the user's adjustment intent, multiple image extension schemes are generated for the target image. These image extension schemes are improvements or variations on the original generated image, providing a variety of options.

[0142] For example, in Figure 2 In the corresponding example, the user's picture adjustment instruction instructs to put a hat on the elephant, so the picture expansion adjustment solution can be to change the type of hat for the elephant.

[0143] For example, in Figure 2 In the user's example, if we further change the background of the picture, then the picture expansion adjustment solution can be to change the background of the picture to a city background with a different painting style. For example, please refer to Figure 3 and Figure 4 , Figure 3 The middle one is a cartoon-style city background picture. Figure 4 In the middle is a Van Gogh-style city background.

[0144] S33: Display the target adjustment picture and the picture expansion adjustment plan to the user.

[0145] In this embodiment, after generating multiple image expansion and adjustment proposals, the generated target image and the proposed image expansion and adjustment proposals are presented to the user. The user can then select the image that best meets their expectations from these options or further adjust their description based on these suggestions to achieve a more satisfactory image generation result. This embodiment demonstrates the auxiliary role of artificial intelligence in artistic creation. By understanding and predicting user intent, it provides personalized image generation and optimization suggestions, thereby enhancing the user's creative experience and satisfaction.

[0146] In some embodiments, step S32 of "using the painting processing macro model to determine the user adjustment intention corresponding to the scene context information, and generating multiple picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention" may include the following steps S321-S322:

[0147] S321: Determine context-related scene features and a predicted painting style using a large painting processing model based on scene context information, and infer potential painting intentions based on the context-related scene features and the predicted painting style to obtain user adjustment intentions;

[0148] Contextual scene-related features refer to key attributes or indicators that reflect specific information about the user's current drawing scene or task. These features are extracted based on contextual scene information and are used to help understand the user's drawing environment and needs. In the context of drawing processing, contextual scene-related features may include, but are not limited to: User-input reference generated images: The user may provide one or more reference images, which may contain elements such as the style, color, and composition that the user wants to emulate; Image adjustment instructions: The user may provide specific adjustment requirements through text, voice, or a graphical interface, such as "increase brightness" or "change color style"; Target adjustment image: The user's desired final drawing style or effect, which can be imagined by the user or expressed through other means (such as text descriptions or example images); System output information: Intermediate results or suggestions generated by the system when processing user input, such as automatically generated sketches and color schemes. These features collectively constitute the context of the user's current drawing task and serve as the basis for understanding user intent and subsequent processing.

[0149] Among them, the predicted painting style refers to the painting style or technique that the user may want to adopt, predicted by the large painting processing model based on the contextual scene association features. Painting style generally refers to the unique characteristics or trends of a painting in terms of color, lines, composition, etc. Predicting painting style involves analyzing the user's historical preferences, current contextual information, and the model's learning and understanding of the painting style database. For example, if a user often paints landscapes and prefers to use a watercolor style, the system may predict based on these historical data that the user also tends to use a watercolor style for painting in the current scenario. At the same time, the system may also combine the user's current input information (such as the style of the reference-generated painting) to further fine-tune the prediction results.

[0150] In an embodiment of the present application, a large painting processing model is used based on scene context information to determine context scene association features and predicted painting styles. Based on the context scene association features and predicted painting styles, potential painting intentions are inferred to obtain user adjustment intentions. The user's possible needs for adjustments can be predicted, thereby better meeting the user's painting needs and providing personalized painting assistance services.

[0151] S322: Determine an object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item based on the user's adjustment intention, and combine the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item to obtain multiple picture extension adjustment schemes.

[0152] Object Adjustment Extensions allow you to adjust or modify objects in an image or drawing. This can include changing an object's size, shape, position, quantity, type, and other attributes. For example, adding or removing trees in a landscape painting, or changing the style of a building, would all fall under the Object Adjustment Extensions.

[0153] The Lighting and Shadows extension involves adjusting the lighting and shadows within an image. This can include the position, intensity, and color of light sources, as well as the depth and direction of shadows. Adjusting the lighting and shadows can significantly impact the atmosphere and mood of an image. For example, shifting the sunlight from the left to the right in a painting, or adding a warmer glow at dusk, are examples of Lighting and Shadows extensions.

[0154] Color style adjustments refer to adjustments or changes to the overall or local color of an image. This may include changing the hue from cool to warm, increasing or decreasing saturation, adjusting contrast, and more. Color style adjustments can imbue an image with different emotional expressions and artistic styles. For example, converting a color photo to black and white or increasing the saturation of an image to emphasize the colors are both examples of color style adjustments.

[0155] In this embodiment, after determining the user's adjustment intent, the specific adjustments to be made to the objects, lighting, and color scheme in the image are determined based on the user's adjustment intent. Based on these adjustment intents, the system or software then generates multiple possible image expansion adjustment schemes, each combining different adjustment options for the objects, lighting, and color scheme. The user can then select the one from these image expansion adjustment schemes that best suits their needs.

[0156] In one example, let's assume the target image is an interior painting, and the user's predicted adjustment intent is to adapt it to a more modern style. For the object adjustment options, the user can choose to add modern furniture or remove traditional ornaments. For the lighting and shadow environment adjustment options, the user can add cooler lighting to create a modern feel. For the color style adjustment options, the user can reduce color saturation and increase contrast to create a cleaner and brighter image. By combining these adjustment options, the system can generate multiple modern-style image adjustment options for the interior painting, allowing the user to choose.

[0157] In one embodiment, a drawing adjustment processing device is also provided. Figure 5 , Figure 5 This is a schematic diagram of the structure of a painting adjustment processing device 400 provided in an embodiment of the present application. The painting adjustment processing device 400 is applied to an electronic device and includes a receiving module 401, a generating module 402, and a display module 403, as follows:

[0158] The receiving module 401 is used to receive a user's picture adjustment instruction for a reference generated picture in a text-generated picture scene;

[0159] A generation module 402 is configured to parse the picture adjustment instruction using a painting processing macromodel to obtain a picture adjustment description text for the reference generated picture, perform picture adjustment processing on the reference generated picture based on the picture adjustment description text, and generate a target adjusted picture for the reference generated picture;

[0160] The display module 403 is configured to display the target adjustment picture to the user.

[0161] In some embodiments, the generating module 402 is specifically configured to:

[0162] Acquire a first picture description text for the reference generated picture through the painting processing large model, wherein the first picture description text is a picture description after content description processing is performed on the reference generated picture in advance;

[0163] Performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode;

[0164] Based on the graphic adjustment instruction set, the target adjusted picture is obtained by performing picture adjustment through the painting processing model.

[0165] In some embodiments, the generating module 402 is specifically configured to:

[0166] determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for the reference generated picture;

[0167] Using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector;

[0168] The fused adjustment text vector is subjected to a graphic-text adjustment decoding process to obtain a graphic-text adjustment mode, and a description instruction conversion process is performed based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

[0169] In some embodiments, the generating module 402 is specifically configured to:

[0170] performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map;

[0171] A graphic and text adjustment mode is determined based on the distribution type of the adjustment attention map, and a graphic and text adjustment instruction set is obtained by performing instruction generation processing using a fused adjustment text vector based on the graphic and text adjustment mode.

[0172] In some embodiments, the generating module 402 is specifically configured to:

[0173] If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment area based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment area based on the fused adjustment text vector to obtain a first picture and text adjustment instruction set;

[0174] If the attention map is of a global distribution type, the global editing mode is determined by the painting processing model, and based on the global editing mode, the fused adjustment text vector is used to generate a picture reconstruction instruction to obtain a second picture and text adjustment instruction set.

[0175] In some embodiments, the display module 403 is specifically used to:

[0176] Obtaining scene context information of the cultural image scene;

[0177] Determining a user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention;

[0178] The target adjustment picture and the picture expansion adjustment solution are presented to the user.

[0179] In some embodiments, the display module 403 is specifically used to:

[0180] Determining context scene association features and a predicted painting style using the painting processing large model based on the scene context information, and performing potential painting intention inference based on the context scene association features and the predicted painting style to obtain a user adjustment intention;

[0181] An object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item are determined based on the user adjustment intention, and a plurality of picture extension adjustment schemes are obtained by combining the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item.

[0182] It should be noted that the painting adjustment processing device provided in the embodiment of the present application and the painting adjustment processing method in the above embodiment have the same concept. Any method provided in the painting adjustment processing method embodiment can be implemented through the painting adjustment processing device. The specific implementation process is detailed in the painting adjustment processing method embodiment and will not be repeated here.

[0183] In addition, in order to better implement the painting adjustment processing method in the embodiment of the present application, based on the painting adjustment processing method, the present application also provides an electronic device, which can be a smart phone, tablet computer, PDA, laptop computer, or desktop computer. Figure 6 , Figure 6 This is a schematic diagram of a first structure of an electronic device provided in an embodiment of the present application. The electronic device 500 includes a processor 501 and a memory 502. The processor 501 is electrically connected to the memory 502.

[0184] The processor 501 is the control center of the electronic device 500. It connects the various parts of the entire electronic device using various interfaces and lines. By running or calling computer programs stored in the memory 502, as well as calling data stored in the memory 502, it executes various functions of the electronic device and processes data, thereby monitoring the entire electronic device. The processor 501 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0185] Memory 502 can be used to store computer programs and data. The computer programs stored in memory 502 contain instructions that can be executed by the processor. Computer programs can be composed of various functional modules. Processor 401 executes various functional applications and data processing by calling the computer programs stored in memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area can store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created based on the use of the electronic device 500 (such as audio data, video data, etc.). In addition, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0186] In this embodiment, the processor 501 in the electronic device 500 loads instructions corresponding to one or more computer program processes into the memory 502 according to the following steps, and the processor 501 runs the computer program stored in the memory 502 to implement various functions:

[0187] In the Wensheng picture scenario, receiving a user's picture adjustment instruction for a reference generated picture;

[0188] parsing the picture adjustment instruction using a large painting processing model to obtain a picture adjustment description text for the reference generated picture, performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture;

[0189] The target adjustment picture is presented to the user.

[0190] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of a second structure of an electronic device provided in an embodiment of the present application. The electronic device 500 further includes a radio frequency circuit 503, a display screen 504, a control circuit 505, an input unit 506, an audio circuit 507, a sensor 508, and a power supply 509. The processor 501 is electrically connected to the radio frequency circuit 503, the display screen 504, the control circuit 505, the input unit 506, the audio circuit 507, the sensor 508, and the power supply 509, respectively.

[0191] The radio frequency circuit 503 is used to transmit and receive radio frequency signals to communicate with network devices or other electronic devices through wireless communication.

[0192] The display screen 504 may be used to display information input by a user or information provided to a user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces may be composed of images, texts, icons, videos, and any combination thereof.

[0193] The control circuit 505 is electrically connected to the display screen 504 and is used to control the display screen 504 to display information.

[0194] The input unit 506 may be configured to receive input numbers, characters, or user characteristics (e.g., fingerprints), and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control. The input unit 506 may include a fingerprint recognition module.

[0195] The audio circuit 507 can provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 507 includes a microphone. The microphone is electrically connected to the processor 501. The microphone is used to receive voice information input by the user.

[0196] The sensor 508 is used to collect external environment information. The sensor 508 may include one or more sensors such as an ambient brightness sensor, an acceleration sensor, and a gyroscope.

[0197] The power supply 509 is used to supply power to various components of the electronic device 500. In some embodiments, the power supply 509 can be logically connected to the processor 501 through a power management system, so that the power management system can manage charging, discharging, and power consumption.

[0198] Although not shown in the figure, the electronic device 500 may also include a camera, a Bluetooth module, etc., which will not be described in detail here.

[0199] In this embodiment, the processor 501 in the electronic device 500 loads instructions corresponding to one or more computer program processes into the memory 502 according to the following steps, and the processor 501 runs the computer program stored in the memory 502 to implement various functions:

[0200] In the Wensheng picture scenario, receiving a user's picture adjustment instruction for a reference generated picture;

[0201] parsing the picture adjustment instruction using a large painting processing model to obtain a picture adjustment description text for the reference generated picture, performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture;

[0202] The target adjustment picture is presented to the user.

[0203] In some embodiments, when the processor 501 performs the picture adjustment processing on the reference generated picture based on the picture adjustment description text and generates a target adjusted picture for the reference generated picture, the processor 501 may perform the following:

[0204] Acquire a first picture description text for the reference generated picture through the painting processing large model, wherein the first picture description text is a picture description after content description processing is performed on the reference generated picture in advance;

[0205] Performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode;

[0206] Based on the graphic adjustment instruction set, the target adjusted picture is obtained by performing picture adjustment through the painting processing model.

[0207] In some embodiments, when the processor 501 performs the image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image to obtain an image-text adjustment mode and generates an image-text adjustment instruction set based on the image-text adjustment mode, the processor 501 may execute:

[0208] determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for the reference generated picture;

[0209] Using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector;

[0210] The fused adjustment text vector is subjected to a graphic-text adjustment decoding process to obtain a graphic-text adjustment mode, and a description instruction conversion process is performed based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

[0211] In some embodiments, when the processor 501 performs the image-text adjustment decoding process on the fused adjustment text vector to obtain the image-text adjustment mode, and performs the description instruction conversion process based on the image-text adjustment mode to obtain the image-text adjustment instruction set, the following steps may be performed:

[0212] performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map;

[0213] A graphic and text adjustment mode is determined based on the distribution type of the adjustment attention map, and a graphic and text adjustment instruction set is obtained by performing instruction generation processing using a fused adjustment text vector based on the graphic and text adjustment mode.

[0214] In some embodiments, when the processor 501 determines the image-text adjustment mode based on the distribution type of the attention map adjustment, and generates instructions based on the image-text adjustment mode using the fused adjustment text vector to obtain the image-text adjustment instruction set, the processor 501 may execute:

[0215] If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment area based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment area based on the fused adjustment text vector to obtain a first picture and text adjustment instruction set;

[0216] If the attention map is of a global distribution type, the global editing mode is determined by the painting processing model, and based on the global editing mode, the fused adjustment text vector is used to generate a picture reconstruction instruction to obtain a second picture and text adjustment instruction set.

[0217] In some embodiments, when the processor 501 performs the step of displaying the target adjustment picture to the user, the processor 501 may execute:

[0218] Obtaining scene context information of the cultural image scene;

[0219] Determining a user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention;

[0220] The target adjustment picture and the picture expansion adjustment solution are presented to the user.

[0221] In some embodiments, when the processor 501 determines the user adjustment intention corresponding to the scene context information using the painting processing large model and generates multiple picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention, the processor 501 may perform the following:

[0222] Determining context scene association features and a predicted painting style using the painting processing large model based on the scene context information, and performing potential painting intention inference based on the context scene association features and the predicted painting style to obtain a user adjustment intention;

[0223] An object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item are determined based on the user adjustment intention, and a plurality of picture extension adjustment schemes are obtained by combining the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item.

[0224] An embodiment of the present application further provides a computer-readable storage medium, wherein the storage medium stores a computer program. When the computer program is run on a computer, the computer executes the painting adjustment processing method described in any of the above embodiments.

[0225] It should be noted that, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing related hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the storage medium may include but is not limited to: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0226] Furthermore, the terms "first," "second," and "third," etc., in this application are used to distinguish between different objects, not to describe a specific order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but rather some embodiments may include steps or modules not listed, or other steps or modules that are inherent to such process, method, product, or apparatus.

[0227] The above describes in detail the painting adjustment processing method, device, storage medium, and electronic device provided in the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is intended only to help understand the method and core concept of the present application. Furthermore, those skilled in the art will appreciate that variations in the specific implementation methods and scope of application may occur based on the concepts of the present application. In summary, the contents of this specification should not be construed as limiting the present application.

Claims

1. A painting processing method, characterized in that: include: In the Wensheng picture scenario, receiving a user's picture adjustment instruction for a reference generated picture; parsing the picture adjustment instruction using a large painting processing model to obtain a picture adjustment description text for the reference generated picture, performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture; The target adjustment picture is presented to the user.

2. The method according to claim 1, characterized in that The performing picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture includes: Acquire a first picture description text for the reference generated picture through the painting processing large model, wherein the first picture description text is a picture description after content description processing is performed on the reference generated picture in advance; Performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image using the painting processing macro model to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode; Based on the graphic adjustment instruction set, the target adjusted picture is obtained by performing picture adjustment through the painting processing model.

3. The method according to claim 2, characterized in that The step of performing image-text adjustment alignment processing based on the image adjustment description text, the first image description text, and the reference generated image to obtain an image-text adjustment mode, and generating an image-text adjustment instruction set based on the image-text adjustment mode includes: determining a picture adjustment description vector corresponding to the picture adjustment description text and a first picture description vector corresponding to the first picture description text, and determining a reference image coding vector for the reference generated picture; Using an image-text attention mechanism to perform image-text adjustment fusion processing on the reference image encoding vector based on the picture adjustment description vector and the first picture description vector to obtain a fused adjusted text vector; The fused adjustment text vector is subjected to a graphic-text adjustment decoding process to obtain a graphic-text adjustment mode, and a description instruction conversion process is performed based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set.

4. The method according to claim 3, characterized in that The performing a graphic-text adjustment decoding process on the fused adjustment text vector to obtain a graphic-text adjustment mode, and performing a description instruction conversion process based on the graphic-text adjustment mode to obtain a graphic-text adjustment instruction set, including: performing attention mapping processing on the first picture description vector based on the fused adjusted text vector to obtain an adjusted attention map; A graphic and text adjustment mode is determined based on the distribution type of the adjustment attention map, and a graphic and text adjustment instruction set is obtained by performing instruction generation processing using a fused adjustment text vector based on the graphic and text adjustment mode.

5. The method according to claim 4, characterized in that The determining of the image-text adjustment mode based on the distribution type of the attention map adjustment, and performing instruction generation processing based on the image-text adjustment mode using the fused adjustment text vector to obtain an image-text adjustment instruction set, including: If the attention map is of a local distribution type, determining a local editing mode using the painting processing large model, determining a target adjustment area based on the local editing mode using the fused adjustment text vector and the first picture description vector, and performing regional adjustment instruction generation processing on the target adjustment area based on the fused adjustment text vector to obtain a first picture and text adjustment instruction set; If the attention map is of a global distribution type, the global editing mode is determined by the painting processing model, and based on the global editing mode, the fused adjustment text vector is used to generate a picture reconstruction instruction to obtain a second picture and text adjustment instruction set.

6. The method according to claim 1, characterized in that The presenting the target adjustment picture to the user includes: Obtaining scene context information of the cultural image scene; Determining a user adjustment intention corresponding to the scene context information using the painting processing macro model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention; The target adjustment picture and the picture expansion adjustment solution are presented to the user.

7. The method according to claim 6, characterized in that The step of determining the user adjustment intention corresponding to the scene context information by using the painting processing large model, and generating a plurality of picture extension adjustment schemes for the target adjustment picture based on the user adjustment intention, includes: Determining context scene association features and a predicted painting style using the painting processing large model based on the scene context information, and performing potential painting intention inference based on the context scene association features and the predicted painting style to obtain a user adjustment intention; An object adjustment extension item, a light and shadow environment adjustment extension item, and a color style adjustment extension item are determined based on the user adjustment intention, and a plurality of picture extension adjustment schemes are obtained by combining the target adjustment picture based on the object adjustment extension item, the light and shadow environment adjustment extension item, and the color style adjustment extension item.

8. A painting adjustment processing device, characterized in that: include: A receiving module, configured to receive, in a text-generated image scene, an image adjustment instruction from a user for a reference generated image; a generation module configured to parse the picture adjustment instruction using a painting processing macromodel to obtain a picture adjustment description text for the reference generated picture, and perform picture adjustment processing on the reference generated picture based on the picture adjustment description text to generate a target adjusted picture for the reference generated picture; A display module is used to display the target adjustment picture to the user.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed on a computer, the computer is caused to execute the painting adjustment method according to any one of claims 1 to 7.

10. An electronic device comprising a processor and a memory, wherein the memory stores a computer program, wherein: The processor is configured to execute the painting adjustment processing method according to any one of claims 7 to 8 by calling the computer program.