A painting processing method, device, storage medium and electronic equipment
By determining the painting style pattern and guiding the painting style characteristics in the large language model, the painting processing flow is optimized, which solves the problem of the existing technology that the painting effect does not meet the user's expectations, and generates high-quality paintings that better meet user expectations.
Patent Information
- Application Number
- CN202410612576.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-05-16
AI Technical Summary
In the existing technology, the painting effect generated by the large-model Wenshengtu technology is far from the overall effect expected by the user, and it is difficult to meet the user's painting needs.
By obtaining the initial painting prompt words input by the user, the painting style mode is determined, and the painting language model is used to process the text image in this mode. The painting style recommendation model and style guidance information are combined to optimize the painting processing process and generate a target painting that better meets the user's expectations.
The painting processing flow has been optimized to make the generated paintings more effective in painting style mode, more in line with the user's painting expectations, and generate higher quality target paintings.
Smart Images

Figure CN118918199B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a painting processing method, device, storage medium, and electronic device. Background Art
[0002] With the development of computer technology, electronic devices such as smart watches and smart bracelets are rapidly becoming popular. Electronic devices usually have the function of intelligent voice question and answer, and support the function of waking up the device through voice, including making calls, setting alarms, opening applications, playing music stories, etc., and can also perform drawing processing. Summary of the Invention
[0003] The present application provides a method, device, storage medium, and electronic device for processing a painting. The technical solution is as follows:
[0004] In a first aspect, an embodiment of the present application provides a method for processing a painting, the method comprising:
[0005] Obtaining an initial painting prompt word input by a user, determining a painting style mode corresponding to the initial painting prompt word, and performing text-image processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting;
[0006] A painting display process is performed to the user based on the target painting.
[0007] In a feasible implementation manner, determining the painting style mode corresponding to the initial painting prompt word includes:
[0008] Determining painting scene description data for the user;
[0009] A style pattern matching process is performed on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word.
[0010] In a feasible implementation manner, determining the painting scene description data for the user includes:
[0011] Acquiring the user's painting style interest information, and generating painting scene description data for the user based on the painting style interest information; and / or,
[0012] A painting theme for the user is determined based on the initial painting prompt word, user attribute information and a historical painting style for the user are determined, and painting scene description data for the user is generated based on the painting theme, the user attribute information and the historical painting style.
[0013] In a feasible implementation manner, performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word includes:
[0014] Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word;
[0015] The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
[0016] In a feasible implementation manner, the step of performing text-based image processing based on the initial painting prompt word through the painting language model to obtain a target painting in the painting style mode includes:
[0017] determining painting style guidance information corresponding to the painting style mode, and performing painting style prompt enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word;
[0018] In the painting style mode, the target painting prompt word is input into the painting language model for text-image processing to obtain the target painting.
[0019] In a feasible implementation, determining the painting style guidance information corresponding to the painting style mode, and performing painting style guidance enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word includes:
[0020] Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern;
[0021] The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
[0022] In a feasible embodiment, the method further includes:
[0023] Performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or
[0024] A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
[0025] In a feasible embodiment, the method further includes:
[0026] Use the basic large language model to create an initial painting large language model for the painting generation scenario;
[0027] Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels;
[0028] The sample painting prompt words are used to perform model transfer training on the initial painting language model.
[0029] During the model transfer training process, a sample painting style pattern corresponding to the sample painting prompt word is determined, and under the sample painting style pattern, a cultural image processing is performed on the sample painting prompt word through the painting large language model to obtain a predicted painting, and model parameters of the initial painting large language model are adjusted based on the predicted painting and the painting label to obtain a painting large language model after transfer training.
[0030] In a second aspect, an embodiment of the present application provides a drawing processing device, the device comprising:
[0031] a Wenshengtu module, configured to obtain an initial painting prompt word input by a user, determine a painting style mode corresponding to the initial painting prompt word, and perform Wenshengtu processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting;
[0032] The painting display module is used to display the target painting to the user.
[0033] In a feasible implementation, the cultural graph module includes:
[0034] a description determining unit, configured to determine painting scene description data for the user;
[0035] A painting processing unit is configured to perform style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word.
[0036] In a feasible implementation manner, the description determining unit is configured to:
[0037] Acquiring the user's painting style interest information, and generating painting scene description data for the user based on the painting style interest information; and / or,
[0038] A painting theme for the user is determined based on the initial painting prompt word, user attribute information and a historical painting style for the user are determined, and painting scene description data for the user is generated based on the painting theme, the user attribute information and the historical painting style.
[0039] In a feasible implementation manner, the painting processing unit is configured to:
[0040] Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word;
[0041] The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
[0042] In a feasible implementation, the cultural graph module is used to:
[0043] a prompt enhancement unit, configured to determine painting style guidance information corresponding to the painting style mode, and perform painting style prompt enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word;
[0044] In the painting style mode, the target painting prompt word is input into the painting language model for text-image processing to obtain the target painting.
[0045] In a feasible implementation manner, the prompt enhancement unit is configured to:
[0046] Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern;
[0047] The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
[0048] In a feasible embodiment, the device is further used for:
[0049] Performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or
[0050] A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
[0051] In a feasible embodiment, the device is further used for:
[0052] Use the basic large language model to create an initial painting large language model for the painting generation scenario;
[0053] Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels;
[0054] The sample painting prompt words are used to perform model transfer training on the initial painting large language model. During the model transfer training process, the sample painting style mode corresponding to the sample painting prompt words is determined. Under the sample painting style mode, the sample painting prompt words are used to perform cultural image processing through the painting large language model to obtain a predicted painting. Based on the predicted painting and the painting label, the model parameters of the initial painting large language model are adjusted to obtain the painting large language model after transfer training.
[0055] In a third aspect, an embodiment of the present application provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.
[0056] In a fourth aspect, an embodiment of the present application provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
[0057] The beneficial effects of the technical solutions provided by some embodiments of the present application include at least:
[0058] In one or more embodiments of the present application, the initial painting prompt word input by the user is obtained, and the painting style mode corresponding to the initial painting prompt word is determined. In the painting style mode, the target painting is obtained by performing text-generated image processing based on the initial painting prompt word through the painting large language model, rather than directly indicating the initial painting prompt word to instruct the large model to perform text-generated image. In this way, the painting large language model is guided in advance to fully refer to the style characteristics of the painting style mode, and the painting processing flow is optimized, so that the overall effect of generating the painting in this painting style mode is better, more in line with the user's painting expectations, and a target painting with better quality can be generated for display to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0060] Figure 1 This is a flow chart of a painting processing method provided in an embodiment of the present application;
[0061] Figure 2 This is a schematic diagram of a text-based graph interface provided in an embodiment of the present application;
[0062] Figure 3 This is a schematic diagram of a scene for processing a cultural image provided by an embodiment of the present application;
[0063] Figure 4 This is a flow chart of another embodiment of a painting processing method provided in an embodiment of the present application;
[0064] Figure 5 This is a schematic diagram of a model training process for a large painting language model provided in an embodiment of the present application;
[0065] Figure 6 1 is a structural diagram of a painting processing device provided in an embodiment of the present application;
[0066] Figure 7 This is a schematic diagram of the structure of a Vincent diagram provided in an embodiment of the present application;
[0067] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0068] Figure 9 This is a schematic diagram of the structure of the operating system and user space provided in an embodiment of the present application;
[0069] Figure 10 yes Figure 9 The architecture diagram of the Android operating system;
[0070] Figure 11 yes Figure 9 Architecture diagram of the IOS operating system. DETAILED DESCRIPTION
[0071] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0072] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to the specific circumstances. In addition, in the description of this application, unless otherwise specified, "multiple" refers to two or more. "and / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.
[0073] With the development of large-scale models and artificial intelligence (AIGC) technologies, AI is increasingly being used in bank marketing, poster generation, web page background images, and even PowerPoint presentations. However, current large-scale model-based text-generated image technology has some issues, primarily resulting in the overall effect of generated images often differing significantly from user expectations.
[0074] The present application is described in detail below with reference to specific embodiments.
[0075] In one embodiment, Figure 1 As shown, a drawing processing method is proposed. This method can be implemented using a computer program and can be run on a drawing processing device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The drawing processing device can be an electronic device, including but not limited to a personal computer, tablet computer, handheld device, vehicle-mounted device, wearable device, computing device, or other processing device connected to a wireless modem. Terminal devices in different networks can be called different names, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, terminal device in 5G network or future evolution network, etc.; electronic devices can also be service platforms.
[0076] Specifically, the painting processing method includes:
[0077] S102: Obtaining the initial drawing prompt word input by the user;
[0078] The initial painting prompt word is a description of the painting generated by the user input. The initial painting prompt word includes descriptive information of the painting that the user expects to generate, such as the image style description, image content description, image specification description and other descriptive information of the expected image. The initial painting prompt word is used by the user to instruct the smart object service to generate painting content that meets the user's expectations.
[0079] The initial drawing prompt is usually a descriptive text entered by the user. Based on the initial drawing prompt, a target drawing with corresponding content can be automatically generated with or without a sample image (the user can optionally enter a sample image). In other words, the user can choose whether to enter a sample image.
[0080] The example image is a painting reference provided for the painting language model to generate corresponding content based on the painting prompt word. It is used to instruct the painting language model to generate a target painting that matches the example painting characteristics and corresponds to the painting prompt word with reference to the example painting characteristics of the example image. In this way, the target painting can be highly consistent with the example image in terms of example painting characteristics such as image depth and layout.
[0081] Schematically, the electronic device obtains the initial painting prompt word input by the user. The initial painting prompt word can carry an example image when the user inputs an example image. Subsequently, the painting style mode corresponding to the initial painting prompt word is determined. In the painting style mode, the target painting is obtained by performing text-based image processing based on the initial painting prompt word with reference to the example image through the painting large language model.
[0082] Illustratively, the initial drawing prompt word may not carry an example image when the user does not input an example image.
[0083] The initial drawing prompt word can be a text input by the user, or it can be a text converted from the user's voice. Figure 2 As shown, Figure 2 It is a schematic diagram of a cultural graph interface. Figure 2 In the interface shown, the user can choose to manually input or voice input the initial drawing prompt word; this specification does not impose any restrictions on the form of obtaining the initial drawing prompt word and the content of the initial drawing prompt word.
[0084] S104: determining a painting style mode corresponding to the initial painting prompt word, and performing text-image processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting;
[0085] The painting style mode may be one or more of a plurality of pre-configured reference painting style modes, such as a universal painting style mode, a two-dimensional painting style mode, a science fiction painting style mode, a traditional Chinese painting style mode, and the like.
[0086] For example, the user may pre-select a painting style mode for this painting, and then determine the painting style mode corresponding to the initial painting prompt word;
[0087] For example, style analysis can be performed based on the initial painting prompt words and / or example images, and the initial painting prompt words and / or example images provided by the user can be analyzed. These initial painting prompt words and / or example images may contain painting features related to the theme, description, painting details, scene description, and picture scale. By analyzing the painting features, the user's desired painting style mode can be further clarified, rather than directly indicating the initial painting prompt words to instruct the large model to perform text drawing. In this way, the painting large language model is pre-guided into the painting style mode, which can make the overall effect of generating paintings in this painting style mode better.
[0088] Furthermore, a style parsing model is pre-trained based on the machine learning model, and the initial painting prompt words and / or sample images are input into the style parsing model to extract painting features. The painting features are analyzed by the style parsing model to determine the painting style pattern, and the painting style pattern is output.
[0089] Optionally, after determining the painting style mode, the painting language model is controlled to enter the painting style mode, and the target painting is obtained by performing text-image processing based on the initial painting prompt words.
[0090] For example, a large painting language model that can support various painting style texts is pre-trained for different painting style modes. After determining the painting style mode, the style mode prompt words corresponding to the painting style mode are input into the large painting language model. Then, the large painting language model of the current painting style mode can be selected to control the large painting language model to enter the current painting style mode. In the current painting style mode, text processing is performed based on the initial painting prompt words to obtain the target painting.
[0091] For example, Figure 3 As shown, Figure 3 It is a scene diagram of Wensheng picture processing. After the electronic device determines the painting style mode, it can be displayed on the user input interface as follows Figure 3 The text of the painting style mode shown is Figure 3In the "Painting Style: Chinese Style (Mode)", the user can further adjust the painting style mode, such as adjusting from Chinese style (Mode) to realistic style. After the user determines the current painting style mode and enters the initial painting prompt word "draw a monkey catching the moon", he can start painting, such as clicking the start painting button, and then controlling the painting language model to enter the painting style mode. In this painting style mode, the text image is processed based on the initial painting prompt word to obtain the target painting that fits the painting style mode.
[0092] In a feasible implementation, the drawing large language model may directly use the basic large language model.
[0093] In one possible implementation, the painting large language model can be trained on a painting scenario based on a basic large language model (LLM). The basic large language model (LLM) is an artificial intelligence content generation model designed to understand and generate human language. The LLM is trained on a large amount of data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and more.
[0094] Optionally, the basic large language model can be the AIGC model on the public network;
[0095] Optionally, you can obtain a trained basic large language model and adapt the basic large language model to the painting generation scenario. First, obtain the basic large language model to create an initial painting large language model, and obtain sample painting prompt words in the new painting scenario. Use the sample painting prompt words to fine-tune the initial painting large language model. After the model fine-tuning training is completed, you can obtain a painting large language model adapted to the painting generation scenario.
[0096] S106: Performing a painting display process to the user based on the target painting.
[0097] In some embodiments, the user terminal has an input function that can be used for user input of data. For example, typing input, voice input, etc. The service platform can be centralized or distributed. In some embodiments, the user terminal can send the initial painting prompt word input by the user and the determined painting style mode to the service platform to instruct the service platform to perform text-based image processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting. The user terminal then obtains the target painting and displays the target painting to the user.
[0098] In some embodiments, a user terminal (e.g., a smart terminal) receives an initial painting prompt word input by a user, determines a painting style mode, performs text-image processing based on the initial painting prompt word using a painting language model in the painting style mode to obtain a target painting, and displays the target painting to the user;
[0099] In a feasible implementation, after obtaining the initial drawing prompt word input by the user, a prompt sensitivity detection may be performed on the initial drawing prompt word to obtain prompt sensitivity information, and prompt correction processing may be performed on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing;
[0100] Optionally, a content-sensitive constraint library is set up. The content-sensitive constraint library is established and maintained based on the content-sensitive constraint target. The content-sensitive constraint target may be to avoid generating drawings involving sensitive content. By writing some keywords involving the content-sensitive constraint target into the content-sensitive constraint library, the content-sensitive constraint library can be used in actual applications to perform prompt sensitivity detection on the initial drawing prompt words to obtain prompt sensitivity information.
[0101] Then, based on the prompt sensitive information, the initial drawing prompt word is modified to obtain the initial drawing prompt word after the prompt modification;
[0102] For example, if it is determined that the initial drawing prompt contains sensitive keywords based on the prompt sensitive information, the sensitive keywords of the initial drawing prompt are deleted.
[0103] For example, sensitive keywords can be replaced with regular keywords. When establishing a content-sensitive constraint library, corresponding regular keywords are set for keywords related to content-sensitive constraint targets. In actual applications, regular keywords corresponding to corresponding sensitive keywords can be obtained from the content-sensitive constraint library based on prompt sensitive information, and then sensitive keywords can be replaced with regular keywords.
[0104] In a feasible implementation, painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and painting correction processing is performed on the target painting based on the painting sensitivity information to obtain the target painting after painting correction processing.
[0105] In one or more embodiments of the present application, the initial painting prompt word input by the user is obtained, and the painting style mode corresponding to the initial painting prompt word is determined. In the painting style mode, the target painting is obtained by performing text-generated image processing based on the initial painting prompt word through the painting large language model, rather than directly indicating the initial painting prompt word to instruct the large model to perform text-generated image. In this way, the painting large language model is guided in advance to fully refer to the style characteristics of the painting style mode, and the painting processing flow is optimized, so that the overall effect of generating the painting in this painting style mode is better, more in line with the user's painting expectations, and a target painting with better quality can be generated for display to the user.
[0106] See Figure 4 , Figure 4 This is a flow chart of another embodiment of a painting processing method proposed in this application. Specifically:
[0107] S202: Obtaining the initial drawing prompt word input by the user;
[0108] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here.
[0109] S204: Determine painting scene description data for the user;
[0110] The painting scene description data can be understood as metadata that describes the user's cultural scene. The painting scene description data can include painting themes, painting style interest information, historical painting styles, user attributes (such as age, identity, knowledge level, etc.), and other data that can describe the user's cultural scene.
[0111] In this specification, by introducing painting scene description data into the initial painting prompt words originally input by the user to adjust the prompt words, the target painting prompt words finally input into the large model can include components describing the user's literary and artistic scene and the matching painting style guidance information. Subsequently, the large model can be instructed that the literary and artistic scene and painting style can be converted into literary and artistic scenes more accurately, thereby improving the efficiency and accuracy of painting generation.
[0112] In a feasible implementation, obtaining the user's painting style interest information, and generating painting scene description data for the user based on the painting style interest information;
[0113] The user's historical painting data is collected and painting style interest information is identified from this historical painting data. This information is identified by analyzing the user's past painting data within the historical painting data. For example, the user's preferred colors, lines, composition, themes, and other elements are identified and summarized as the user's unique painting style interest information. In some embodiments, this painting style interest information can be stored as the painting style interest information matched to the current user. In practical applications, the user's painting style interest information can be directly obtained to generate a series of painting scene description data that matches the user's preferences. This description data includes elements such as colors of interest, lighting of interest, composition of interest, objects of interest, and themes of interest, thereby forming a painting scene description data unique to the user.
[0114] It can be understood that by obtaining the user's painting style interest information to generate painting scene description data, the initial painting prompt words can be modified later to provide the user with more personalized painting prompt words that are in line with their painting style interest information, so as to better express the user's creativity and ideas during the painting process.
[0115] Exemplary:
[0116] In a feasible implementation, a painting theme for the user can be determined based on the initial painting prompt words, user attribute information and historical painting style for the user can be determined, and painting scene description data for the user can be generated based on the painting theme, the user attribute information and the historical painting style.
[0117] Painting theme: The theme semantics of the initial painting prompt words can be analyzed to obtain the painting theme generated by the user's intention;
[0118] User attribute information: including the (current) user's age, gender, occupation, educational background, interests, etc. User attribute information can provide clues about the user's possible interests, knowledge level, and experience;
[0119] Historical painting style: Identify the user's painting style based on the user's historical painting data. Specifically, the historical painting style can be determined by performing a style analysis on the painting elements such as color, line, composition, and theme in the historical painting data.
[0120] The painting theme, user attribute information, and historical painting styles are then combined to generate the user's painting scene description data. This painting scene description data is then used to determine the painting style pattern that matches the user's interests.
[0121] S206: Performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word.
[0122] For example, multiple reference painting style patterns are pre-defined, and a painting style library is established by collecting paintings of different artistic painting styles. Key painting features, such as color matching, line usage, composition, and subject matter, are extracted from the paintings in the style library. Each reference painting style pattern is then defined based on the extracted key painting features.
[0123] Specifically, key description elements are extracted from the painting scene description data, and the extracted key description elements are matched with the defined reference painting style patterns to obtain style pattern similarity, and the painting style pattern indicated by the highest style pattern similarity is selected.
[0124] In a feasible implementation, performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word may be:
[0125] Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word;
[0126] The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
[0127] For example, the following illustrates the model training process of a painting style recommendation model:
[0128] Model creation: Create a painting style recommendation model based on the machine learning model;
[0129] Sample data acquisition: Acquire sample painting scene description data, which is the painting scene description data in the model training phase;
[0130] Sample data annotation: Expert services are used to annotate sample painting scene description data with painting style pattern labels;
[0131] Model training process: Sample painting scene description data is input into the initial painting style recommendation model for at least one round of model training. During each round of model training, the initial painting style recommendation model recommends a painting style for the sample painting scene description data to obtain a predicted painting style pattern. Based on the predicted query scene description and the painting style pattern label, a model loss value is determined using a model loss function. Based on the model loss value, model parameters of the initial painting style recommendation model are adjusted until the model training end conditions are met to obtain a painting style recommendation model.
[0132] Optionally, the model loss function can be a hinge loss function, a cross entropy loss function, a contrast loss function, etc.
[0133] Optionally, the model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. Specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.
[0134] It should be noted that the machine learning models involved in one or more embodiments of this specification include but are not limited to the fitting of one or more machine learning models such as convolutional neural network (CNN) model, deep neural network (DNN) model, recurrent neural network (RNN) model, embedding model, gradient boosting decision tree (GBDT) model, logistic regression (LR) model, etc.
[0135] S208: Determine painting style guidance information corresponding to the painting style mode, and perform painting style guidance enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word;
[0136] Illustratively, for one or more reference painting style modes, reference painting style guidance information of each reference painting style mode is predefined, such as color preference guidance words, line and brushstroke guidance words, composition suggestion guidance words, light and shadow guidance words, and texture guidance words, etc.;
[0137] After the painting style pattern is determined, a predefined set of painting style guidance information is extracted or obtained for the painting style pattern, and then the painting style guidance information is added to the initial painting prompt word to enhance the painting style prompt, thereby obtaining the target painting prompt word.
[0138] Optionally, the painting style guidance information may include painting style positive guidance words and / or painting style negative guidance words;
[0139] In a feasible implementation, the following method may be used:
[0140] A2: Determine a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern;
[0141] For example, based on the painting style guidance information and the content of the initial painting prompt words, the following prompt word enhancement processing is performed:
[0142] Color enhancement: Add color description words that match the painting style to the prompt words, such as the positive guide words "use warm colors to depict the sunset" or the negative guide words "don't use cold colors to depict the sunset"
[0143] Brushstroke and line tips: Add tips and guides for brushstroke and line usage, such as "Use rough brushstrokes to express the texture of the rocks";
[0144] Composition guidance: Give specific composition suggestions, such as positive guidance words "focus on the center of the picture, surrounded by trees and rivers"; negative guidance words "focus on the periphery of the picture"
[0145] Light and shadow and texture description: Add descriptive words about light and shadow and texture, such as the positive guide word "emphasize the contrast of light and shadow to highlight the three-dimensional sense of the building";
[0146] Emphasis on themes and elements: If the painting style has specific themes or elements, you can emphasize these elements in the prompts, such as the positive guide "Add flying birds to the picture to express the theme of freedom"; the negative guide "Do not add fallen leaves to the picture to express the theme of desolation"
[0147] The enhanced content is then integrated into the initial drawing prompt to generate the target drawing prompt. This ensures that the target prompt contains both the core content of the initial drawing prompt and elements of drawing style guidance, thereby providing more accurate guidance for drawing creation.
[0148] A4: Based on the painting style positive guiding words and / or the painting style negative guiding words, the initial painting prompt words are optimized in terms of prompt word style to obtain target painting prompt words.
[0149] Exemplary:
[0150] Initial drawing prompt: Draw a cabin in the forest with sunlight shining through the leaves.
[0151] Painting style mode:
[0152] Impressionism
[0153] Painting style guidance information:
[0154] Color preference: warm and bright colors, emphasizing the changes in light and shadow, the colors are not cold, and it is forbidden to paint a cold atmosphere.
[0155] Brushstrokes and lines: Use loose brushstrokes to avoid fine brushwork and capture the fluidity of light and shadow.
[0156] Drawing prompt: Use warm, bright colors to depict a cabin in the forest. Sunlight filtering through the leaves creates a dappled effect. Use loose brushstrokes to capture the fluidity of light and shadow, emphasizing the play of light and shadow between the sunlight and the leaves.
[0157] S210: Inputting the target painting prompt word into the painting language model in the painting style mode to perform text-image processing to obtain a target painting.
[0158] For example, the generated target painting prompt words have advantages for the painting language model, such as improved accuracy, enhanced style consistency, stimulated creativity, increased flexibility, and enhanced interpretability. These advantages enable the painting language model to more accurately understand user needs, generate images that meet the requirements, and play an important role in a wide range of cultural image applications.
[0159] Improved Accuracy: By combining the target painting prompt with the initial description and style guidance, the large model can more accurately understand the user's intent. This accuracy is reflected not only in the accuracy of the content but also in the accuracy of the style. The large model can more accurately grasp the scene, emotion, and style that the user wants to express, thereby generating images that better meet the user's needs.
[0160] Enhanced style consistency: By adding guidance on painting style, the target painting prompt word ensures that the generated images are stylistically consistent. The model can generate images with consistent stylistic characteristics based on the stylistic requirements of the prompt word, improving the overall integrity and aesthetics of the output.
[0161] Creative Stimulation: The style guidance provided by the target painting prompt can also stimulate the model's creativity. By introducing different painting styles and elements, the large model can incorporate more creativity and inspiration into the image generation process. This not only enriches the image's expressiveness but also meets users' demand for novel and unique images.
[0162] Improved Flexibility: Because the target image prompts are generated based on the initial description and style guidance, they are highly flexible. Users can modify the initial description or add different style guidance as needed to generate images with different content and styles. This flexibility makes the model applicable to a wider range of application scenarios and user needs.
[0163] Enhanced interpretability: By generating target image prompts, the model can more clearly explain the process and basis for generating images. This helps enhance the model's interpretability, allowing users to better understand its working principles and outputs. It also helps identify potential problems and flaws in the model's generation process, allowing for targeted optimization and improvement.
[0164] S212: Performing a painting display process to the user based on the target painting.
[0165] For details, please refer to the method steps in other embodiments of this specification, which will not be repeated here.
[0166] In one or more embodiments of the present application, the target painting prompt words generated using the above method have advantages for the painting language model, such as improved accuracy, enhanced style consistency, stimulated creativity, increased flexibility, and enhanced interpretability. These advantages enable the painting language model to more accurately understand user needs, generate images that meet the requirements, and play an important role in a wide range of cultural image application scenarios.
[0167] See Figure 5 , Figure 5 This is a schematic diagram of the model training process of a large painting language model proposed in this application. Specifically:
[0168] S302: creating an initial painting large language model for the painting generation scenario using the basic large language model;
[0169] Specifically, a basic large language model is obtained; based on the basic large language model, an initial painting large language model is created, which includes at least a painting generation scene adaptation module and a large language generation module.
[0170] Optionally, different painting style generation scene adaptation modules can be created for different reference painting style modes; for example, a universal painting style generation scene adaptation module corresponding to the universal painting style mode, a two-dimensional painting style generation scene adaptation module corresponding to the two-dimensional painting style mode, and so on; subsequently, sample painting prompt words of different reference painting styles are used for model transfer training.
[0171] A trained basic large language model can be obtained, and the basic large language model can be adapted to the painting generation scenario to obtain a painting large language model. Usually, the basic large language model is directly applied to generate different painting styles in the painting generation scenario, which is difficult to adapt to the new painting scenario. Based on this, the basic large language model is first obtained to create an initial painting large language model, and sample painting prompt words in the new painting scenario or sample painting prompt words of different reference painting styles in the obtained painting scenario are obtained to perform model migration training for different painting styles. Since the basic large language model is usually an open source AIGC model that has been trained and has content generation capabilities, based on this, this manual only needs to adapt it to the painting scenario. Specifically, the sample data can be used to perform model fine-tuning training on the initial painting large language model, and after the model fine-tuning training is completed, a painting large language model adapted to the painting scenario is obtained.
[0172] S304: Obtain sample painting prompt words and label the sample painting prompt words with painting labels;
[0173] Sample data acquisition: obtain sample painting prompt words, call expert services to label the sample painting prompt words with painting labels (usually painting images drawn by expert services based on the sample painting prompt words); the sample painting prompt words can be sourced from historical cultural image databases, the Internet, etc.
[0174] S306: Using the sample painting prompt words to perform model transfer training on the initial painting language model,
[0175] S308: During the model transfer training process, a sample painting style pattern corresponding to the sample painting prompt word is determined, and based on the sample painting prompt word, a cultural image processing is performed through the painting large language model to obtain a predicted painting in the sample painting style pattern, and model parameters of the initial painting large language model are adjusted based on the predicted painting and the painting label to obtain a painting large language model after transfer training.
[0176] The model transfer training process involves inputting sample painting prompts into the initial painting large language model for at least one round of model training to obtain a predicted painting. A comprehensive model loss is calculated based on the predicted painting and the painting label. Based on this comprehensive model loss, the model parameters of the painting scene adaptation module in the initial painting large language model are adjusted, while the model parameters of the large language generation module remain unchanged. This process continues until the model training termination conditions are met, resulting in the large language generation module and the painting scene adaptation module. This completes the model fusion of the large language generation module and the painting scene adaptation module, resulting in a trained painting large language model.
[0177] In some embodiments, different painting style generation scene adaptation modules can be created for different reference painting style modes; for example, a universal painting style generation scene adaptation module corresponding to a universal painting style mode, a two-dimensional painting style generation scene adaptation module corresponding to a two-dimensional painting style mode, and so on; subsequently, model transfer training is performed in stages using sample painting prompt words of different reference painting styles, and a painting style generation scene adaptation module that completes the end condition of the model training in the completion stage is obtained. At this time, a combination of "n painting style generation scene adaptation modules + 1 large language generation module" can be obtained. After the combination is completed, a trained painting large language model can be obtained. In actual application deployment, the current painting style generation scene adaptation module corresponding to the painting style mode is selected according to "determining the painting style mode corresponding to the initial painting prompt word", and the current painting style generation scene adaptation module is linked and integrated with the painting scene adaptation module, thereby completing the control of the painting large language model to enter the current painting style mode;
[0178] The link fusion of the large language generation module and the painting scene adaptation module is to fuse the model structure layer weights of a painting scene adaptation module with the large language generation module, and by determining the target model structure layer corresponding to the model structure layer weight in the large language generation module, the model structure layer parameters of the target model structure layer are fused with the model structure layer weights. The model structure layer weights of the painting scene adaptation module may only partially correspond to and have model structure layer weights in all model structure layers in the basic large language model. By completing the parameter update of the model structure layer based on the model structure layer weights for this part of the target model structure layer, the reference update process of all model structure layer weights is completed by analogy, thereby completing the control of the painting large language model to enter the current painting style mode.
[0179] Optionally, the model training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, the number of iterations reaching a preset number threshold, etc. Specific model training termination conditions may be determined based on actual conditions and are not specifically limited here.
[0180] In one or more embodiments of the present application, the above-mentioned method can be used to train a large painting language model. The large painting language model can be used to perform text-generated image processing based on the initial painting prompt words through the large painting language model in the painting style mode to obtain a target painting, rather than directly indicating the initial painting prompt words to instruct the large model to perform text-generated image. In this way, the large painting language model is guided in advance to fully refer to the style characteristics of the painting style mode, and the painting processing flow is optimized, so that the overall effect of generating paintings in this painting style mode is better, more in line with the user's painting expectations, and a target painting with better quality can be generated for display to the user.
[0181] The following will be combined Figure 6 , the painting processing device provided in the embodiment of the present application is introduced in detail. It should be noted that, Figure 6 The painting processing device shown is used to execute the present application Figures 1 to 5 For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the present application. Figures 1 to 5 The embodiment shown.
[0182] See Figure 6 , which shows a schematic diagram of the structure of a painting processing device according to an embodiment of the present application. The painting processing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the painting processing device 1 includes a Wensheng image module 11 and a painting display module 12, which are specifically used to:
[0183] The Wenshengtu module 11 is configured to obtain an initial painting prompt word input by a user, determine a painting style mode corresponding to the initial painting prompt word, and perform Wenshengtu processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting;
[0184] The painting display module 12 is configured to display the target painting to the user.
[0185] Optional, such as Figure 7 As shown, the cultural graph module 11 includes:
[0186] A description determining unit 111, configured to determine painting scene description data for the user;
[0187] The painting processing unit 112 is configured to perform style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word.
[0188] Optionally, the description determining unit 111 is configured to:
[0189] Acquiring the user's painting style interest information, and generating painting scene description data for the user based on the painting style interest information; and / or,
[0190] A painting theme for the user is determined based on the initial painting prompt word, user attribute information and a historical painting style for the user are determined, and painting scene description data for the user is generated based on the painting theme, the user attribute information and the historical painting style.
[0191] Optionally, the painting processing unit 112 is configured to: input the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and output a painting style model corresponding to the initial painting prompt word;
[0192] The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
[0193] Optionally, the cultural graph module 11 is used to:
[0194] a prompt enhancement unit 113 configured to determine painting style guidance information corresponding to the painting style mode, and perform painting style prompt enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word;
[0195] In the painting style mode, the target painting prompt word is input into the painting language model for text-image processing to obtain the target painting.
[0196] Optionally, the prompt enhancement unit 113 is configured to:
[0197] Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern;
[0198] The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
[0199] Optionally, the device 1 is further used for:
[0200] Performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or
[0201] A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
[0202] Optionally, the device 1 is further used for:
[0203] Use the basic large language model to create an initial painting large language model for the painting generation scenario;
[0204] Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels;
[0205] The sample painting prompt words are used to perform model transfer training on the initial painting large language model. During the model transfer training process, the sample painting style mode corresponding to the sample painting prompt words is determined. Under the sample painting style mode, the sample painting prompt words are used to perform cultural image processing through the painting large language model to obtain a predicted painting. Based on the predicted painting and the painting label, the model parameters of the initial painting large language model are adjusted to obtain the painting large language model after transfer training.
[0206] It should be noted that the above-described embodiments of the image processing device, when executing the image processing method, illustrate the division of the aforementioned functional modules only as an example. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, i.e., the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the above-described embodiments of the image processing device and the image processing method embodiments share the same concept. Their implementation is detailed in the method embodiments and will not be further elaborated here.
[0207] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0208] In an embodiment of the present application, by obtaining the initial painting prompt word input by the user, the painting style mode corresponding to the initial painting prompt word is determined, and the target painting is obtained by performing text-generated image processing based on the initial painting prompt word through the painting large language model in the painting style mode, rather than directly indicating the initial painting prompt word to instruct the large model to perform text-generated image. In this way, the painting large language model is guided in advance to fully refer to the style characteristics of the painting style mode, and the painting processing flow is optimized, so that the overall effect of generating the painting in this painting style mode is better, more in line with the user's painting expectations, and a target painting with better quality can be generated for display to the user.
[0209] The present application also provides a computer storage medium that can store multiple instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1 to 5 The specific execution process of the painting processing method in the embodiment shown can be found in Figures 1 to 5 The detailed description of the illustrated embodiment will not be repeated here.
[0210] The present application also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 5 The specific execution process of the painting processing method in the embodiment shown can be found in Figures 1 to 5 The detailed description of the illustrated embodiment will not be repeated here.
[0211] Please refer to Figure 8 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. The electronic device in the present application may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
[0212] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device 100 and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.
[0213] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.
[0214] See also Figure 9As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
[0215] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0216] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows: Figure 10As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware components of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides major feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as games, instant messaging programs, and photo enhancement programs.
[0217] Taking the operating system as the IOS system as an example, the programs and data stored in the memory 120 are as follows: Figure 11As shown, the IOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by applications, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides applications with audio-visual interfaces, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technologies. The touchable layer 480 provides various commonly used interface-related frameworks for application development. The touchable layer 480 is responsible for user touch interaction operations on electronic devices. For example, local notification service, remote push service, advertising framework, game tool framework, message user interface (UI) framework, user interface UIKit framework, map framework, etc.
[0218] exist Figure 11 Among the frameworks shown, those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480. The Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI. The classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
[0219] Among them, the method and principle of implementing data communication between third-party applications and operating system in the IOS system can be referred to the Android system, and this application will not go into details here.
[0220] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are a touch screen display, which is used to receive touch operations on or near it by the user using any suitable object such as a finger or a touch pen, and to display the user interface of each application. The touch screen display is usually provided on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of the present application.
[0221] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which are not described in detail here.
[0222] In the embodiments of the present application, the execution subject of each step can be the electronic device described above. Optionally, the execution subject of each step is the operating system of the electronic device. The operating system can be an Android system, an iOS system, or other operating systems, which are not limited in the embodiments of the present application.
[0223] The electronic device of the embodiment of the present application may further be equipped with a display device, which may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device 101 to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing, and the like.
[0224] exist Figure 8 In the electronic device shown, the processor 110 may be configured to call an application stored in the memory 120 and specifically perform the following operations:
[0225] Obtaining an initial painting prompt word input by a user, determining a painting style mode corresponding to the initial painting prompt word, and performing text-image processing based on the initial painting prompt word using the painting language model in the painting style mode to obtain a target painting;
[0226] A painting display process is performed to the user based on the target painting.
[0227] In one embodiment, the processor 110 performs the following operations when determining the painting style mode corresponding to the initial painting prompt word:
[0228] Determining painting scene description data for the user;
[0229] A style pattern matching process is performed on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word.
[0230] In one embodiment, the processor 110 performs the following steps when determining the painting scene description data for the user:
[0231] Acquiring the user's painting style interest information, and generating painting scene description data for the user based on the painting style interest information; and / or,
[0232] A painting theme for the user is determined based on the initial painting prompt word, user attribute information and a historical painting style for the user are determined, and painting scene description data for the user is generated based on the painting theme, the user attribute information and the historical painting style.
[0233] In one embodiment, the processor 110 performs the following steps when performing the style pattern matching process on the painting scene description data to obtain the painting style pattern corresponding to the initial painting prompt word:
[0234] Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word;
[0235] The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
[0236] In a feasible implementation manner, in the painting style mode, based on the initial painting prompt word, the target painting is obtained by performing text-image processing using the painting language model, and the following steps are performed:
[0237] determining painting style guidance information corresponding to the painting style mode, and performing painting style prompt enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word;
[0238] In the painting style mode, the target painting prompt word is input into the painting language model for text-image processing to obtain the target painting.
[0239] In one embodiment, the processor 110, after determining the painting style guidance information corresponding to the painting style mode and performing painting style enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word, performs the following steps:
[0240] Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern;
[0241] The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
[0242] In one embodiment, the processor 110 further performs the following steps when executing the method:
[0243] Performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or
[0244] A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
[0245] In one embodiment, when executing the drawing processing method, the processor 110 further performs the following steps:
[0246] Use the basic large language model to create an initial painting large language model for the painting generation scenario;
[0247] Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels;
[0248] The sample painting prompt words are used to perform model transfer training on the initial painting language model.
[0249] During the model transfer training process, a sample painting style pattern corresponding to the sample painting prompt word is determined, and under the sample painting style pattern, a cultural image processing is performed on the sample painting prompt word through the painting large language model to obtain a predicted painting, and model parameters of the initial painting large language model are adjusted based on the predicted painting and the painting label to obtain a painting large language model after transfer training.
[0250] In an embodiment of the present application, by obtaining the initial painting prompt word input by the user, the painting style mode corresponding to the initial painting prompt word is determined, and the target painting is obtained by performing text-generated image processing based on the initial painting prompt word through the painting large language model in the painting style mode, rather than directly indicating the initial painting prompt word to instruct the large model to perform text-generated image. In this way, the painting large language model is guided in advance to fully refer to the style characteristics of the painting style mode, and the painting processing flow is optimized, so that the overall effect of generating the painting in this painting style mode is better, more in line with the user's painting expectations, and a target painting with better quality can be generated for display to the user.
[0251] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0252] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.
Claims
1. A painting processing method, characterized in that: The method comprises: Get the initial drawing prompt word entered by the user; determining a painting theme for the user based on the initial painting prompt word, determining user attribute information and a historical painting style for the user, generating painting scene description data for the user based on the painting theme, the user attribute information, and the historical painting style, and performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word; determining painting style guidance information corresponding to the painting style mode, and performing painting style prompt enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word; In the painting style mode, the target painting prompt word is input into the painting language model to perform text-image processing to obtain a target painting; A painting display process is performed to the user based on the target painting.
2. The method according to claim 1, characterized in that The performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word includes: Key description elements are extracted from the painting scene description data, and the extracted key description elements are matched with the defined reference painting style patterns to obtain style pattern similarity, and the painting style pattern indicated by the highest style pattern similarity is selected.
3. The method according to claim 1, characterized in that The performing style pattern matching processing on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word includes: Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word; The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
4. The method according to claim 1, wherein The determining of the painting style guidance information corresponding to the painting style mode, and performing painting style guidance enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word, includes: Reference painting style guidance information for each reference painting style mode is predefined for multiple reference painting style modes, and the reference painting style guidance information includes color preference guidance words, line and brushstroke guidance words, composition suggestion guidance words, light and shadow guidance words, and texture guidance words; After the painting style mode is determined, a predefined set of painting style guidance information is obtained for the painting style mode, and the painting style guidance information is added to the initial painting prompt words to perform painting style prompt enhancement processing to obtain target painting prompt words.
5. The method according to claim 1, wherein The guide words in the reference painting style guide information include positive painting style guide words and / or negative painting style guide words, and determining the painting style guide information corresponding to the painting style mode, and performing painting style prompt enhancement processing on the initial painting prompt words based on the painting style guide information to obtain target painting prompt words, including: Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern; The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
6. The method according to claim 1, characterized in that The method further comprises: Performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
7. The method according to claim 1, characterized in that The method further comprises: Use the basic large language model to create an initial painting large language model for the painting generation scenario; Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels; The sample painting prompt words are used to perform model transfer training on the initial painting large language model. During the model transfer training process, the sample painting style mode corresponding to the sample painting prompt words is determined. Under the sample painting style mode, the sample painting prompt words are used to perform cultural image processing through the painting large language model to obtain a predicted painting. Based on the predicted painting and the painting label, the model parameters of the initial painting large language model are adjusted to obtain the painting large language model after transfer training.
8. A painting processing device, characterized in that: The device comprises: a Wenshengtu module configured to obtain an initial painting prompt word input by a user, determine a painting theme for the user based on the initial painting prompt word, determine user attribute information and a historical painting style for the user, generate painting scene description data for the user based on the painting theme, the user attribute information, and the historical painting style, and perform style pattern matching on the painting scene description data to obtain a painting style pattern corresponding to the initial painting prompt word; a Wenshengtu module, configured to determine the painting style guidance information corresponding to the painting style mode, and perform painting style guidance enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word; A Wenshengtu module is used to input the target painting prompt word into the painting language model in the painting style mode to perform Wenshengtu processing to obtain a target painting; The painting display module is used to display the target painting to the user.
9. The device according to claim 8, characterized in that The drawing processing unit is used to: Key description elements are extracted from the painting scene description data, and the extracted key description elements are matched with the defined reference painting style patterns to obtain style pattern similarity, and the painting style pattern indicated by the highest style pattern similarity is selected.
10. The device according to claim 8, characterized in that The drawing processing unit is used to: Inputting the painting scene description data into a painting style recommendation model to perform painting style recommendation processing, and outputting a painting style model corresponding to the initial painting prompt word; The painting style recommendation model is obtained by training an initial painting style recommendation model using sample painting scene description data that has been annotated with painting style pattern labels.
11. The device according to claim 8, characterized in that The determining of the painting style guidance information corresponding to the painting style mode, and performing painting style guidance enhancement processing on the initial painting prompt word based on the painting style guidance information to obtain a target painting prompt word, includes: Reference painting style guidance information for each reference painting style mode is predefined for multiple reference painting style modes, and the reference painting style guidance information includes color preference guidance words, line and brushstroke guidance words, composition suggestion guidance words, light and shadow guidance words, and texture guidance words; After the painting style mode is determined, a predefined set of painting style guidance information is obtained for the painting style mode, and the painting style guidance information is added to the initial painting prompt words to perform painting style prompt enhancement processing to obtain target painting prompt words.
12. The device according to claim 8, characterized in that The prompt enhancement unit is used to: Determining a painting style positive guiding word and / or a painting style negative guiding word corresponding to the painting style pattern; The initial painting prompt word is optimized based on the painting style positive guide word and / or the painting style negative guide word to obtain a target painting prompt word.
13. The device according to claim 8, characterized in that The device is also used for: performing prompt sensitivity detection on the initial drawing prompt word to obtain prompt sensitivity information, and performing prompt correction processing on the initial drawing prompt word based on the prompt sensitivity information to obtain the initial drawing prompt word after prompt correction processing; and / or A painting sensitivity detection is performed on the target painting to obtain painting sensitivity information, and a painting correction process is performed on the target painting based on the painting sensitivity information to obtain the target painting after the painting correction process.
14. The device according to claim 8, characterized in that The device is also used for: Use the basic large language model to create an initial painting large language model for the painting generation scenario; Obtaining sample painting prompt words, and labeling the sample painting prompt words with painting labels; The sample painting prompt words are used to perform model transfer training on the initial painting large language model. During the model transfer training process, the sample painting style mode corresponding to the sample painting prompt words is determined. Under the sample painting style mode, the sample painting prompt words are used to perform cultural image processing through the painting large language model to obtain a predicted painting. Based on the predicted painting and the painting label, the model parameters of the initial painting large language model are adjusted to obtain the painting large language model after transfer training.
15. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 7.
16. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 7.
Citation Information
Patent Citations
Image generation method and device, computer equipment and storage medium
CN116363242A
Ancient poetry artistic conception map generation method based on chain analysis
CN117274411A
Image generation method and device, terminal and storage medium
CN117392254A