Image generation method and device, equipment, storage medium and product
By acquiring user-input control keywords and environmental information, and combining the target image generation model and physical simulation layer, the problem of insufficient user control in existing technologies is solved, and more accurate and realistic antique painting generation is achieved.
Patent Information
- Application Number
- CN202510988265.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-21
AI Technical Summary
Existing image generation models lack flexible user control for non-professional users, resulting in generated images that fail to accurately reflect the user's specific intentions.
By acquiring user-input prompts, including control keywords such as the degree of aging, the number and location of seals, and the color and texture of the canvas, and inputting them into a preset image generation model, the model generates an image that conforms to the control keywords, using the target image generation model and physical simulation layer.
It enhances the user's control over image generation, making the generated antique-style images more accurately reflect the user's specific intentions and increasing the authenticity and historical feel of the generated images.
Smart Images

Figure CN120997628A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image generation technology, and in particular to an image generation method, apparatus, device, storage medium and product. Background Technology
[0002] In related technologies, common image generation models can synthesize high-quality images, but they often rely heavily on carefully designed prompts. For non-professional users, the lack of flexible user control makes it difficult for the generated images to accurately reflect the user's specific intentions. Summary of the Invention
[0003] The main objective of this application is to provide an image generation method, apparatus, device, storage medium, and product, aiming to solve the technical problem that images generated by related technologies are difficult to accurately reflect the user's specific intentions.
[0004] To achieve the above objectives, this application proposes an image generation method, which includes:
[0005] Obtain prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, the control keyword including the degree of aging, the number and position of stamps, and the color and texture of the canvas;
[0006] The prompt words are input into a preset image generation model to generate an image, resulting in a target antique image that matches the control keywords.
[0007] In one embodiment, the step of inputting the prompt words into a preset image generation model to generate an image that conforms to the control keywords includes:
[0008] The collected current environmental information and the prompt words are input into a preset image generation model to generate an image that conforms to the control keywords. The environmental information includes humidity and light intensity.
[0009] In one embodiment, the preset image generation model includes a target image generation model and a physical simulation layer. The step of inputting the collected current environmental information and the prompt words into the preset image generation model to generate an image that has a higher degree of aging simulation and conforms to the control keywords includes:
[0010] The target image generation model extracts features from the prompt words to obtain feature vectors. Based on the feature vectors, an initial antique image is generated. The feature vectors include keyword vectors, era vectors, and environment vectors.
[0011] The physical simulation layer determines the material properties of the canvas material corresponding to the age vector, and the aging mask is determined based on the age vector, the environment vector, and the material properties.
[0012] The aging mask is fused with the initial antique image to obtain an antique painting image with a higher degree of aging simulation than the initial simulation image and that meets the control keywords.
[0013] In one embodiment, the step of inputting the prompt word into a preset image generation model to generate an image and obtain a target antique image that conforms to the control keyword includes:
[0014] Obtain a set of ancient images with a consistent style, and extract content keywords from the ancient image set;
[0015] The design factors related to the control keywords in each image of the ancient image set are precisely labeled to obtain the ancient image set with the control keywords after precise labeling;
[0016] Based on the control keywords and the content keywords, generate sample prompt words corresponding to each image;
[0017] The sample prompts and the set of ancient images containing the control keywords are input into the first image generation model for iterative training to obtain the second image generation model after preliminary adjustment.
[0018] The second image generation model is fine-tuned to obtain a target image generation model that satisfies local precise control.
[0019] In one embodiment, the step of fine-tuning the second image generation model to obtain the target image generation model that satisfies local precise control includes:
[0020] In each image, a target mask of the corresponding type is generated based on the control keywords for the region corresponding to the design factor. The target mask includes an aging mask, a stamp mask, and a canvas mask.
[0021] Determine the target loss function, wherein the target loss function is constructed based on the base loss function and the decoupling loss function;
[0022] Based on the ancient image set with the target mask and the control keywords, the content keywords, and the target loss function, the second image generation model is fine-tuned to obtain a target image generation model that satisfies local precise control.
[0023] In one embodiment, the step of fine-tuning the second image generation model to obtain the target image generation model that satisfies local precise control includes:
[0024] A material attribute database is constructed based on the canvas material and the physical properties corresponding to the canvas material.
[0025] Based on the material property database and the initial aging model, a physical simulation layer is constructed.
[0026] Furthermore, to achieve the above objectives, this application also proposes an image generation apparatus, the image generation apparatus comprising:
[0027] The acquisition module acquires prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, and the control keyword includes the degree of aging, the number and position of stamps, and the color and texture of the canvas;
[0028] The generation module is used to input the prompt words into a preset image generation model to generate an image and obtain a target antique image that conforms to the control keywords.
[0029] In addition, to achieve the above objectives, this application also proposes an image generation apparatus, the apparatus comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image generation method as described above.
[0030] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the image generation method described above.
[0031] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the image generation method described above.
[0032] One or more technical solutions proposed in this application have at least the following technical effects:
[0033] While common image generation models in related technologies can synthesize high-quality images, they often heavily rely on carefully designed prompts. For non-professional users, this lack of flexible user control makes it difficult for the generated images to accurately reflect the user's specific intentions. In contrast, this application obtains user-input prompts, where each prompt includes at least one control keyword targeting preset design factors, such as the degree of aging, the number and location of stamps, and the canvas color and texture. The prompts are then input into a preset image generation model to generate a target antique image that conforms to the control keywords. This application pre-sets control keywords targeting preset design factors, including the degree of aging, the number and location of stamps, and the canvas color and texture. By inputting prompts containing these control keywords into the preset image generation model, a target antique image conforming to the control keywords is obtained. This preset control keywords help users select design factors without requiring carefully designed prompts, improving the user's ability to control prompts and preventing the generated image from failing to accurately reflect the user's specific intentions. Attached Figure Description
[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a flowchart illustrating an embodiment of the image generation method of this application.
[0037] Figure 2 This is a scene diagram illustrating the data acquisition process for the image generation method described in this application.
[0038] Figure 3 Examples of mild, moderate, and severe image generation methods of this application are shown;
[0039] Figure 4 The diagram illustrates the white, khaki, and dark brown colors used in the image generation method of this application.
[0040] Figure 5 This is a flowchart illustrating Embodiment 2 of the image generation method of this application;
[0041] Figure 6 This is a schematic diagram of the module structure of the image generation device according to an embodiment of this application;
[0042] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the image generation method in the embodiments of this application.
[0043] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0044] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0045] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0046] The main solution of this application embodiment is: to obtain prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, the control keyword including the degree of aging, the number and position of the seals, and the color and texture of the canvas; to input the prompt words into a preset image generation model to generate an image, thereby obtaining a target antique image that conforms to the control keyword.
[0047] In related technologies, common image generation models can synthesize high-quality images, but they often rely heavily on carefully designed prompts. For non-professional users, the lack of flexible user control makes it difficult for the generated images to accurately reflect the user's specific intentions.
[0048] This application pre-sets control keywords for preset design factors, including the degree of aging, the number and position of stamps, and the color and texture of the canvas. It then inputs prompts containing these control keywords into a preset image generation model to generate an image that conforms to the control keywords. This results in a target antique image that matches the control keywords. By using preset control keywords, the application helps users select design factors without requiring carefully designed prompts, thus improving the user's ability to control prompts and preventing the generated image from failing to accurately reflect the user's specific intentions.
[0049] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or image generation device capable of performing the above functions. The following description uses an image generation device as an example to illustrate this embodiment and the subsequent embodiments.
[0050] Based on this, embodiments of this application provide an image generation method, referring to... Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the image generation method of this application.
[0051] In this embodiment, the image generation method includes steps S10 to S20:
[0052] Step S10: Obtain prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, and the control keyword includes the degree of aging, the number and position of stamps, and the color and texture of the canvas.
[0053] It should be noted that the execution subject in this embodiment is an image generation device. This image generation device has a front-end page for receiving prompts input by the user. These prompts include not only control keywords but also the user's description of the main content of the image and the era of the antique finish. For example, the user-input prompts might be: "A meticulous brushwork painting from the Song Dynasty of China, a sparrow perched on a branch of a night-blooming jasmine," and control keywords including design elements such as "the canvas color is khaki," "there is a red seal in the right corner of the painting," and "lightly antiqued."
[0054] Understandably, antique-style portraits, as an art form combining modern technology and traditional culture, have a wide range of applications. Taking virtual museums and online exhibitions as examples, antique-style portraits, as core visual content, can greatly enrich the digital cultural experience, break the limitations of time and space, and allow more people to access and understand history and culture in an immersive and interactive way.
[0055] It should also be noted that the user-related data involved in this application (e.g., user-input prompts) was obtained with the user's permission or consent; that is, when this application is applied to a specific product or technology, user permission is required to obtain and process the relevant data, and the processing of the relevant data must comply with the relevant laws, regulations and regulatory standards of the relevant countries and regions.
[0056] For example, when it is necessary to obtain a prompt word input by the user, a prompt word retrieval prompt can be provided on the user's terminal. After receiving confirmation from the user regarding the prompt word retrieval prompt, the terminal can obtain the prompt word input by the user. As follows. Figure 2 As shown, Figure 2 A diagram illustrating the data acquisition scenario is provided.
[0057] It is understandable that, since this application has fixed control keywords, users only need to select control keywords or describe the content of the image to be generated after using preset control keywords, including a simple description of the era information, to generate an antique image that accurately meets the user's expectations. Therefore, step S10 can avoid the problem that the generated image is difficult to accurately reflect the user's specific intention, thereby improving the user's ability to control prompts.
[0058] Furthermore, the degree of distressing can be categorized into: light, medium, and heavy, as shown in the reference. Figure 3 , Figure 3 It provides sample images of light, medium and heavy ink, and the number and position of the stamps can be set according to the user's needs, as can the canvas color.
[0059] Step S20: Input the prompt words into a preset image generation model to generate an image, thereby obtaining a target antique image that conforms to the control keywords.
[0060] Understandably, controlling keywords helps the model more accurately understand the requirements and generate images that meet expectations. After receiving the prompt words, the image generation device inputs them into a preset image generation model to generate an image that meets the user's expectations.
[0061] In one feasible implementation, the step of inputting the prompt words into a preset image generation model to generate an image that conforms to the control keywords includes the following steps:
[0062] The collected current environmental information and the prompt words are input into a preset image generation model to generate an image that conforms to the control keywords. The environmental information includes humidity and light intensity.
[0063] It should be noted that humidity can affect the aging process of paper or fabric materials. In high humidity environments, paper or fabrics are more prone to mold and discoloration, while low humidity may cause them to become brittle. Long-term exposure to strong light can lead to color fading and material embrittlement. Different wavelengths of light also have different effects on materials. Therefore, to make the generated image more realistically reflect the aging effect, the model input can be adjusted based on the actual collected humidity and light intensity data. The image generation device combines the current environmental information and prompts to generate an image using a preset image generation model, resulting in an antique painting image with a higher degree of aging simulation and conforming to the control keywords.
[0064] In one feasible implementation, the steps of inputting the collected current environmental information and the prompt words into a preset image generation model to generate an image that has a higher degree of aging simulation and conforms to the control keywords include:
[0065] The target image generation model extracts features from the prompt words to obtain feature vectors. Based on the feature vectors, an initial antique image is generated. The feature vectors include keyword vectors, era vectors, and environment vectors.
[0066] Understandably, keyword vectors refer to the vector representations of content-related keywords extracted from prompts. Era vectors are specific to the time context of prompts, encoding era information into vector form. For example, the era vector for "Song Dynasty" reflects the artistic style, social culture, and other characteristics of that period, enabling the generated image to accurately reflect the unique features of that era. Environment vectors convert environmental information into vector form.
[0067] It should be noted that the target image generation model uses these feature vectors to guide the generation process, ensuring that the final output image not only conforms to the content description of the prompt words, but also has the corresponding time context characteristics and takes into account the influence of specific environmental factors.
[0068] Specifically, after receiving the complete text prompts provided by the user, the target image generation model performs necessary preprocessing. This includes text cleaning and word segmentation to remove irrelevant information that may interfere with the parsing process, ensuring that subsequent steps can proceed accurately. The preprocessed prompts are then fed into the model's core component for deep parsing, extracting control keywords related to preset design factors. These control keywords are used as control instructions to guide subsequent image generation. The identified control instructions are incorporated as conditions into the image generation process. Furthermore, based on these instructions, the target image generation model may adopt different strategies:
[0069] Modify latent space representation: Adjust the initial representation or generation path of the image in the latent space based on control commands.
[0070] Guided attention mechanism: Makes the model pay more attention to features related to instructions, so as to enhance the expressiveness of specific regions.
[0071] Activate a specific set of parameters: Select an appropriate subset or mode of parameters based on the control state to achieve the desired artistic effect.
[0072] Furthermore, guided by control commands, the target image generation model begins to synthesize images step by step. It not only generates the main content described by the user, but also accurately renders the required aging effect, seal features, and canvas color and texture according to the control commands, thereby generating an initial antique-style image.
[0073] Furthermore, the guidance provided by each control command to the target image generation model is as follows:
[0074] Aging Effects:
[0075] User input: The prompt words include instructions such as "[Antique Prompt Word] Mild", "[Antique Prompt Word] Moderate" or "[Antique Prompt Word] Severe".
[0076] Model Response: Based on training on images (and their labels) with different levels of aging, the model adjusts the overall tone of the generated image (e.g., a light aging might result in slightly yellowed paper, while a heavy aging would be significantly yellowed and darkened) and texture details (e.g., a light aging might have no obvious flaws, a moderate aging might have a few cracks and stains, while a heavy aging would have more dense and obvious cracks and stains). (Refer to...) Figure 3 The example outputs show "light", "medium" and "heavy" aging effects respectively.
[0077] Number and placement of seals (Seals):
[0078] User input: The prompts include commands such as "No stamp", "[Stamp prompt] There is a red stamp on the right side of the screen", or "[Stamp prompt] There are three red stamps on the left side of the screen".
[0079] Model Response: The model determines whether to generate seals, as well as the number and approximate location of the seals, based on the instructions. For example, for "three red seals on the left side of the image," the model will attempt to place three red seal-like shapes in the left area of the generated image. The style of these seals will also be kept as consistent as possible with the overall style of Song Dynasty paintings. Sample outputs of "no seals," "one red seal on the right side of the image," and "three red seals on the left side of the image" are shown respectively.
[0080] Canvas Color:
[0081] User input: The prompts include instructions such as "[Canvas color prompt] White", "[Canvas color prompt] Khaki" or "[Canvas color prompt] Dark brown".
[0082] Model Response: The model adjusts the dominant color tone of the background canvas area in the generated image. Additionally, since the canvas color labels in the training data may be associated with specific paper types (such as Xuan paper) and their conditions (such as cracks or stains), the model may implicitly generate corresponding textures and details when generating a canvas of a specified color. (See reference...) Figure 4 , Figure 4 Illustrations of white, khaki, and dark brown canvases are provided, showing sample outputs for the "white," "khaki," and "dark brown" canvases, respectively.
[0083] The physical simulation layer determines the material properties of the canvas material corresponding to the age vector, and the aging mask is determined based on the age vector, the environment vector, and the material properties.
[0084] Understandably, environmental factors reflect the impact of preservation conditions (such as temperature, humidity, and light) on the artwork. Material properties indicate how the canvas and paint respond to these environmental factors. An aging mask is used to represent the location and extent of color changes, crack formation, and other aging phenomena on the painting due to the passage of time and environmental influences. A physical simulation layer is used to simulate the physical properties of a specific material under various conditions, taking into account the material's basic properties (such as flexibility and absorbency) and the aging process over time (such as fading and embrittlement). By inputting a date vector, the image generation device allows the physical simulation layer to determine the specific material properties of the corresponding canvas material. Combining the date vector, environmental vector, and material properties, the physical simulation layer, through the resulting aging mask, can more accurately predict the changes a painting may undergo over the centuries following its creation.
[0085] Furthermore, although image generation devices can preliminarily define the degree of aging by controlling keywords, keywords are often based on mapping relationships to determine the corresponding aging style. This makes it difficult to accurately simulate the real effect of artworks aging over time by relying solely on keywords. The physical simulation layer can comprehensively consider various environmental factors (such as humidity and light intensity) and material properties (such as canvas type and pigment composition), thereby more accurately predicting and simulating the actual aging effect. Through physical simulation technology, not only can the aging characteristics on the appearance be simulated, but also changes in the internal structure can be taken into account, such as the brittleness of paper or fabric and the peeling of color layers, making the final antique image more visually realistic and possessing a higher sense of history and artistic value.
[0086] The aging mask is fused with the initial antique image to obtain an antique painting image with a higher degree of aging simulation than the initial simulation image and that meets the control keywords.
[0087] It should be noted that the aging mask is used to guide the aging of different areas of the image. The image generation device uses the aging mask as a weight map to locally modify the initial antique image, resulting in the modified and merged image. This modified image is an antique painting image with a higher degree of aging simulation and conforms to the control keywords. While retaining the original artistic style, reasonable aging details are added, making the work more narrative and visually appealing.
[0088] In this implementation, users can adjust different control keywords to achieve different needs for the final image, such as changing the degree of aging or the position of the stamp, which reduces the difficulty of designing prompts while providing a highly personalized creative experience.
[0089] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5Before step S20, the image generation method further includes steps S01 to S05:
[0090] Step S01: Obtain a set of ancient images with a consistent style, extract content from the ancient image set, and obtain content keywords;
[0091] It should be noted that the image generation device collects a dataset of traditional painting images covering multiple historical periods, art schools, and forms of expression. The images only need to cover one style (such as Song Dynasty meticulous brushwork) and composition type (with flowers, birds, fish, and insects as the theme), but can involve multiple materials (silk, paper, murals, etc.) and cover multiple types. Selecting high-quality images with a unified theme (such as "flowers, birds, fish, and insects") and consistent style (such as Song Dynasty meticulous brushwork paintings) as training samples helps the model focus on a specific art style, avoids style confusion caused by messy data, and ensures the professionalism and consistency of the generated content.
[0092] Furthermore, since Song Dynasty meticulous brushwork paintings are known for their exquisite depiction, precise marking allows the model to capture and reproduce these subtle details, such as the texture of feathers and the color gradations of petals, greatly enhancing the realism and artistic value of the work.
[0093] Furthermore, based on the understanding gained from in-depth learning of specific themes (such as flowers, birds, fish, and insects), the model may be able to more easily transfer the relevant skills to other similar but not exactly the same tasks, such as extending to the generation of landscape or figure paintings.
[0094] The image generation device analyzes the collected ancient images and extracts representative features or concepts, namely "content keywords". These keywords may include compositional information of the image, such as "sparrow", "standing", "night-blooming jasmine", and "branch".
[0095] Step S02: Accurately label the design factors related to the control keywords in each image of the ancient image set to obtain the ancient image set with the control keywords after accurate labeling;
[0096] Understandably, precise labeling refers to the process of annotating artistic features (i.e., design elements) in an image that can be specified by the user and that the user wishes to incorporate into the generated image. The image generation device manually or automatically labels certain design elements that can be clearly controlled by the user (such as the degree of aging, the number and location of stamps, the color and texture of the canvas, etc.), forming a data sample with tagged information.
[0097] Furthermore, the tag information is in a structured form: these precise tags can be organized into a structured metadata format, such as JSON or XML. The metadata for each image can contain the following fields:
[0098]
[0099] This structured, multi-dimensional labeling information is far richer than simple keyword labels. It provides the model with clear and distinguishable signals about specific artistic elements and their variations. The "precise labeling" of this invention enables the model to adjust its generation strategy based on these detailed "image labels." By "precisely tagging," visually perceptible artistic features (such as the degree of aging, the specific layout of the seal, and the specific color and state of the canvas) are associated with discrete or continuous parameters that the model can learn. This is a key prerequisite for achieving subsequent fine-grained control.
[0100] Specifically, the precise marking method is as follows:
[0101] Aging Effects:
[0102] The labeling should not only indicate whether there are signs of artificial aging, but also the degree of aging, such as "light," "moderate," or "severe." These degrees can be defined and distinguished based on visual characteristics (such as the degree of yellowing of the paper, wrinkles, damage, and the density and intensity of stains).
[0103] Labeling methods: Each training image or a specific region in an image (if applicable) can be assigned a category label (e.g., aging_degree: heavy) or a continuous value (if the model supports more granular control).
[0104] Seals:
[0105] The annotation content includes whether the seal appears, its approximate location, and the number of seals. For example, you can indicate the number of seals (e.g., 0, 1, 2, 3, etc.), the approximate location of the seals (e.g., upper left, lower right, right side of the image, left side of the image, etc.), and even the shape (square, round) and color (mainly red) of the seals.
[0106] Labeling method: You can use a bounding box to mark the position of the seal, and supplement it with attribute labels (such as seal_count:2, seal_location_1:top_right, seal_color_1:red).
[0107] Canvas Color and Properties:
[0108] The annotations should include: the background color of the canvas (e.g., "white", "khaki", "dark brown"), as well as the canvas texture and condition related to the color, such as the texture of the rice paper, "cracks", "minor stains", etc.
[0109] Labeling method: You can specify a primary color label for the background of the entire painting (such as canvas_color:khaki) and supplement it with descriptive labels to indicate the material and state of the canvas (such as canvas_texture:xuan_paper_cracked_stained).
[0110] Step S03: Based on the control keywords and the content keywords, generate sample prompt words corresponding to each image;
[0111] It should be noted that the image generation device has a deep understanding and organization of control keywords (such as the degree of aging, seal features, etc.) and content keywords (such as painting style, theme, etc.). For example, "Song Dynasty meticulous flower and bird painting" is used as a content keyword, combined with the control keyword "light aging" and the specific "seal position". The prompt word template designed according to different combination methods should have a certain degree of flexibility and adaptability so as to cover various possible combinations.
[0112] Step S04: Input the sample prompt words and the set of ancient images containing the control keywords into the first image generation model for iterative training to obtain the second image generation model after preliminary adjustment;
[0113] Understandably, the first image generation model is built upon Flux-based matching. Flux, as a relatively new architecture, with its rectified flowtransformer blocks, demonstrates strong capabilities in terms of image realism and adherence to cues. The image generation device trains the first image generation model on sample cues and the set of ancient images containing the control keywords to obtain a second image generation model adapted to the style of antique paintings.
[0114] Step S05: Fine-tune the second image generation model to obtain a target image generation model that satisfies local precise control.
[0115] It should be noted that by fine-tuning the image generation device with a larger than second image generation model, the model can be made to have the ability to precisely control and edit specific regions of the image, thereby obtaining the target image generation model.
[0116] Furthermore, the target image generation model supports generating images of any aspect ratio within the range of 1 million to 2 million pixels. This provides users with flexibility to generate images of appropriate size and aspect ratio according to different application needs (such as social media sharing, printing output, large-format display, etc.).
[0117] In one feasible implementation, the following steps are included after step S05:
[0118] In each image, a target mask of the corresponding type is generated based on the control keywords for the region corresponding to the design factor. The target mask includes an aging mask, a stamp mask, and a canvas mask.
[0119] Understandably, a target mask is a binary or grayscale image that marks the region containing a specific artistic element in an image at the pixel level, used to guide the model to perform specified control operations within these regions. The image generation device enhances the local control and semantic understanding capabilities of the image generation model by generating corresponding target masks (aging masks, stamp masks, canvas masks) based on control keywords in each image.
[0120] Determine the target loss function, wherein the target loss function is constructed based on the base loss function and the decoupling loss function;
[0121] It should be noted that the decoupling loss function includes aging control loss, stamp control loss, and canvas color control loss. When fine-tuning the image generation model, the image generation device must not only ensure that the generated image is semantically consistent with the prompt and visually realistic, but also achieve accurate local responses to specific artistic elements. Therefore, a target loss function that integrates the basic loss and decoupling loss is designed to guide the model to learn more effectively how these design factors are represented.
[0122] Based on the ancient image set with the target mask and the control keywords, the content keywords, and the target loss function, the second image generation model is fine-tuned to obtain a target image generation model that satisfies local precise control.
[0123] Understandably, before fine-tuning the model, image generation devices freeze some of the model's parameters (usually early layers of the model that are responsible for extracting general features) to help preserve the knowledge in the pre-trained model and reduce the risk of performance degradation due to overfitting.
[0124] Furthermore, the image generation device inputs the target mask and the set of ancient images and content keywords of control keywords into the second image generation model for training, so that the model can accurately generate or modify specific artistic elements in the specified area according to the user's instructions while maintaining the consistency of the overall style.
[0125] In one feasible implementation, after the step of fine-tuning the second image generation model to obtain the target image generation model that satisfies local precise control, the following steps are included:
[0126] A material attribute database is constructed based on the canvas material and the physical properties corresponding to the canvas material.
[0127] It should be noted that the image generation device establishes a database containing information on different canvas materials and their physical properties, providing data support for subsequent aging simulation and physical simulation.
[0128] Based on the material property database and the initial aging model, a physical simulation layer is constructed.
[0129] Understandably, the initial aging model is based on the AMaterial Aging Model, a simulation system based on materials science principles and historical aging data. It predicts and reproduces the physical and chemical changes that occur in different materials over time and under environmental influences. These changes include, but are not limited to, yellowing, oxidation, cracking, fading, and insect damage. The image generation device simulates age vectors and environmental variables at different aging levels. It can train the initial aging model using the age vectors, the corresponding material properties, and the environmental variables to obtain a physical simulation layer that accurately reflects the aging effect.
[0130] In this embodiment, by training a target image generation model with high controllability and artistic expression, it can not only respond to complex control commands from users, but also achieve highly realistic generation of antique paintings at both the visual and physical levels.
[0131] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image generation method of this application. Any simple transformations based on this technical concept are all within the protection scope of this application.
[0132] This application also provides an image generation apparatus, please refer to... Figure 6 The image generation apparatus includes:
[0133] The acquisition module 10 acquires the prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, and the control keyword includes the degree of aging, the number and position of the stamps, and the color and texture of the canvas;
[0134] The generation module 20 is used to input the prompt words into a preset image generation model to generate an image and obtain a target antique image that conforms to the control keywords.
[0135] Optionally, the generation module includes:
[0136] The input submodule is used to input the collected current environmental information and the prompt words into a preset image generation model to generate an image that conforms to the control keywords. The environmental information includes humidity and light intensity.
[0137] The training submodule is used to acquire a set of ancient images with a consistent style, extract content from the ancient image set to obtain content keywords, accurately label the design factors related to the control keywords in each image in the ancient image set to obtain the accurately labeled ancient image set with the control keywords, generate sample prompt words corresponding to each image based on the control keywords and the content keywords, input the sample prompt words and the ancient image set with the control keywords into the first image generation model for iterative training to obtain a preliminarily adjusted second image generation model, and fine-tune the second image generation model to obtain the target image generation model that satisfies local precise control.
[0138] Optionally, the input submodule includes:
[0139] The fusion unit is used to extract features from the prompt words using the target image generation model to obtain feature vectors, and generate an initial antique image based on the feature vectors. The feature vectors include keyword vectors, era vectors, and environment vectors. The physical simulation layer determines the material properties of the canvas material corresponding to the era vector, and determines an aging mask based on the era vector, environment vector, and material properties. The aging mask is fused with the initial antique image to obtain an antique painting image with a higher degree of aging simulation than the initial simulation image and that conforms to the control keywords.
[0140] Optionally, the training submodule includes:
[0141] A fine-tuning unit is used to generate a target mask of a corresponding type based on the control keywords for the regions corresponding to the design factors in each image, wherein the target mask includes an aging mask, a stamp mask, and a canvas mask; determine a target loss function, wherein the target loss function is constructed based on a base loss function and a decoupling loss function; and fine-tune the second image generation model based on the ancient image set with the target mask and the control keywords, the content keywords, and the target loss function to obtain a target image generation model that satisfies local precise control.
[0142] The construction unit is used to build a material attribute database based on the canvas material and the corresponding physical properties of the canvas material; and to build a physical simulation layer based on the material attribute database and the initial aging model.
[0143] The image generation apparatus provided in this application, employing the image generation method described in the above embodiments, can solve the technical problem of image generation. Compared with the prior art, the beneficial effects of the image generation apparatus provided in this application are the same as those of the image generation method described in the above embodiments, and other technical features in the image generation apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0144] This application provides an image generation apparatus, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the image generation method in Embodiment 1 above.
[0145] The following is for reference. Figure 7 The diagram illustrates a structural schematic of an image generation device suitable for implementing embodiments of this application. The image generation device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, tablets, digital broadcast receivers, PDAs (Personal Digital Assistants), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The image generation device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0146] like Figure 7As shown, the image generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the image generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the image generating device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows image generating devices with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0147] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0148] The image generation device provided in this application, employing the image generation method described in the above embodiments, can solve the technical problem of image generation. Compared with the prior art, the beneficial effects of the image generation device provided in this application are the same as those of the image generation method described in the above embodiments, and other technical features of the image generation device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0149] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0150] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0151] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the image generation method described in the above embodiments.
[0152] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0153] The aforementioned computer-readable storage medium may be included in the image generating apparatus or may exist independently without being assembled into the image generating apparatus.
[0154] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by an image generation device, cause the image generation device to: acquire prompts input by a user, wherein the prompts include at least one control keyword for preset design factors, the control keyword including the degree of aging, the number and position of stamps, and the canvas color and texture; input the prompts into a preset image generation model to generate an image, thereby obtaining a target antique image that conforms to the control keywords.
[0155] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0156] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0157] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0158] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described image generation method, thereby solving the technical problem of image generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the image generation method provided in the above embodiments, and will not be repeated here.
[0159] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the image generation method described above.
[0160] The computer program product provided in this application can solve the technical problem of image generation. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the image generation method provided in the above embodiments, and will not be repeated here.
[0161] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.
Claims
1. An image generation method, characterized in that, The image generation method includes: Obtain prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, the control keyword including the degree of aging, the number and position of stamps, and the color and texture of the canvas; The prompt words are input into a preset image generation model to generate an image, resulting in a target antique image that matches the control keywords.
2. The image generation method as described in claim 1, characterized in that, The step of inputting the prompt words into a preset image generation model to generate an image that conforms to the control keywords includes: The collected current environmental information and the prompt words are input into a preset image generation model to generate an image that conforms to the control keywords. The environmental information includes humidity and light intensity.
3. The image generation method as described in claim 2, characterized in that, The preset image generation model includes a target image generation model and a physical simulation layer. The step of inputting the collected current environmental information and the prompt words into the preset image generation model to generate an image that has a higher degree of aging simulation and conforms to the control keywords includes: The target image generation model extracts features from the prompt words to obtain feature vectors. Based on the feature vectors, an initial antique image is generated. The feature vectors include keyword vectors, era vectors, and environment vectors. The physical simulation layer determines the material properties of the canvas material corresponding to the age vector, and the aging mask is determined based on the age vector, the environment vector, and the material properties. The aging mask is fused with the initial antique image to obtain an antique painting image with a higher degree of aging simulation than the initial simulation image and that meets the control keywords.
4. The image generation method as described in claim 1, characterized in that, Before the step of inputting the prompt words into a preset image generation model to generate an image that conforms to the control keywords, the following steps are included: Obtain a set of ancient images with a consistent style, and extract content keywords from the ancient image set; The design factors related to the control keywords in each image of the ancient image set are precisely labeled to obtain the ancient image set with the control keywords after precise labeling; Based on the control keywords and the content keywords, generate sample prompt words corresponding to each image; The sample prompts and the set of ancient images containing the control keywords are input into the first image generation model for iterative training to obtain the second image generation model after preliminary adjustment. The second image generation model is fine-tuned to obtain a target image generation model that satisfies local precise control.
5. The image generation method as described in claim 4, characterized in that, The step of fine-tuning the second image generation model to obtain the target image generation model that satisfies local precise control includes: In each image, a target mask of the corresponding type is generated based on the control keywords for the region corresponding to the design factor. The target mask includes an aging mask, a stamp mask, and a canvas mask. Determine the target loss function, wherein the target loss function is constructed based on the base loss function and the decoupling loss function; Based on the ancient image set with the target mask and the control keywords, the content keywords, and the target loss function, the second image generation model is fine-tuned to obtain the fine-tuned target image generation model.
6. The image generation method as described in claim 4, characterized in that, The step of fine-tuning the second image generation model to obtain the target image generation model that satisfies local precise control includes: A material attribute database is constructed based on the canvas material and the physical properties corresponding to the canvas material. Based on the material property database and the initial aging model, a physical simulation layer is constructed.
7. An image generation apparatus, characterized in that, The device includes: The acquisition module acquires prompt words input by the user, wherein the prompt words include at least one control keyword for preset design factors, and the control keyword includes the degree of aging, the number and position of stamps, and the color and texture of the canvas; The generation module is used to input the prompt words into a preset image generation model to generate an image and obtain a target antique image that conforms to the control keywords.
8. An image generation device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image generation method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the image generation method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the image generation method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN118314229A
Automobile modeling automatic generation method, device and equipment and storage medium
CN118332698A
Painting processing method and device, storage medium and electronic equipment
CN118918199A
Image generation method and device, equipment and medium
CN120198526A
Defect picture generation method and device, electronic equipment and storage medium
CN120198549A