Image generation method, electronic equipment, storage medium and computer program product

By decomposing image generation into a multi-stage process and receiving user feedback in real time, the problem of uncontrollable generation quality is solved, and an efficient image generation method is achieved.

CN121937574APending Publication Date: 2026-04-28SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN COOCAA NETWORK TECH CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing image generation methods based on large models suffer from uncontrollable generation quality, requiring users to make repeated adjustments and reducing image generation efficiency.

Method used

The system extracts painting information from the received natural language description, breaks it down into multiple painting stages, and receives user feedback in real time at each stage. The system adjusts the painting content through voice, remote control, and body language to generate the target image.

Benefits of technology

It significantly improves the controllability and efficiency of the generation process, reduces the waste of time and computing resources caused by regeneration, and ensures that the generated results meet user expectations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937574A_ABST
    Figure CN121937574A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method, electronic equipment, a storage medium and a computer program product, and relates to the technical field of image processing, and the method comprises the steps: extracting drawing information from a received natural language description; based on the drawing information, drawing contents corresponding to the drawing stages are generated in sequence; and in the process of generating the drawing content, receiving user feedback information in real time, and adjusting the drawing content based on the user feedback information to obtain a target image. The technical problem that the generation quality is uncontrollable in the current image generation method based on a large model, so that a user needs to adjust repeatedly, and the image generation efficiency is reduced is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image generation method, electronic device, storage medium and computer program product. Background Technology

[0002] With the rapid development of artificial intelligence technology, AI painting applications based on large models are becoming increasingly popular. In existing image generation technologies, mainstream solutions typically input the user's complete text description into the model in one go, directly generating the final image result. During this process, if the user is dissatisfied with the generated result, they can usually only try to obtain a new image by modifying the text prompts and triggering the model to perform a complete re-inference, resulting in a long adjustment cycle and low efficiency. Therefore, current image generation methods based on large models suffer from uncontrollable generation quality, requiring users to make repeated adjustments and reducing image generation efficiency.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main purpose of this application is to provide an image generation method, electronic device, storage medium, and computer program product, which aims to solve the technical problem in current image generation methods based on large models where the generation quality is uncontrollable, requiring users to make repeated adjustments and reducing image generation efficiency.

[0005] To achieve the above objectives, this application proposes an image generation method, which includes: Extract painting information from the received natural language description; Based on the painting information, the painting content corresponding to each painting stage is generated sequentially. During the generation of the painting content, user feedback information is received in real time, and the painting content is adjusted based on the user feedback information to obtain the target image.

[0006] In one embodiment, each drawing stage includes a sketch generation stage, a detail addition stage, and a coloring optimization stage; The step of sequentially generating the painting content corresponding to each painting stage based on the painting information includes: During the sketch generation stage, a first drawing including the outlines of core elements is generated based on the drawing information; During the detail addition stage, a second drawing including detail elements is generated based on the painting information and the first drawing; During the coloring optimization stage, a third drawing, including color elements, is generated based on the painting information and the second drawing.

[0007] In one embodiment, the user feedback information includes at least one of the user's voice information, remote control information, and body language information; The step of receiving user feedback information in real time and adjusting the drawing content based on the user feedback information includes at least one of the following: The voice information is captured, a first adjustment intention is extracted from the voice information, and the drawing content is adjusted based on the first adjustment intention. Receive the remote control information, extract the second adjustment intention from the remote control information, and adjust the painting content based on the second adjustment intention; The system receives the limb information, extracts a third adjustment intention from the limb information, and adjusts the drawing content based on the third adjustment intention.

[0008] In one embodiment, the first adjustment intention includes the user's adjustment object and adjustment content of the drawing content; The step of extracting a first adjustment intent from the voice information and adjusting the drawing content based on the first adjustment intent includes: Identify the voice commands in the voice information, and determine the adjustment object and the adjustment content based on the voice commands; Based on the adjustment object, a corresponding first adjustment area is located in the painting content, and based on the adjustment content, a first adjustment operation is determined for the first adjustment area. The first adjustment operation is then performed on the first adjustment area to adjust the painting content.

[0009] In one embodiment, the second adjustment intent includes the user's adjustment options for the drawing content and the adjustment area; The step of extracting the second adjustment intention from the remote control information and adjusting the painting content based on the second adjustment intention includes: The button signals and cursor signals are extracted from the remote control information; the adjustment options are determined based on the button signals; and the adjustment area is determined based on the cursor signals. Based on the adjustment area, a corresponding second adjustment area is located in the painting content, and based on the adjustment options, a second adjustment operation is determined for the second adjustment area. The second adjustment operation is then performed on the second adjustment area to adjust the painting content.

[0010] In one embodiment, the third adjustment intent includes the user's implicit evaluation of the painting content and the adjustment target; The step of extracting a third adjustment intention from the limb information and adjusting the drawing content based on the third adjustment intention includes: Extract facial expression information and body posture information from the limb information; The implicit evaluation is assessed based on the facial expression information and the body posture information; Based on the facial expression information, the user's gaze focus is determined, as well as the user's image visual features at the moment the expression is triggered. The user's gaze focus and the image visual features are then analyzed to obtain the adjustment target. When the implicit evaluation is negative feedback, a third adjustment area and a third adjustment operation are determined in the painting content based on the adjustment target, and the third adjustment operation is performed on the third adjustment area to adjust the painting content.

[0011] In one embodiment, the step of determining the third adjustment area and the third adjustment operation in the painting content based on the adjustment target includes: Based on the adjustment target, an initial adjustment area is determined in the painting content, wherein the initial adjustment area is the area corresponding to the user's visual focus and the image visual features in the painting content; The initial adjustment area is matched with the user's historical style preferences to determine the core area in the initial adjustment area that does not match the user's historical style preferences, thus obtaining the third adjustment area. The user's historical style preferences include visual style features, which are features recorded from the painting content when the implicit evaluation is positive feedback. The corresponding third adjustment operation is determined based on the third adjustment area and the user's historical style preferences.

[0012] Furthermore, to achieve the above objectives, this application also proposes an image generation system, which includes: The information extraction module is used to extract painting information from the received natural language description; The content generation module is used to sequentially generate painting content corresponding to each painting stage based on the painting information. The image adjustment module is used to receive user feedback information in real time during the generation of the painting content, and adjust the painting content based on the user feedback information to obtain the target image.

[0013] In addition, to achieve the above objectives, this application also proposes an electronic device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image generation method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the image generation method described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the image generation method described above.

[0016] This application provides an image generation method, which includes: extracting painting information from a received natural language description; generating painting content corresponding to each painting stage based on the painting information; receiving user feedback information in real time during the generation of painting content, and adjusting the painting content based on the user feedback information to obtain a target image.

[0017] This application extracts painting information from the received natural language description, thereby transforming user intent into structured generation conditions. This solves the technical problem of model misunderstanding caused by vague or complex descriptions, ensuring the initial accuracy of the generation process. Based on this painting information, the application sequentially generates painting content corresponding to each painting stage, decomposing the one-time generation process into multiple controllable progressive stages. This addresses the technical problems of unpredictable results and the need for users to passively accept all details in existing technologies, which require one-time generation. Users can intervene in intermediate stages, achieving initial control over the generation direction. During the generation of these staged contents, user feedback is received in real time, and the painting content is dynamically adjusted accordingly. Real-time intervention and correction are allowed during generation, significantly reducing the time and computational resource waste caused by completely regenerating the image, ultimately resulting in a target image that meets the user's expectations. Compared to related solutions that input the user's complete text description into the model at once and directly generate the final image result, this solution transforms the traditional single black-box generation into a multi-stage, interactive, guided generation, thereby significantly improving the controllability and overall efficiency of the generation process while ensuring generation quality. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating an embodiment of the image generation method of this application. Figure 2 This is a flowchart illustrating Embodiment 2 of the image generation method of this application; Figure 3 A simplified flowchart illustrating the image generation method provided in Embodiment 2 of this application; Figure 4 This is a schematic diagram of the module structure of the image generation system according to an embodiment of this application; Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the image generation method in the embodiments of this application.

[0021] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the first embodiment described herein is merely used to explain the technical solution of this application and is not intended to limit this application.

[0023] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0024] The main solution of the first embodiment of this application is: extracting painting information from the received natural language description; generating painting content corresponding to each painting stage in sequence based on the painting information; receiving user feedback information in real time during the process of generating painting content, and adjusting the painting content based on the user feedback information to obtain the target image.

[0025] In the first embodiment, for ease of description, the following description will focus on the image generation system as the execution subject.

[0026] With the rapid development of artificial intelligence technology, AI painting applications based on large models are becoming increasingly popular. In existing image generation technologies, mainstream solutions typically input the user's complete text description into the model at once and directly generate the final image result. In this process, if the user is not satisfied with the generated result, they can usually only try to obtain a new image by modifying the text prompts and triggering the model to perform a complete inference again, which is a long adjustment cycle and inefficient.

[0027] This application provides a solution that extracts painting information from received natural language descriptions, thereby transforming user intent into structured generation conditions. This solves the technical problem of model comprehension bias caused by vague or complex descriptions, ensuring the initial accuracy of the generation process. Based on this painting information, painting content corresponding to each painting stage is generated sequentially, decomposing the one-time generation process into multiple controllable progressive stages. This solves the technical problems of unpredictable results and the need for users to passively accept all details in existing technologies, which require one-time generation. Users can intervene in intermediate stages, achieving initial control over the generation direction. During the generation of these staged contents, user feedback is received in real time, and the painting content is dynamically adjusted accordingly. Real-time intervention and correction are allowed during generation, significantly reducing the time and computational resource waste caused by completely regenerating the image, ultimately obtaining the target image that meets the user's expectations. Compared to related solutions that input the user's complete text description into the model all at once and directly generate the final image result, this solution transforms the traditional single black-box generation into a multi-stage, interactive, guided generation, thereby significantly improving the controllability and overall efficiency of the generation process while ensuring generation quality.

[0028] It should be noted that the executing entity in the first embodiment can be a computing service device with data processing, network communication, and program execution functions, such as an electronic device like a television, tablet computer, personal computer, or mobile phone, or a system, application, or program capable of implementing the above functions. The first embodiment and the following embodiments will be described using an image generation system as an example.

[0029] All actions involving the acquisition of signals, information, or data in this application are carried out in accordance with the relevant data protection laws and policies of the country where the application is located, and with the authorization of the owner of the relevant device.

[0030] Based on this, embodiments of this application provide an image generation method, referring to... Figure 1 , Figure 1 This is a schematic flowchart of the first embodiment of the image generation method of this application.

[0031] In this embodiment, the image generation method includes steps S01 to S03: Step S01: Extract painting information from the received natural language description; It should be noted that natural language descriptions are statements entered by the user to express their painting intentions. Painting information refers to semantic elements used to guide the generation of the target image, which may include core elements, attributes, and scene relationships.

[0032] For example, the system receives a natural language description input by the user (e.g., "Draw me a picture of Princess Elsa's party"), calls a natural language processing model to parse it, first performs word segmentation and part-of-speech tagging to identify key entities (such as "Princess Elsa" and "party"); then it uses dependency parsing to understand the modification relationships between key entities ("party" modifies "picture"); finally, it structures the key entities and modification relationships to extract painting information including dimensions such as core subject, scene, and style.

[0033] Understandably, step S01 transforms free text into explicit and operable painting elements by performing structured information extraction, laying a precise foundation for subsequent generation, significantly improving the alignment between the generation starting point and the user's original intent, reducing completely invalid generation caused by misunderstanding of intent from the source, and creating a premise for improving the efficiency of the entire process.

[0034] Step S02: Generate the painting content corresponding to each painting stage in sequence based on the painting information; It should be noted that the painting stage is a sequence of sub-processes with different generation goals, executed sequentially. It deconstructs the overall image generation task into multiple progressive and logically coherent stages, such as the sketch generation stage, the detail addition stage, and the color optimization stage. The painting content refers to the intermediate or final visual output generated in each painting stage based on the generation goal and input conditions of the current stage (such as the output or painting information of the previous painting stage).

[0035] Understandably, step S02 decomposes a single generation task into multiple ordered drawing stages, thereby enabling a progressive unfolding of the generation process. This gives users the opportunity to review and intervene at key intermediate nodes, allowing potential directional deviations to be detected and corrected in the early stages. It avoids consuming all computing resources on generation paths that may ultimately be completely rejected, thus improving the controllability and resource utilization efficiency of the generation process.

[0036] Step S03: During the process of generating the painting content, user feedback information is received in real time, and the painting content is adjusted based on the user feedback information to obtain the target image.

[0037] It should be noted that user feedback refers to input data captured through the interactive interface during or after any stage of the drawing process, reflecting the user's evaluation or intention regarding the current drawing content. Examples include voice commands (such as "make the pine trees denser"), remote control operations (such as selecting the sky area and choosing the "adjust tone" option), or user facial expressions and body language captured by the camera. The target image is the final, completed digital image output after all drawing stages, incorporating the user's feedback at each stage.

[0038] Understandably, step S03 receives user feedback in real time during the generation of the drawing content and adjusts the drawing content based on the user feedback. This addresses the core shortcomings of existing technologies, which rely on a single adjustment method and require re-entering prompts and regenerating the entire process, resulting in long cycles and sluggish feedback loops. By allowing users to inject feedback through natural interaction during the generation process, the traditional trial-and-error regeneration is optimized into guided iterative optimization. This allows for targeted correction of the generated content in the current or subsequent stages based on feedback, greatly reducing the number of complete generation iterations required to obtain satisfactory results, thereby improving the overall efficiency of image generation and user satisfaction.

[0039] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 2 In step S02, each painting stage includes a sketch generation stage, a detail addition stage, and a coloring optimization stage. The steps for generating the painting content corresponding to each painting stage based on the painting information include steps S11 to S13: Step S11: In the sketch generation stage, a first drawing including the outlines of core elements is generated based on the drawing information. It should be noted that the sketch generation stage is the initial structured output stage in the entire target image generation process. Based on the parsed painting information, this stage infers and outputs the most basic spatial structure and main form of the image through an image generation model. The core element outline refers to the boundary representation generated in the sketch generation stage after shape inference of the key entities (such as figures, main objects, and scene landmarks) described by the painting information. This representation is expressed as lines, blocks, or low-resolution semantic segmentation. It does not contain texture, internal details, or true color; it only defines the position, basic proportions, and geometric shape of each main component. The first drawing is the direct output of the sketch generation stage; it is a low-detail image data primarily composed of the core element outlines.

[0040] Additionally, it should be noted that the image generation model is an artificial intelligence model trained on an image-text pair dataset. Its core function is to generate digital images that meet semantic and visual requirements based on given input information (such as text descriptions, sketches, semantic labels, or other images).

[0041] Step S12, in the detail addition stage, a second drawing including detail elements is generated based on the painting information and the first drawing; It's important to note that the detailing stage follows the sketch generation stage and is a processing phase designed to enrich the visual content. This stage uses the first drawing as a foundation and references painting information, employing an image generation model to fill the first drawing with microstructures that conform to the semantics of the painting information. Detail elements refer to the visual components generated within and around the outline of core elements during the detailing stage, used to represent object textures, decorative parts, or environmental embellishments; for example, the clusters of leaves and bark texture within the outline of a tree. The second drawing is the output after the detailing stage is completed. It approaches the visual richness of the final image, but is usually still monochrome or only has a basic color tone, without full color rendering.

[0042] Step S13: In the color optimization stage, a third drawing including color elements is generated based on the painting information and the second drawing.

[0043] It should be noted that the color optimization stage is the final processing stage in the target image generation process. It is primarily responsible for applying and coordinating colors and rendering the overall image. This stage uses the second drawing as a base and integrates the color descriptions and style instructions from the drawing information. Color elements refer to the color attributes assigned to different areas and details in the second drawing during the color optimization stage, including but not limited to inherent color, light and shadow color, and ambient reflection color. For example, giving leaves green and adding highlights of yellow under sunlight. The third drawing is the output after the color optimization stage is completed, and it is also the final target image.

[0044] For example, in the sketch generation stage, when the drawing information includes "windmill grassland at sunset", the core elements such as "windmill", "grassland" and "sunset" are identified, and a line drawing style image is generated through the image generation model. This image only includes the approximate shape, relative position and spatial relationship of these elements, without involving any detailed texture or color. In the detail addition stage, the structural texture of the blades and the material of the tower are added to the outline of the windmill. The layers and undulations of the grass blades are depicted within the outline of the grassland, and the shape of the clouds is initially outlined. In the color optimization stage, the golden yellow gradient of the sunset is rendered for the sky, and the lighting and shadows are applied to the windmill and grassland to determine the overall color harmony and atmosphere.

[0045] In this embodiment, by generating the first drawing of the core element outline, the technical problem of the initial composition potentially deviating fundamentally from the user's intention is solved, allowing the user to confirm or correct the overall layout in the early stages. Based on the confirmed outline, detailed elements are added to generate the second drawing, which solves the technical problem that the detailed depiction may deviate from the user's expectations or be contrary to common sense. This solidifies the basic content of the image before coloring, reducing rework caused by improper details. Based on the determined details, coloring optimization is performed to generate the third drawing, which solves the technical problem that the final effect of color matching or style rendering is unsatisfactory, greatly improving the controllability of the final output and the success rate of one-time generation.

[0046] In one feasible implementation, in step S03, the user feedback information includes at least one of the user's voice information, remote control information, and body language information. The step of receiving user feedback information in real time and adjusting the drawing content based on the user feedback information includes at least one of steps A01 to A03: Step A01: Capture voice information, extract the first adjustment intention from the voice information, and adjust the drawing content based on the first adjustment intention; It should be noted that voice information refers to continuous acoustic signals emitted by the user, which contain the user's intention to modify the currently generated painting content in natural language form, such as "darken the color of the trees on the left." The first adjustment intention is a structured, executable modification instruction extracted after recognizing and semantically understanding the voice information. This intention clarifies the object of adjustment (such as "trees on the left") and the specific adjustment operation (such as "darken the color"), thereby transforming the vague user speech into precise image processing parameters.

[0047] Step A02: Receive remote control information, extract the second adjustment intention from the remote control information, and adjust the drawing content based on the second adjustment intention; It should be noted that remote control information refers to control signals generated by the user through a physical remote control device (such as a TV remote control), including button trigger signals and cursor movement / positioning signals. For example, the user presses the "brightness+" button or uses the directional keys to select an area on the screen. The second adjustment intent is the direct operation command determined after decoding and contextualizing the remote control information. This intent usually directly corresponds to predefined adjustment options (such as "increase overall brightness") and / or spatial operations (such as "blur the selected area"), achieving fast and precise directional control.

[0048] Step A03: Receive body information, extract the third adjustment intention from the body information, and adjust the drawing content based on the third adjustment intention.

[0049] It should be noted that body language information refers to nonverbal physiological signals captured by image acquisition devices (such as cameras) that reflect the user's subconscious reaction to the current image. This includes facial expressions (such as frowning, smiling, etc.) and body posture information (such as crossed arms, "OK" gesture, etc.). For example, the system detects a user staring at a certain part of the image for a long time while shaking their head. The third adjustment intention is the implicit evaluation and adjustment target of the user inferred by the system after analyzing and recognizing the body language information. For example, "frowning + focusing the gaze on the sky area of ​​the image" is interpreted as "possibly dissatisfied with the sky area," thereby triggering adjustments to that area.

[0050] Additionally, it should be noted that after adjusting the drawing content, new drawing content will be obtained. For the new drawing content, user interaction information can be captured again. If the user's adjustment request is received, a new round of adjustments will be started until no further adjustment requests are received (for example, capturing a signal from the user that the adjustment is complete and can proceed to the next stage).

[0051] Additionally, it should be noted that the three adjustment methods in steps A01 to A03 can be used individually or in combination. If a certain adjustment method cannot determine the content to be adjusted, it can be combined with other adjustment methods for joint identification. For example, adjustments based on both voice information and body language information can be combined simultaneously.

[0052] In this implementation, by parsing voice commands to extract adjustment intentions, the problem of unintuitive and inefficient expression of adjustment intentions in traditional text input is solved, improving the naturalness of interaction and the efficiency of intention transmission. By parsing remote control signals to extract operation intentions, the problem of cumbersome operation of complex interfaces and difficulty in accurately controlling specific areas is solved. By parsing body language to extract emotional and attentional intentions, the problem of difficulty in capturing implicit user feedback and lack of immediacy in adjustments is solved, realizing intelligent perception and pre-adjustment without explicit commands. These multimodal feedback mechanisms together realize an efficient and accurate human-computer interaction closed loop, significantly reducing the waiting time caused by repeated trial adjustments, and improving the efficiency of the overall generation process and user satisfaction.

[0053] In one feasible implementation, in step A01, the first adjustment intention includes the user's desired adjustment object and the content to be adjusted in the drawing. The steps of extracting the first adjustment intention from the voice information and adjusting the drawing content based on the first adjustment intention include steps A11-A12: Step A11: Identify the voice commands in the voice information, and determine the adjustment object and adjustment content based on the voice commands; It should be noted that voice commands are extracted from speech information through speech recognition, representing the user's intended modifications to the current drawing content. The "object to be adjusted" refers to the element or component the user intends to modify, parsed from the voice command using semantic understanding technology. This can be an entity or visual feature within the drawing content, such as "the trees on the left." The "content to be adjusted" refers to the specific modification requests the user wishes to make to the object to be adjusted, identified from the voice command using semantic understanding technology. This defines the nature and direction of the modification, such as "darken the colors."

[0054] Step A12: Locate the corresponding first adjustment area in the drawing content based on the adjustment object, determine the first adjustment operation for the first adjustment area based on the adjustment content, and perform the first adjustment operation on the first adjustment area to adjust the drawing content.

[0055] It should be noted that the first adjustment region is one or more pixel areas located and mapped within the painting content through object recognition and spatial position analysis. This region is the specific scope of the first adjustment operation. The first adjustment operation converts the user's natural language description into operation instructions that can be executed by the image processing engine, based on the content to be adjusted. The image processing engine is a pre-trained image adjustment model that can perform local adjustments to the first adjustment region according to the first adjustment operation. Because it is a local adjustment, there is no need to regenerate the entire painting content, and the user can more intuitively identify the adjusted part, thereby improving the efficiency of image adjustment.

[0056] Additionally, it's worth noting that when using an image processing engine, adjustments to local areas of the image are involved. To ensure the adjusted image doesn't appear abrupt or harsh, boundary-aware feathering and blending techniques can be employed. This involves progressively blending the edges of the adjusted area at the pixel level to prevent sharp seams. For example, if a tree is adjusted to an autumnal, withered yellow, the engine will simultaneously adjust the color of its cast shadow and may fine-tune the surrounding affected leaves or ground reflections to maintain the rationality and consistency of the global lighting relationship.

[0057] For example, a speech command is extracted from the speech information by a speech recognition engine and converted into text. The text is then semantically parsed using a natural language processing model to identify the adjustment object (e.g., "the house in the lower left corner of the picture") and the adjustment content (e.g., "turn it red") in the text. Through image recognition and semantic matching technology, a first adjustment area that matches the description of the adjustment object is located in the painting content, such as the image area where "the house in the lower left corner of the picture" is located. At the same time, the adjustment content is parsed into a specific first adjustment operation, such as replacing the color of the first adjustment area in the painting content with red. The image processing engine is then called to perform the first adjustment operation (i.e., the color replacement operation) on the first adjustment area, thereby completing the adjustment of the painting content.

[0058] In this embodiment, speech recognition and semantic understanding technologies solve the technical problem of ambiguous natural language instructions that cannot be directly converted into executable operations, achieving precise structuring of user intent. By mapping the structured intent to specific image regions and operation parameters, the technical problem of a gap between adjustment instructions and image modification actions is solved, realizing an automated closed loop from voice instructions to pixel-level image modification. This avoids multiple rounds of repeated modifications caused by unclear intent or inconvenient operation, thereby significantly improving the adjustment efficiency and controllability of the overall generation process.

[0059] In one feasible implementation, in step S01, the second adjustment intention includes the user's adjustment options for the drawing content and the adjustment area. The steps of extracting the second adjustment intention from the remote control information and adjusting the drawing content based on the second adjustment intention include steps A21-A22: Step A21: Extract button signals and cursor signals from the remote control information, determine the adjustment options based on the button signals, and determine the adjustment area based on the cursor signals; It should be noted that button signals refer to discrete electronic signals generated when a user operates the physical buttons on the remote control or the touch panel, used to select or activate predefined adjustment options. Cursor signals refer to the continuous coordinate data stream generated when a user controls the movement of the screen pointer using the buttons on the remote control or the touch panel, used to track the pointer's position and movement trajectory on the displayed screen in real time, allowing the user to intuitively define adjustment areas; for example, the user can move the cursor to select a region on the screen. Adjustment options are preset image processing or content generation operation categories, triggered by button signals, such as "saturation enhancement," "partial redraw," and "filter addition." Adjustment areas refer to one or more spatial ranges directly specified by the user on the displayed drawing content screen using cursor signals.

[0060] Step A22: Locate the corresponding second adjustment area in the painting content based on the adjustment area, determine the second adjustment operation for the second adjustment area based on the adjustment options, and perform the second adjustment operation on the second adjustment area to adjust the painting content.

[0061] It should be noted that the second adjustment region is the actual image area within the internal pixel matrix of the painting content that matches the user's selected area, determined through coordinate mapping and image semantic segmentation techniques. For example, the roughly selected rectangular area is precisely adapted to the pixel set corresponding to the target area in the painting content using image recognition. The second adjustment operation is executed in conjunction with the specific image processing instructions defined by the user's selected adjustment option. For example, if the adjustment option is "increase saturation," the second adjustment operation applies a saturation enhancement algorithm to the pixels within the second adjustment region.

[0062] For example, remote control information from a remote control device is parsed, and the key signals and cursor signals are separated. Based on a preset key-function mapping table, the key press event corresponding to the key signal (e.g., pressing the number "1" key) is parsed into a specific adjustment option (e.g., "increase saturation"). Simultaneously, the cursor signal (e.g., coordinate data generated by cursor pointer movement) is tracked and decoded, and its movement trajectory on the display interface is converted into a selection or pointing to a specific graphic area on the screen, thereby determining the area the user intends to adjust. The adjustment area is mapped to the corresponding pixel coordinate space of the current drawing content, locating the second adjustment area to be modified (e.g., the sky area circled by the user with the cursor). Then, based on the adjustment option, a corresponding second adjustment operation is matched or dynamically generated from a preset image processing instruction library (e.g., applying a pixel-level algorithm to increase saturation to the second adjustment area). During adjustment, the second adjustment area is isolated, and the second adjustment operation is performed only within this local area, achieving targeted modification without regenerating the entire drawing content, thus achieving precise point-to-point modification of the drawing content.

[0063] In this embodiment, by parsing the button signals and cursor signals of the remote control, the user's physical operations are directly mapped to clear adjustment options and precise adjustment areas. This solves the technical problem that traditional TV-end interaction cannot perform intuitive and efficient local adjustments to the large-screen image, making the intention transmission more direct. Based on the precisely defined second adjustment area and the selected adjustment option, the corresponding image processing algorithm is invoked to perform the second adjustment operation on that area. This solves the core technical problem in existing solutions where any local modification requires triggering a global regeneration of the large model, resulting in wasted computing resources and long waiting times. It enables point-to-point, rapid, and controllable modification of the drawing content, avoiding the efficiency loss caused by global regeneration, and significantly improving the accuracy of the adjustment and the efficiency of the overall generation process.

[0064] In one feasible implementation, in step S01, the third adjustment intention includes the user's implicit evaluation of the painting content and the adjustment goal. The steps of extracting the third adjustment intention from the body information and adjusting the painting content based on the third adjustment intention include steps A31 to A34: Step A31: Extract facial expression information and body posture information from the limb information; It should be noted that facial expression information refers to the morphological pattern data formed by the movement of the user's facial muscles in the user's facial image captured by the image acquisition device, such as a combination of features like upturned corners of the mouth, furrowed eyebrows, or widened eyes. Body posture information is the spatial position, orientation, or movement trajectory data of the user's body parts captured by the image acquisition device, such as nodding, shaking the head, pointing an arm, or leaning forward / backward.

[0065] Step A32: Evaluate the implicit evaluation based on facial expression information and body posture information; It should be noted that implicit evaluation is the user's attitude towards the current drawing content inferred by analyzing facial expression information and body posture information, including positive feedback (such as satisfaction, interest) and negative feedback (such as dissatisfaction, confusion).

[0066] Additionally, it's worth noting that implicit evaluations can be assessed using a pre-trained multimodal fusion evaluation model. This model establishes a mapping between facial expressions, body postures, and pre-defined sentiment-intent labels (such as "highly approving," "generally satisfied," "neutral," "slightly dissatisfied," and "strongly disapproving"), enabling it to infer implicit evaluations from a user's real-time state. By inputting facial expression and body posture information into this multimodal fusion evaluation model, the corresponding implicit evaluation can be output. For example, when the model detects a user frequently frowning and shaking their head, it might classify it as "strongly disapproving"; while when it detects a smile and nodding, it might classify it as "highly approving." Ultimately, a structured implicit evaluation result is output, which includes not only qualitative sentiment classification but also a corresponding confidence score.

[0067] Step A33: Determine the user's gaze focus based on facial expression information, as well as the user's visual features at the moment the expression is triggered, and analyze the user's gaze focus and visual features to obtain the adjustment target; It should be noted that the user's gaze focus is estimated based on the analysis of the user's eye orientation and head posture, representing the area of ​​visual attention concentrated on the display screen, such as the upper and lower halves of the drawing content. The expression trigger moment is the specific point in time when the user's facial expression undergoes a significant change (e.g., from calm to frowning). Image visual features refer to the regional attribute information output or adjusted in the drawing content at the expression trigger moment, such as added details, colors, etc. The adjustment goal is a directional description of the user's potential modification expectations, synthesized from the user's gaze focus and corresponding image visual features, combined with inferences about implicit evaluations. For example, when the user responds to the color filling the sky area in the drawing content, the analyzed adjustment goal is "to darken the color of that area."

[0068] Additionally, it should be noted that when determining the user's gaze focus, the spatial orientation of the user's head can be calculated using a head pose estimation algorithm, and the user's pupils can be located using image recognition to calculate the gaze direction of the user's eyes. The head orientation and gaze direction are then fused together and converted into a gaze area in the display screen coordinate system. This gaze area is the user's gaze focus, such as the left / right area of ​​the upper / lower half of the drawing content.

[0069] Additionally, it's important to note that since the drawing content isn't generated entirely at once, the drawing itself isn't generated in a single, continuous process for each stage. For example, in the sketch generation stage, the output might be printed from top to bottom. Therefore, user feedback (a change in facial expression) might be received when the upper half of the drawing content is output. In the detail addition stage, since adjustments are made based on the first drawing obtained in the sketch generation stage, user feedback might be received after any local adjustment. For instance, adding a "star" element to the first drawing might elicit a frown from the user indicating dissatisfaction. Similarly, in the color optimization stage, since adjustments are made based on the second drawing obtained in the detail addition stage, user feedback might be received after any area's color is adjusted. For example, filling the sky with blue in the second drawing might elicit a frown from the user indicating dissatisfaction. Thus, because the drawing content isn't generated all at once, users will react immediately (i.e., their facial expressions change) when they observe content in the drawing that doesn't meet their expectations during image generation. The parts of the drawing content generated / adjusted at this point become the key areas for adjustment, requiring the extraction of their visual features to aid in subsequent image adjustments.

[0070] Additionally, it should be noted that by combining the user's visual focus (the part the user mainly observes) with the visual features of the image (the area where the user reacts or may be dissatisfied), it is possible to accurately identify the area the user wants to adjust, i.e., the adjustment target.

[0071] Step A34: In the case of implicit negative feedback, determine the third adjustment area and the third adjustment operation in the painting content based on the adjustment goal, and perform the third adjustment operation on the third adjustment area to adjust the painting content.

[0072] It should be noted that negative feedback refers to an implicit evaluation of dissatisfaction, rejection, or disapproval. The third adjustment area refers to the specific image region in the drawing content that the system identifies as needing modification when the implicit evaluation is negative feedback, based on the adjustment target. The third adjustment operation refers to the specific image processing or generation instructions planned and executed for the third adjustment area to correct the negative feedback.

[0073] In one feasible implementation, step A34, which involves determining the third adjustment area and the third adjustment operation in the painting content based on the adjustment target, includes steps A41-A43: Step A41: Determine the initial adjustment area in the painting content based on the adjustment target, wherein the initial adjustment area is the area corresponding to the user's visual focus and the visual features of the image in the painting content; It should be noted that the initial adjustment area is one or more potential areas to be modified based on the user's gaze focus and the visual features of the image mapped onto the painting content.

[0074] Step A42: Match the initial adjustment area with the user's historical style preferences, determine the core area in the initial adjustment area that does not match the user's historical style preferences, and obtain the third adjustment area. The user's historical style preferences include visual style features, which are visual area features recorded from the painting content when the implicit evaluation is positive feedback. It should be noted that user historical style preferences refer to a personalized visual generation tendency database built for users through continuous learning. This includes visual style features, which characterize the statistical patterns in color, composition, brushstrokes, etc., of images that users have explicitly preferred or ultimately adopted in the past. Visual style features are specific image attribute parameters that constitute user historical style preferences. For example, in terms of color, this may be reflected in a preference for high saturation and warm tones; in terms of composition, it may be reflected in a preference for centrally centered subjects. These features are analyzed, extracted, and recorded from the approved painting content when the user's implicit evaluation shows positive feedback (such as smiling, nodding, or other explicit positive body signals). Positive feedback is a positive emotional state or approval intention determined by analyzing the user's facial expression and body posture information, and its manifestations include, but are not limited to, smiling, relaxed posture, and nodding. The core area is one or more sub-regions identified within the initial adjustment area by comparing the current visual performance (such as color, brightness, and structure) with the corresponding visual style features in the user's historical style preferences item by item, and identifying differences exceeding a preset error tolerance threshold. For example, if the initial adjustment area is the sky area in the painting content, and comparison reveals that its blue saturation is significantly lower than the user's historical preference for a vibrant style, then this area is determined to be the core area, i.e., the third adjustment area.

[0075] Step A43: Match the corresponding third adjustment operation in the preset adjustment operation library according to the third adjustment area.

[0076] For example, based on the adjustment target parsed from body information (e.g., the user's gaze is focused on the sunset area in the image for a long time, accompanied by negative micro-expressions), the mapping range of the gaze focus in the image and its associated visual features (such as the color and texture of the area) are defined as the initial adjustment area. Introducing the user's historical style preferences as a decision-making basis, the visual features of the initial adjustment area (e.g., the current orange-red hue of the sunset area) are matched against a pre-built historical style preference database for the user (this database records the strong purplish-red hues and high contrast tended in the sunsets of several paintings the user previously expressed satisfaction with). By calculating feature similarity, the core parts of the initial adjustment area that significantly differ from historical preferences (e.g., color inconsistencies) are identified and designated as the third adjustment area. Based on the identified third adjustment area and the user's historical style preferences, a specific third adjustment operation is generated or matched (e.g., applying a color transformation filter that maps orange-red to purplish-red and enhances contrast).

[0077] In this implementation, the initial adjustment area is located based on the adjustment target, solving the technical problem of ambiguous adjustment range definition when relying solely on explicit instructions. This provides a clear spatial range for subsequent operations. By matching and analyzing the visual features of the initial adjustment area with the user's historical style preferences, the non-matching core area is automatically identified as the third adjustment area. This solves the technical problem that the adjustment direction may deviate from the user's long-term aesthetic preferences, leading to repeated corrections. It achieves personalized prediction of the adjustment. Based on the specific deviation of the non-matching area and the user's preferences, an appropriate third adjustment operation is generated, solving the problem of the traditional method requiring the user to manually specify specific parameters, resulting in a cumbersome and inefficient adjustment process. By introducing an intelligent matching and decision-making mechanism based on historical preferences, the reliance on the user's immediate and precise instructions is significantly reduced, achieving more accurate and efficient automated adjustment, effectively improving the certainty of the generation process and the satisfaction of the final result.

[0078] For example, to aid in understanding the technical concept or principles of this application, please refer to Figure 3 , Figure 3 A simplified flowchart of the image generation method is provided. First, the system extracts drawing information from user input. Then, the process sequentially enters three orderly drawing generation stages: the first drawing is generated in the sketch generation stage, the second drawing is generated based on the previous results in the detail addition stage, and the third drawing is generated in the color optimization stage. During this process, the system adjusts the drawing content of the current stage according to the user feedback information received in real time. Finally, the target image that meets the user's expectations is output at the end of the process.

[0079] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the image generation method of this application. Any simple transformations based on this technical concept are all within the protection scope of this application.

[0080] This application also provides an image generation system; please refer to [link / reference]. Figure 4 The image generation system includes: The information extraction module 10 is used to extract painting information from the received natural language description; Content generation module 20 is used to generate painting content corresponding to each painting stage based on painting information. The image adjustment module 30 is used to receive user feedback information in real time during the process of generating the painting content, and adjust the painting content based on the user feedback information to obtain the target image.

[0081] The image generation system provided in this application, employing the image generation method described in the above embodiments, can solve the technical problem in current large-model-based image generation methods where uncontrollable generation quality necessitates repeated adjustments by the user, thus reducing image generation efficiency. Compared with the prior art, the beneficial effects of the image generation system provided in this application are the same as those of the image generation method described in the above embodiments, and other technical features of the image generation system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0082] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the image generation method in Embodiment 1 above.

[0083] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, and PADs (Portable Application Description: Tablet computers), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0084] like Figure 5As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the electronic device to communicate wirelessly or wiredly with other devices to exchange data. Although the diagrams show electronic devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0085] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0086] The electronic device provided in this application, employing the image generation method described in the above embodiments, solves the technical problem in current image generation methods based on large models where uncontrollable generation quality necessitates repeated adjustments by the user, thus reducing image generation efficiency. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the image generation method described in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0087] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0088] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0089] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the image generation method described in the above embodiments.

[0090] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0091] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0092] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the image generating device to: extract painting information from a received natural language description; sequentially generate painting content corresponding to each painting stage based on the painting information; and, during the generation of painting content, receive user feedback information in real time and adjust the painting content based on the user feedback information to obtain the target image.

[0093] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0095] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0096] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described image generation method. This solves the technical problem in current large-model-based image generation methods where uncontrollable generation quality leads to repeated adjustments by the user, reducing image generation efficiency. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the image generation method provided in the above embodiments, and will not be repeated here.

[0097] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the image generation method described above.

[0098] The computer program product provided in this application can solve the technical problem in current image generation methods based on large models where the generation quality is uncontrollable, requiring users to make repeated adjustments and reducing image generation efficiency. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the image generation methods provided in the above embodiments, and will not be repeated here.

[0099] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. An image generation method, characterized in that, The image generation method includes: Extract painting information from the received natural language description; Based on the painting information, the painting content corresponding to each painting stage is generated sequentially. During the generation of the painting content, user feedback information is received in real time, and the painting content is adjusted based on the user feedback information to obtain the target image.

2. The image generation method as described in claim 1, characterized in that, Each painting stage includes the sketch generation stage, the detail addition stage, and the coloring optimization stage; The step of sequentially generating the painting content corresponding to each painting stage based on the painting information includes: During the sketch generation stage, a first drawing including the outlines of core elements is generated based on the drawing information; During the detail addition stage, a second drawing including detail elements is generated based on the painting information and the first drawing; During the coloring optimization stage, a third drawing, including color elements, is generated based on the painting information and the second drawing.

3. The image generation method as described in claim 2, characterized in that, The user feedback information includes at least one of the user's voice information, remote control information, and body language information; The step of receiving user feedback information in real time and adjusting the drawing content based on the user feedback information includes at least one of the following: The voice information is captured, a first adjustment intention is extracted from the voice information, and the drawing content is adjusted based on the first adjustment intention. Receive the remote control information, extract the second adjustment intention from the remote control information, and adjust the painting content based on the second adjustment intention; The system receives the limb information, extracts a third adjustment intention from the limb information, and adjusts the drawing content based on the third adjustment intention.

4. The image generation method as described in claim 3, characterized in that, The first adjustment intent includes the user's adjustment target and the content of the adjustment to the drawing content; The step of extracting a first adjustment intent from the voice information and adjusting the drawing content based on the first adjustment intent includes: Identify the voice commands in the voice information, and determine the adjustment object and the adjustment content based on the voice commands; Based on the adjustment object, a corresponding first adjustment area is located in the painting content, and based on the adjustment content, a first adjustment operation is determined for the first adjustment area. The first adjustment operation is then performed on the first adjustment area to adjust the painting content.

5. The image generation method as described in claim 3, characterized in that, The second adjustment intent includes the user's adjustment options for the drawing content and the adjustment area; The step of extracting the second adjustment intention from the remote control information and adjusting the painting content based on the second adjustment intention includes: The button signals and cursor signals are extracted from the remote control information; the adjustment options are determined based on the button signals; and the adjustment area is determined based on the cursor signals. Based on the adjustment area, a corresponding second adjustment area is located in the painting content, and based on the adjustment options, a second adjustment operation is determined for the second adjustment area. The second adjustment operation is then performed on the second adjustment area to adjust the painting content.

6. The image generation method as described in claim 3, characterized in that, The third adjustment intent includes the user's implicit evaluation of the painting content and the adjustment goal; The step of extracting a third adjustment intention from the limb information and adjusting the drawing content based on the third adjustment intention includes: Extract facial expression information and body posture information from the limb information; The implicit evaluation is assessed based on the facial expression information and the body posture information; Based on the facial expression information, the user's gaze focus is determined, as well as the user's image visual features at the moment the expression is triggered. The user's gaze focus and the image visual features are then analyzed to obtain the adjustment target. When the implicit evaluation is negative feedback, a third adjustment area and a third adjustment operation are determined in the painting content based on the adjustment target, and the third adjustment operation is performed on the third adjustment area to adjust the painting content.

7. The image generation method as described in claim 6, characterized in that, The steps of determining the third adjustment area and the third adjustment operation in the painting content based on the adjustment target include: Based on the adjustment target, an initial adjustment area is determined in the painting content, wherein the initial adjustment area is the area corresponding to the user's visual focus and the image visual features in the painting content; The initial adjustment area is matched with the user's historical style preferences to determine the core area in the initial adjustment area that does not match the user's historical style preferences, thus obtaining the third adjustment area. The user's historical style preferences include visual style features, which are features recorded from the painting content when the implicit evaluation is positive feedback. The corresponding third adjustment operation is determined based on the third adjustment area and the user's historical style preferences.

8. An electronic device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the image generation method as described in any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the image generation method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the image generation method as described in any one of claims 1 to 7.