Image processing method, device and equipment
By detecting and correcting image defects in real time during the image generation process, the problem that users cannot intervene in the prior art is solved, and higher quality and accurate image generation are achieved.
Patent Information
- Application Number
- CN202510395598.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
Existing image generation systems are difficult to ensure the accuracy of the generated results and meet user expectations, and users cannot intervene in the actual model generation process.
During the image generation process, image quality defects are detected simultaneously and correction prompts are displayed, allowing the user to enter correction instructions to correct the image.
Improve the controllability and accuracy of image generation, make the generated results closer to user's intentions, and improve image quality.
Smart Images

Figure CN120278976A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology. More specifically, it relates to an image processing method, device, and equipment. Background Art
[0002] Currently, with the development of artificial intelligence technology, generative artificial intelligence (Generative AI) models represented by diffusion models (Diffusion Models) and generative adversarial networks (GANs) have achieved high-precision text-to-image and image-to-image functions.
[0003] In the prior art, image generation systems generally adopt an end-to-end generation paradigm. Users cannot intervene in the actual model generation process, making it difficult to ensure the accuracy of the generated results and their compliance with user expectations. Summary of the Invention
[0004] Based on the above technical problems, this application proposes an image processing method, device, and equipment, which can effectively improve the quality and accuracy of image generation.
[0005] To achieve the above technical objectives, this application specifically proposes the following technical solutions:
[0006] In the first aspect of this application, an image processing method is provided, and the method further includes:
[0007] Obtain the creative input provided by the user;
[0008] During the process of generating an image in the image generation area based on the creative input, synchronously detect whether the generated image has quality defects;
[0009] When it is detected that the image has quality defects, display a correction prompt for the image with quality defects;
[0010] In response to receiving a correction instruction input by the user based on the correction prompt, correct the image displayed in the image generation area.
[0011] In the second aspect of this application, an image processing device is provided, and the device includes:
[0012] An acquisition unit for obtaining the creative input provided by the user;
[0013] A first processing unit for synchronously detecting whether the generated image has quality defects during the process of generating an image in the image generation area based on the creative input;
[0014] A display unit, configured to display a correction prompt for the image with quality defects in case that quality defects are detected in the image;
[0015] A correction unit, configured to correct the image displayed in the image generation area in response to receiving a correction instruction input by a user based on the correction prompt.
[0016] A third aspect of the present application provides a handwriting device, including:
[0017] A handwriting function component, and a processor connected to the handwriting function component;
[0018] The handwriting function component is configured to sense a creation input provided by a user;
[0019] The processor is configured to execute the image processing method as described in the first aspect of the present application.
[0020] A fourth aspect of the present application provides a control device, including a processor and an interface circuit, wherein the processor is connected to a handwriting function component through the interface circuit;
[0021] The handwriting function component is configured to sense a creation input provided by a user;
[0022] The processor is configured to execute the image processing method as described in the first aspect of the present application.
[0023] A fifth aspect of the present application provides an electronic device, including a memory and a processor;
[0024] The memory is connected to the processor and is configured to store a program;
[0025] The processor is configured to implement the image processing method as described in the first aspect of the present application by running the program in the memory.
[0026] A sixth aspect of the present application provides a chip, including a processor and a data interface, wherein the processor reads and runs a program stored on a memory through the data interface to execute the image processing method as described in the first aspect of the present application.
[0027] A seventh aspect of the present application provides a storage medium, on which a computer program is stored, and when the computer program is run by a processor, the image processing method as described in the first aspect of the present application is implemented.
[0028] The technical solution provided by this application obtains the creative input provided by the user. During the process of generating an image in the image generation area based on the creative input, it synchronously detects whether there are quality defects in the generated image, and when it detects that there are quality defects in the image, it displays a correction prompt for the image with quality defects; and in response to receiving a correction instruction input by the user based on the correction prompt, it corrects the image displayed in the image generation area. In this way, compared with the end-to-end image generation method, the embodiments of this application monitor the image quality in real time during the image generation process and correct the image during the generation process. At the same time, it allows the user to intervene during the image generation process and correct the deviation in the intermediate steps of image generation. Therefore, it can enhance the controllability of the image generation result, ensure that the image generation result is closer to the user's intention, and thus effectively improve the quality and accuracy of image generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only the embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on the provided accompanying drawings.
[0030] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of this application.
[0031] Figure 2 It is a schematic flowchart of another image processing method provided by an embodiment of this application.
[0032] Figure 3 It is a schematic flowchart of still another image processing method provided by an embodiment of this application.
[0033] Figure 4 It is a schematic flowchart of yet another image processing method provided by an embodiment of this application.
[0034] Figure 5 It is a schematic flowchart of yet another image processing method provided by an embodiment of this application.
[0035] Figure 6 It is a schematic flowchart of yet another image processing method provided by an embodiment of this application.
[0036] Figure 7a It is a schematic diagram of the AI painting interface of the image generation system provided by an embodiment of this application Figure 1 。
[0037] Figure 7b It is a schematic diagram of the AI painting interface of the image generation system provided by an embodiment of this application Figure 2 。
[0038] Figure 7c Schematic diagram of the AI painting interface of the image generation system provided by the embodiment of the present application Figure 3 。
[0039] Figure 7d Schematic diagram of the AI painting interface of the image generation system provided by the embodiment of the present application Figure 4 。
[0040] Figure 7e Schematic diagram of the AI painting interface of the image generation system provided by the embodiment of the present application Figure 5 。
[0041] Figure 7f Schematic diagram of the AI painting interface of the image generation system provided by the embodiment of the present application Figure 6 。
[0042] Figure 7g Schematic diagram VII of the AI painting interface of the image generation system provided by the embodiment of the present application
[0043] Figure 8 Schematic diagram of the structure of an image processing device provided by the embodiment of the present application
[0044] Figure 9 Schematic diagram of the structure of a handwriting device provided by the embodiment of the present application
[0045] Figure 10 Schematic diagram of the structure of an electronic device provided by the embodiment of the present application Detailed implementation manners
[0046] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0047] The embodiments of the present application are not exhaustive. They are only schematic diagrams of some embodiments and do not constitute specific limitations on the protection scope of the present disclosure. Without contradiction, each step in an embodiment can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in an embodiment can also be implemented as an independent embodiment, and the order of the steps in an embodiment can be exchanged arbitrarily. Additionally, the optional implementation manners in an embodiment can be combined arbitrarily; furthermore, the embodiments can be combined arbitrarily. For example, some or all of the steps in different embodiments can be combined arbitrarily, and an embodiment can be combined arbitrarily with the optional implementation manners of other embodiments.
[0048] In the embodiments of the present application, if there is no special description or logical conflict, the terms and / or descriptions among the embodiments are consistent and can be cited from each other. The technical features in different embodiments can be combined to form new embodiments according to their internal logical relationships.
[0049] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure.
[0050] Exemplary implementation environment
[0051] The image processing method provided by the present application can be applied to an electronic device running an image generation system. The electronic device can be a handwriting device, for example, it can be executed by a processor in the handwriting device, or the electronic device can also be any processor, server, terminal device, etc. that can perform image correction processing during the image generation process. The above-mentioned handwriting device, or the above-mentioned processor, server, terminal device, etc. are applicable to execute the technical solutions of the embodiments of the present application.
[0052] Exemplary method
[0053] Figure 1 is a schematic flowchart of an image processing method provided by the embodiments of the present application. As Figure 1 shown, the method includes the following steps:
[0054] S101: Obtain the creative input provided by the user.
[0055] In this embodiment, the creative input provided by the user can be an instruction, data, or interaction signal for guiding and controlling image generation. For example, the relevant input content can be input content such as a natural language description, a hand-drawn picture, a style parameter, or a reference image.
[0056] Exemplarily, the above-mentioned creative input includes: a natural language description input by the user, and the natural language description can be text content or text content converted from voice data. For example, the natural language description is "Please generate an image of a boy bending down and picking up a heavy object with both hands".
[0057] Again exemplarily, the above-mentioned creative input is a hand-drawn trajectory input by the user on the image drawing area, and the hand-drawn trajectory is a trajectory drawn by the user on the image drawing area in a hand-drawn manner. The content of the hand-drawn trajectory is, for example, "a boy bending down and picking up a heavy object with both hands".
[0058] Here, both the natural language description and the content of the hand-drawn trajectory can be converted into prompt words or creative themes for generating images.
[0059] The hand-drawn trajectory may include shape information, spatial layout, and position information for generating an image. The shape information may be, for example, a contour or a line.
[0060] In some other examples, the above-mentioned creative input includes: a natural language description input by the user, and a hand-drawn trajectory input by the user on the image drawing area.
[0061] The hand-drawn trajectory can be used as input and combined with the natural language description to jointly guide and control the generation of the image. For example, the user draws a rough contour, and this contour combined with the user's text description can be used to generate a more refined image.
[0062] The above step S101 may include: obtaining the hand-drawn trajectory input by the user in the image drawing area, and / or the natural language description input by the user.
[0063] In this embodiment, the user interaction interface of the image generation system (i.e., the AI painting interface) may provide an image drawing area and an image generation area. The image drawing area and the image generation area may be arranged in a left-right layout or an up-down layout.
[0064] Here, the image drawing area can be used for the user to perform creative input through multi-modal input (gestures, touches, strokes, voice), and can also be used for the user to import a sketch from a local file as creative input. The image generation area can be used to generate an image corresponding to the creative input based on the creative input.
[0065] In some examples, the image generation area can support spatial manipulation of the generated image according to gesture operations, such as zooming, rotating, or panning, etc.
[0066] S102: During the process of generating an image in the image generation area based on the creative input, synchronously detect whether there are quality defects in the generated image.
[0067] In some examples, an image generation model is used to generate an image in the image generation area based on the creative input. The image generation model may include a text-to-image model and an image-to-image model.
[0068] The text-to-image model uses text-to-image technology to generate visual content semantically matching the text description, achieving semantic alignment and content generation from the text modality to the image modality. Its process spans different modalities and transforms abstract text into a specific image.
[0069] The image-to-image model uses image-to-image technology to generate a new image meeting specific target requirements based on the structural and semantic information of the input image, while retaining some key features of the source image, such as the composition method, artistic style, etc.
[0070] Exemplarily, when the creation input is the stroke information (i.e., hand-drawn trajectory) input by the user in the image drawing area, the image generation system triggers a local diffusion generation operation at regular intervals according to the stroke information input by the user, tracks the user's drawing area, and automatically adjusts the generation intensity according to the drawing situation, and corresponding images will be synchronously generated in the image generation area. That is to say, the logic for the image generation system to implement "generating while drawing" is: perform a local diffusion generation once within a certain period of time, synchronize the user's stroke area, and dynamically adjust the generation intensity.
[0071] In some examples, synchronously detecting whether the generated image has quality defects in step S102 may include: using a pre-trained neural network to evaluate whether the generated image has quality defects to obtain a detection result, or calling different intelligent agents capable of implementing different detection functions to detect whether the generated image has quality defects. Here, the intelligent agent can be a small model dedicated to a specific detection task.
[0072] In some examples, when the detection result of the image indicates that the image has quality defects, the detection result includes quality defect description information of the image, for example, it may include specific defect descriptions, defect location information (such as the upper left corner, central area of the image, etc.).
[0073] S103: When it is detected that the image has quality defects, display a correction prompt for the image with quality defects.
[0074] In some examples, when it is detected that the image has quality defects (such as color imbalance, semantic inconsistency, incorrect human body structure or object position, etc.), a correction prompt for the image with quality defects can be displayed, and the correction prompt may include but is not limited to at least one of the following: defect location, defect cause analysis, adjustment direction, user-operable parameters, or correction suggestions.
[0075] S104: In response to receiving a correction instruction input by the user based on the correction prompt, correct the image displayed in the image generation area.
[0076] In some examples, the correction instruction can be an adoption instruction input for the correction prompt.
[0077] In some examples, a correction instruction input by the user based on the correction prompt can be received during the image generation process (i.e., intermediate steps of image generation). In response to the correction instruction, image correction parameters are generated, and local diffusion regeneration is performed according to the image correction parameters, while retaining the generation features of adjacent regions; after edge fusion of the regeneration result and the original image, it is displayed in the image generation area. Among them, the image correction parameters may include at least one of a conditional guidance scale and the number of denoising steps.
[0078] Exemplarily, when the user is inputting an image of a person bending down to lift a heavy object into the drawing area, during the process of generating an image synchronously with the handwriting, the image generation system understands the input content in real time and gradually determines that the theme is "a person bending down to lift a heavy object". If a quality defect is detected in the image during the image generation process, a correction prompt will be actively displayed to the user. For example, when it is recognized that the "center of gravity is unstable" in the image, the correction suggestion "move the head and neck forward to shift the center of gravity forward" will be synchronized in the image. After the user inputs an adoption instruction for the correction suggestion, the image generation system will correct the generated image, and the one-key replacement generation effect can be achieved to optimize the image effect.
[0079] It can be understood that the system corrects the image with quality defects during the image generation process. For example, it can modify the composition of the image or correct the image elements that do not meet the expectations (such as incorrect human body structures or object positions), which can ensure that the final image generation result is closer to the user's intention.
[0080] In some other examples, in the case where no quality defect is detected in the image, an affirmative prompt for the image without quality defects is displayed. For example, the affirmative prompt is "the image is qualified". Another example is that the specific dimension indicators for the image to be qualified can be listed, such as reasonable image layout, consistent image content with the theme, etc. In this way, positive feedback and more drawing knowledge can be transmitted to the user, which is convenient for further enhancing the user's understanding and drawing ability.
[0081] The embodiment of the present application provides an image processing method. By obtaining the creative input provided by the user, during the process of generating an image in the image generation area based on the creative input, it is synchronously detected whether the generated image has quality defects, and in the case where a quality defect is detected in the image, a correction prompt for the image with quality defects is displayed; and in response to receiving a correction instruction input by the user based on the correction prompt, the image displayed in the image generation area is corrected. In this way, compared with the end-to-end image generation method, the embodiment of the present application monitors the image quality in real time during the image generation process and corrects the image during the generation process. At the same time, the user is allowed to intervene during the image generation process to correct the deviation in the intermediate steps of image generation. Therefore, the controllability of the image generation result can be enhanced, ensuring that the image generation result is closer to the user's intention, thereby effectively improving the quality and accuracy of image generation.
[0082] In one embodiment, an image processing method is provided, as Figure 2 shown, the method includes the following steps:
[0083] S201: When the creative input provided by the user includes a handwritten trajectory and a natural language description, extract spatio-temporal features from the handwritten trajectory and extract text semantics from the natural language description;
[0084] S202: Align and then fuse the spatio-temporal features and text semantics, and determine the user's creative intention based on the fusion result;
[0085] S203: Generate an image in the image generation area based on the user's creative intention.
[0086] In some examples, for the way of obtaining the creative input, reference can be made to Figure 1 the exemplary implementation of step S101 in Figure 1 and other related parts in the involved embodiments, which will not be elaborated here.
[0087] In some examples, in step S201, a temporal convolutional network (TCN) or Transformer can be used to extract the spatio-temporal features of the brushstrokes from the hand-drawn trajectory (brushstroke information), and a pre-trained language model (such as BERT) can be used to extract the text semantics from the text described in natural language.
[0088] In step S202, after aligning the spatio-temporal features of the brushstrokes and the text semantics using the cross-attention mechanism, the spatio-temporal features and the text semantic vectors can be fused using a fusion method such as feature concatenation or weighted summation to obtain a fusion result. Then, based on the fusion result, a classifier is used to predict the user's creative intention (such as image style, creative theme, etc.).
[0089] In step S203, taking the user's creative intention as the conditional input, a diffusion model or a generative adversarial network can be used to generate an image.
[0090] In some examples, the image generation system can adopt a two-stream Transformer architecture to support the capture and synchronization of hand-drawn trajectories, and can obtain the brushstroke information of the user in real time during the drawing process. At the same time, it supports text semantic input and parsing, and can assist the image generation process by combining the text description input by the user.
[0091] In this way, by fusing the spatio-temporal features extracted from the hand-drawn trajectory and the text semantics extracted from the natural language description to form multi-modal features, the user's intention can be understood more comprehensively and accurately, and the deviation of the generated image (such as semantic inconsistency, incorrect human body structure or object position, etc.) can be reduced.
[0092] In some embodiments, steps S201 to S203 can be combined into an independent embodiment for implementation to realize the process of generating an image in the image generation area based on the creative input.
[0093] In other embodiments, steps S201 to S203 can be combined with Figure 1 one or more steps of steps S101 to S104 in
[0094] In one embodiment, in the above step S102, during the process of generating an image in the image generation area based on the creative input, synchronously detecting whether there are quality defects in the generated image may include the steps:
[0095] During the process of generating an image in the image generation area based on the creative input, synchronously detecting whether there are quality defects in the generated image according to at least one preset detection dimension; wherein, the preset detection dimension includes at least one of the following dimensions: spatial rationality, physical correctness, aesthetic coordination, and semantic consistency.
[0096] Here, spatial rationality means that the spatial layout of objects or elements in the image conforms to the spatial logic of the real world, including three-dimensional spatial attributes such as position, proportion, and perspective. For example, the trees on both sides of the road should show a gradual change from large to small in the distance.
[0097] Here, physical correctness means that the image content conforms to the physical laws of nature, including reasonable manifestations of physical attributes such as lighting, material, and movement. For example, the direction of the shadow should be consistent with the position of the light source.
[0098] Here, aesthetic coordination refers to the visual aesthetic quality of the image elements, including the coordination of visual effect degrees such as color, composition, and style.
[0099] Here, semantic consistency refers to the degree of matching between the image content and the expected semantic description, including the accuracy of object categories, attributes, and relationships. For example, when the creative theme is "a person bending down to pick up a bucket", both a person and a bucket must exist in the generated image.
[0100] For example, when generating a human image, synchronously detecting whether the body proportion of the human in the generated human image is reasonable, whether the actions conform to physical laws, and whether the color matching is coordinated, etc.
[0101] In this embodiment, by performing real-time monitoring on the image generation process from dimensions such as spatial rationality, physical correctness, and aesthetic coordination, a real-time quality evaluation mechanism is provided for the entire image generation process, which can improve the accuracy of the image generation result and the degree of meeting user expectations.
[0102] In one embodiment, during the process of generating an image in the image generation area based on the creative input, synchronously detecting whether there are quality defects in the generated image according to at least one preset detection dimension may include:
[0103] During the process of generating an image in the image generation area based on the creative input, performing a detection operation on the generated image to determine whether there are quality defects in the generated image.
[0104] As Figure 3 shown, the above detection operation includes:
[0105] S301: Extract the content features of the image;
[0106] S302: According to the category of the content features, call the target agent from multiple agents to detect the content features, and obtain the detection result of the content features and the corresponding confidence level; wherein, the agents or agent combinations corresponding to different categories of content features are different, and different agents are used to implement the image detection functions under different preset detection dimensions;
[0107] S303: When the confidence level corresponding to the detection result of the content features reaches the confidence level threshold and the detection result indicates that the content features are abnormal features, it is determined that the image has a quality defect.
[0108] Here, the agents or agent combinations corresponding to different categories of content features are different, and an agent combination may include two or more agents.
[0109] Here, the relationship between the category to which the content features belong and the agents is determined based on the image detection functions of the agents.
[0110] An agent, that is, an entity with intelligence, can be an algorithm model that can complete specified intelligent analysis and decision-making tasks, such as a large model agent (LLM Agent), or an expert rule agent (Rule-based Agent), or a "human agent (Human-Agent)" implemented based on an artificial collaboration method. Here, the agent has the ability to perform complex tasks such as image recognition, natural language processing, and predictive analysis by identifying patterns and features.
[0111] For example, the agent types can include: a physical agent, which can be used to detect the biomechanical rationality in real time; an aesthetic agent, which can be used to evaluate the coordination of the picture style; a semantic agent, which can be used to verify the consistency between the content and the theme; a resource scheduling agent, which can be used to dynamically allocate computing resources; a cache preloading agent, which can be used to predict subsequent generation requirements and preload models.
[0112] In some examples, the category of the content features includes at least one of the following:
[0113] The human joint area, and the corresponding combined agent to be called is: a physical agent and a semantic agent. The physical agent is used to perform biomechanical evaluation on the human joint area, and the semantic agent is used to perform pose semantic verification on the human joint area;
[0114] The background area, and the corresponding combined agent to be called is: an aesthetic agent and a cache agent. The aesthetic agent is used to provide style suggestions, and the cache agent is used to cache the preloaded image resources;
[0115] The text annotation area corresponds to the invocation of combined intelligent agents: semantic intelligent agent and physical intelligent agent. The semantic intelligent agent is used for checking the consistency between text and images, and the physical intelligent agent is used for evaluating the mechanics of image layout.
[0116] Here, the confidence level of the detection result represents the reliability of the detection result. The higher the confidence level of the detection result output by the intelligent agent, the more reliable the detection result.
[0117] In some examples, the detection result characterizes the content feature as an abnormal feature. The detection result also includes: the location information of the abnormal feature. The detection result can be fed back to the diffusion model or the generative adversarial network to trigger local regeneration in the area where the location information of the abnormal feature is located.
[0118] In one embodiment, the image generation area is divided into multiple sub-areas, and each sub-area is configured with a detection unit; in the above steps, performing a detection operation on the generated image to determine whether there are quality defects in the generated image may include:
[0119] Performing the detection operation on the image generated by the sub-area to which the detection unit belongs through the detection unit of any one sub-area;
[0120] In the case where the detection unit of any one of the sub-areas determines that there are quality defects in the image generated by the sub-area to which the detection unit belongs, it is determined that there are quality defects in the generated image.
[0121] Here, through spatial division, for example, the image generation area can be evenly divided into multiple sub-areas by a grid, and each sub-area is configured with an independent detection unit. The detection unit can dynamically call one or more of the distributed intelligent agents as needed.
[0122] In some other examples, the granularity of sub-area division can be dynamically adjusted according to the type of image generation task. For example, a higher density of detection unit deployment is adopted in the key area compared to the non-key area. Here, the key area can be the area for presenting the foreground of the image, and the non-key area can be the area for presenting the background of the image.
[0123] In some examples, distributed intelligent agents can be pre-deployed in the image generation system. Each detection unit can call one or more intelligent agents in the distributed intelligent agent. In this way, after the detection unit extracts the content features from the image generated by the sub-area to which the detection unit belongs, it can call the target intelligent agent from multiple intelligent agents to detect the content features according to the category of the content features to determine whether there are quality defects in the image generated by the sub-area.
[0124] In this embodiment, during the image generation process, for any sub-region in the image generation region, the detection unit configured for this sub-region performs a detection operation on the image generated by this sub-region to determine whether there are quality defects in the image generated by this sub-region. And in the case where the detection unit of any sub-region determines that there are quality defects in the image generated by the sub-region to which the detection unit belongs, it is determined that there are quality defects in the generated image. In this way, by dividing the image generation region into multiple sub-regions and using the detection units of each sub-region to parallelly call the corresponding intelligent agents to perform detection tasks, if any detection unit detects a quality defect in the image, it can be determined that there are quality defects in the generated image. Thus, through the parallel processing method, the computing load can be greatly reduced, and efficient and real-time quality monitoring during the image generation process can be achieved. In addition, compared with the fixed pipeline intelligent agent architecture for evaluating image quality, adopting the distributed intelligent agent architecture can load intelligent agents on demand and reduce redundant calculations.
[0125] In one embodiment, the image processing method further includes:
[0126] In the case where it is detected that there are quality defects in the image, a floating window is displayed in the image generation region, where the floating window is used to display the quality defect description information of the image; in response to an interaction operation acting on the floating window, the analysis content of the quality defect of the image is displayed.
[0127] In this embodiment, in the case where it is detected that there are quality defects in the generated image during the image generation process, a series of operations for actively serving users will be triggered. For example, the system will dynamically display a floating window in the image generation region, which can display the quality defect description information of the image to the user without disturbing the user's viewing of the generated image.
[0128] In some examples, the position, size, and style of the floating window can be intelligently adjusted according to the layout of the image generation region and the user's operation habits to ensure its visibility.
[0129] In some examples, the quality defect description information of the image may include specific defect descriptions, defect positions (such as the upper left corner, center region, etc. of the image), and in addition, may also include the severity of the defect (such as minor, medium, severe), so that users can quickly understand the problems existing in the image.
[0130] When the user clicks on or hovers over the floating window, the system will trigger an interaction event, and then display the analysis content of the quality defect of the image. The analysis content may include the cause of the defect, the impact on the overall quality of the image, and possible repair suggestions.
[0131] In this embodiment, when the system detects that the generated image has quality defects during the image generation process, it provides the user with relevant defect descriptions and analysis content. Through proactive user services, it helps the user better understand the image problems, thereby enhancing the user's interaction experience.
[0132] In one embodiment, when it is detected that the image has quality defects, a correction prompt for the image with quality defects is displayed, including:
[0133] When it is detected that the image has quality defects, an operation control is displayed, where the operation control is used to provide a correction prompt for the image with quality defects when triggered; in response to the operation control being triggered, a correction prompt is displayed in the area where the hand-drawn trajectory is located and / or the image generation area.
[0134] In some examples, when it is detected that the image has quality defects, the system will automatically display an operation control in the image generation area. For example, the operation control can be dynamically displayed at the image defect location, and the operation control can be a button or an icon.
[0135] In this embodiment, when the image has quality defects, an operation control is provided, and in response to the operation control being triggered, a correction prompt is displayed in the area where the hand-drawn trajectory is located and / or the image generation area. In this way, there is no need for the user to manually select or switch the interface, saving efficiency and being more able to provide problem analysis and guiding suggestions to the user in a targeted manner.
[0136] In one embodiment, the correction prompt includes: correction suggestions and / or correction lines; displaying the correction prompt in the area where the hand-drawn trajectory is located and / or the image generation area includes:
[0137] Floating the correction suggestions in the image generation area; and / or,
[0138] Displaying the correction lines on the image and / or the hand-drawn trajectory as the creative input.
[0139] Exemplarily, the correction suggestions are displayed in text form in the non-image area within the image generation area.
[0140] Exemplarily, the correction lines are displayed on the hand-drawn trajectory or the generated image in the form of semi-transparent lines or highlighted contour lines to be corrected, so as to indicate the correction direction and scope.
[0141] In this embodiment, by providing the user with image modification suggestions and correction lines, such a "what you see is what you get" visual guidance can more intuitively guide the user to understand the image problems and image correction compared to abstract parameter adjustment.
[0142] In one embodiment, as Figure 4As shown, the image processing method may include:
[0143] S401: According to the creative input and the image content already generated in the image generation area, predict the resource requirements for the content to be generated for image generation within a preset future time period, and preload the resources corresponding to the resource requirements from the resource library.
[0144] Exemplarily, a machine learning model (e.g., a spatio-temporal prediction model or a generative adversarial network) may be used to predict the resources required for image generation in a future time period (e.g., within 3 seconds). According to the prediction result, resources (e.g., image materials, model components, etc.) are loaded from the resource library according to priority and cached in the memory.
[0145] In this embodiment, an active service mechanism based on creative trajectory, text, and image analysis may be constructed to deeply analyze the user's creative trajectory (such as painting strokes), input text, and generated images, so as to implement a mechanism for actively providing services to the user.
[0146] In some embodiments, step S401 may be implemented as an independent embodiment.
[0147] In other embodiments, step S401 may be combined with one or more of the steps Figures 1 to 3 shown to form an independent embodiment for implementation.
[0148] In one embodiment, an image processing method is provided. As Figure 5 shown, the method may include the following steps:
[0149] S501: During the process of generating an image in the image generation area based on the user's creative input, synchronously detect whether the generated image has quality defects;
[0150] S502: When it is detected that the image has quality defects, locate the defect position in the image and generate corresponding knowledge guidance in the image generation area.
[0151] In some examples, the acquisition method of the creative input may refer to the exemplary implementation manner of step S101 in Figure 1 and other related parts in the embodiments involved in Figure 1 , which will not be elaborated here.
[0152] In some examples, the implementation manner of step S501 may refer to the exemplary implementation manner of step S102 in Figure 1 and other related parts in the embodiments involved in Figure 1 , which will not be elaborated here.
[0153] In some examples, in the case where a quality defect is detected in the image, step S502 can locate the defect position in the image, and analyze the associated knowledge that causes the quality defect based on the quality defect of the image through a graph neural network to generate corresponding knowledge guidance.
[0154] In this embodiment, when a quality problem of the image is detected during the image generation process, by locating the defect position in the image and generating corresponding knowledge guidance in the image generation area, such a method of "active teaching" is adopted, so that the generation result is not separated from the teaching guidance. The user can obtain the teaching guidance without switching between interfaces, which is convenient for the user to understand and can help improve the user's creative ability, thereby enhancing the user experience.
[0155] In some embodiments, steps S501 to S502 can be combined into an independent embodiment for implementation.
[0156] In some other embodiments, steps S501 to S502 can be combined with Figures 1 to 4 one or more steps in the steps shown into an independent embodiment for implementation.
[0157] In one embodiment, as Figure 6 shown, generating corresponding knowledge guidance in the image generation area in the above step S502 may include the following steps:
[0158] S601: Query the target knowledge resources related to the quality defect in the knowledge graph according to the quality defect of the image;
[0159] S602: Schedule each target knowledge resource using a priority algorithm;
[0160] S603: Generate corresponding knowledge guidance in the image generation area according to each scheduled target knowledge resource.
[0161] In some examples, in step S601, defect keywords can be extracted according to the quality defect of the image, and multi-hop query tracing can be performed in the knowledge graph according to the defect keywords. For example, the associated target knowledge resources can be queried in the knowledge graph along a preset relationship path according to the defect keywords.
[0162] Here, the knowledge graph includes the association relationships between multiple entity nodes. Among them, the entity nodes can include defect nodes, cause nodes, and resource nodes. The preset relationship path can be, for example, defect node -> cause node -> algorithm node.
[0163] Here, the knowledge resources can include but are not limited to: algorithm resources, tool resources, literature resources, etc.
[0164] In some examples, in step S602, a resource scheduling agent can be called to execute a priority decision algorithm to dynamically schedule each target knowledge resource. For example, the priority of algorithm resources is higher than that of tool resources, and the priority of tool resources is higher than that of literature resources.
[0165] In some examples, in step S603, the scheduled target knowledge resources can be converted into structured data, and corresponding knowledge guidance can be generated based on the structured data to be displayed in the non-image area within the image generation area.
[0166] Compared with the situation where the image generation result and the teaching guidance are separated from each other and the user needs to independently switch to an independent function module to obtain the teaching guidance, in this embodiment, when an error occurs during the image generation process, not only can the error of the image be accurately located, but also knowledge explanations can be provided to the user without the user switching the interface, which is convenient for the user to understand and modify.
[0167] In one embodiment, generating corresponding knowledge guidance in the image generation area according to each scheduled target knowledge resource in the above step S603 may include:
[0168] Performing multimodal content assembly on each scheduled target knowledge resource to be converted into a guidance plan for the user's reference in the image generation area; and / or, anchoring the knowledge elements extracted from the scheduled target knowledge resources to the defect positions of the image and floatingly displaying the corresponding knowledge guidance on the image.
[0169] In some examples, multiple different heterogeneous target knowledge resources can be structurally integrated to be assembled into a unified knowledge guidance plan and displayed in the non-image area within the image generation area.
[0170] In some examples, the knowledge elements extracted from the target knowledge resources can be spatially aligned with the defect positions of the image, and relevant processing suggestions can be floatingly displayed on the image.
[0171] In some examples, an improved SLAM (Simultaneous Localization and Mapping) engine can be used to anchor the knowledge elements extracted from the scheduled target knowledge resources to the defect positions of the image.
[0172] In this embodiment, by converting into a guidance plan for the user's reference in the image generation area, anchoring the knowledge elements extracted from the scheduled target knowledge resources to the defect positions of the image, and floatingly displaying the corresponding knowledge guidance on the image, without the need for the user to switch the interface, three-dimensional teaching guidance and enhanced result presentation are realized, which is further convenient for the user to understand and modify and improves the user experience.
[0173] In one embodiment, the method further includes:
[0174] In response to receiving an operation instruction for the corrected image, perform an operation corresponding to the operation instruction on the corrected image.
[0175] Wherein, the operation instruction includes at least one of image rotation, image scaling, image translation, image parameter adjustment, and image contrast analysis.
[0176] Here, the operation instruction can be triggered by a user gesture. For example, pinch-to-zoom, swipe-to-pan, rotation gesture; or, the operation instruction can also be triggered by user touch. For example, the user adjusts the image by touching a button or a slider on the screen.
[0177] Here, the image parameter adjustment can be a parameter related to the human posture or object structure in the image. For example, the image parameters are adjusted through a real-time mechanics simulator to facilitate the user to observe the force change.
[0178] Here, the image contrast analysis can be an analysis of the pose difference between the correct image and the incorrect image generated in the image library.
[0179] In this embodiment, by responding to the operation instruction of the user on the corrected image, dynamic and intelligent image interaction and knowledge collaboration can be achieved.
[0180] In one embodiment, the method further includes:
[0181] When it is detected that the user draws painting content associated with the user's historical incorrect painting, output a painting prompt to guide the user to draw correctly; wherein, the painting prompt is determined based on the historical incorrect painting.
[0182] For a single painting process, the image generation system can synchronize the correction opinions in the picture for the user to learn, so as to achieve AI's autonomous learning for the user and timely correction during the process.
[0183] For multiple painting processes, the image generation system can record the types of errors that the user often makes (such as always drawing crooked shoulders). When the user draws similar content again, targeted prompts will be given in advance. In this way, by establishing a personalized learning file, as the user's painting level improves, the prompts will gradually decrease, realizing adaptive feedback, and thus can effectively help the user improve the painting skill.
[0184] In one embodiment, the creative input provided by the user includes: the hand-drawn trajectory content input by the user in the image drawing area; generating an image in the image generation area based on the creative input, including:
[0185] Generate an image in the image generation area based on the text content and / or image content included in the hand-drawn trajectory content; wherein, the image is an enhanced image including text content and / or image content.
[0186] In some examples, the text content included in the hand-drawn trajectory content can be, for example, handwritten letters, numbers or symbols, and the image content included in the hand-drawn trajectory content can be, for example, hand-drawn graphics, patterns, outlines of people or objects, etc.
[0187] Here, the enhanced image is an enhancement or extension of the hand-drawn trajectory content, including richer details, colors, textures or other visual elements.
[0188] The enhancement of the image can include improving clarity, increasing color contrast, adding shadow or highlight effects, introducing new image elements (such as background, border or decorative patterns), etc.
[0189] It should be noted that the enhanced image can also include content that is associated with the hand-drawn trajectory content but not directly drawn by the user, such as translation text automatically generated based on the handwritten text by the user or similar patterns automatically generated based on the hand-drawn graphics.
[0190] In this embodiment, generating an enhanced image in the image generation area based on the text content and / or image content included in the hand-drawn trajectory content can improve the image generation effect and provide a more vivid and rich visual experience for the user.
[0191] In one embodiment, the creative input includes: hand-drawn trajectory content input by the user in the image drawing area; the method further includes:
[0192] In response to receiving a modification operation for the target content in the hand-drawn trajectory, synchronously adjust the image generated in the image generation area.
[0193] In some examples, the user may make local modifications (such as erasing, moving, scaling, redrawing) to the hand-drawn trajectory that has been drawn. The modified content may involve text, graphics or a combination of both.
[0194] When the image generation system detects a modification operation by the user for the target content in the hand-drawn trajectory, extract the modified content, including but not limited to: modified area coordinates, modification type (addition, deletion, modification), and modified trajectory information, etc., and dynamically adjust the image generated in the image generation area according to the modified content.
[0195] For example, it can be a local update of the image generated in the image generation area during the image generation process: only regenerate the image part related to the modified area, or, it can also be a global update of the image generated in the image generation area. For example, if the modification affects the overall composition (such as a change in the semantic meaning of the text), then regenerate the complete image.
[0196] In this embodiment, it is possible to achieve dynamic synchronous adjustment of the modification of hand-drawn content and the generated image.
[0197] Next, the image processing method provided by this application will be described by way of example. Figures 7a to 7g Exemplarily illustrate the image processing method provided by this application.
[0198] Refer to Figure 7a the schematic diagram of the AI painting interface of the image generation system shown Figure 1 , the user is inputting an image of a person bending down to lift a heavy object into the drawing area. During the process of synchronizing the handwriting, the system understands the input content in real time and gradually determines that the theme is "a person bending down to lift a heavy object".
[0199] Based on the implementation logic of "generating while drawing", during the drawing process, every few new strokes drawn by the user, the AI will improve the image in the background, but each stroke of the user will be preferentially retained. Relying on the underlying model training, the generation idea and the resources to be called behind are adjusted in real time, and active services are prepared to be provided in real time.
[0200] Refer to Figure 7b the schematic diagram of the AI painting interface of the image generation system shown Figure 2 , when the system detects a quality defect in the image during the image generation process, it will actively remind the user. For example, when the system recognizes that the "center of gravity is unstable" in the image, it points out the key problem points, displays the intelligent annotation floating window and prompts relevant content in the intelligent annotation floating window "According to the drawing trajectory, it is found that the center of gravity is unstable...".
[0201] Refer to Figure 7c the schematic diagram of the AI painting interface of the image generation system shown Figure 3 , if the user clicks on Figure 7b the intelligent annotation floating window in, the system will synchronously display the AI warning content and the basis for generating the reminder content. For the problem analysis provided by the AI, the user can perform interactive operations to understand the cause of the problem more deeply. For example, after the user clicks on the intelligent annotation floating window, the system can analyze the center of gravity area of the image based on elements such as the landing point of the person drawn by the user, the trend of the shoulders and neck, the trend of the waist, and the body movement line, and give the suggestion that the user's center of gravity is unstable. At the same time, the relevant expression knowledge about the center of gravity of the human body will be provided to the user in the form of text, video, pictures, interactive software, etc. for reference, to achieve three-dimensional teaching guidance and enhanced result presentation. Further, the user can click on the operation control showing the correction suggestion to trigger the system to display the correction prompt in the drawing area and the image generation area.
[0202] Refer to Figure 7d the schematic diagram of the AI painting interface of the image generation system shown Figure 4 , the user clicks on Figure 7cAfter the operation control with correction suggestions is displayed, the AI correction suggestions "Head and neck forward, shift the center of gravity forward" can be viewed in the image generation area, and the correction suggestions can be used by clicking to achieve one-key replacement of the generation effect.
[0203] From Figures 7b to 7d It can be directly seen that all the guiding information in the AI painting generation process will be directly superimposed on the AI-synchronized screen. Without the need for the user to switch interfaces, it not only saves efficiency but also can specifically point out problem analysis and guiding suggestions.
[0204] Exemplarily, as Figure 7e shown in the schematic diagram of the AI painting interface of the image generation system Figure 5 When the system detects that the image has a quality defect of "poor performance of bones and muscles", it will locate the defect position in the image and generate corresponding knowledge guidance in the image generation area.
[0205] It can be understood that when the system detects that the image has no quality defect, it can also give knowledge guidance for the generated image. For example, Figure 7f shown in the schematic diagram of the AI painting interface of the image generation system Figure 6 for the AI image with the theme of "a person holding a heavy object", knowledge explanations related to physics are generated in the image generation area.
[0206] Referring to Figure 7g shown in the seventh schematic diagram of the AI painting interface of the image generation system, the system can generate enhanced images for the hand-drawn trajectory in the form of text and images through the method of "text + image + automatic text content supplementation + automatic image expansion", and the image generation result supports overall and partial redrawing.
[0207] In summary, the technical solution provided by the embodiments of the present application deeply integrates technologies such as generative AI and augmented reality, and constructs the first intelligent creation system that supports process intervention, error traceability, and knowledge embedding. It realizes real-time process recognition of content, provides active suggestions, and conducts content analysis and correction without departing from the original content of the user's screen. This design can not only record the creation process for the entire AI painting generation but also guide the correction direction in real time, making learning and creation more intuitive and efficient.
[0208] Exemplary device
[0209] Embodiments of the present application provide an image processing device, as Figure 8 shown, the device includes:
[0210] An acquisition unit 101, configured to acquire the creation input provided by the user;
[0211] The first processing unit 102 is configured to synchronously detect whether there are quality defects in the generated image during the process of generating an image in the image generation area based on the creation input;
[0212] The display unit 103 is configured to display a correction prompt for the image with quality defects in case it is detected that the image has quality defects;
[0213] The correction unit 104 is configured to correct the image displayed in the image generation area in response to receiving a correction instruction input by the user based on the correction prompt.
[0214] In some examples, the creation input includes: a natural language description input by the user and / or a hand-drawn trajectory input by the user in the image drawing area; the first processing unit 102 is configured to:
[0215] In the case where the creation input includes a hand-drawn trajectory and a natural language description, extract spatio-temporal features from the hand-drawn trajectory and extract text semantics from the natural language description;
[0216] Align and fuse the spatio-temporal features and the text semantics, and determine the user's creation intention based on the fusion result;
[0217] Generate an image in the image generation area based on the user's creation intention.
[0218] In some examples, the first processing unit 102 is configured to:
[0219] During the process of generating an image in the image generation area based on the creation input, synchronously detect whether there are quality defects in the generated image according to at least one preset detection dimension; wherein, the preset detection dimension includes at least one of the following dimensions: spatial rationality, physical correctness, aesthetic coordination, and semantic consistency.
[0220] In some examples, the image generation area is divided into multiple sub-areas, and each sub-area is configured with a detection unit; the first processing unit 102 is configured to:
[0221] During the process of generating an image in the image generation area based on the creation input, perform a detection operation on the generated image to determine whether there are quality defects in the generated image;
[0222] Wherein, the detection operation includes:
[0223] Extract the content features of the image;
[0224] According to the category of the content features, a target agent is called from multiple agents to detect the content features, and a detection result of the content features and a corresponding confidence level are obtained; wherein, the agents or agent combinations corresponding to different categories of the content features are different, and different agents are used to implement image detection functions under different preset detection dimensions;
[0225] When the confidence level corresponding to the detection result of the content features reaches the confidence level threshold and the detection result indicates that the content features are abnormal features, it is determined that the image has a quality defect.
[0226] In some examples, the image generation area is divided into multiple sub-areas, and each sub-area is configured with a detection unit; the first processing unit 102 is used for:
[0227] Performing the detection operation on the image generated by the sub-area to which the detection unit belongs through the detection unit of any one of the sub-areas;
[0228] When the detection unit of any one of the sub-areas determines that the image generated by the sub-area to which the detection unit belongs has a quality defect, it is determined that the generated image has a quality defect.
[0229] In some examples, the display unit 103 is further used for:
[0230] When it is detected that the image has a quality defect, a floating window is displayed in the image generation area, where the floating window is used to display the quality defect description information of the image;
[0231] In response to an interaction operation on the floating window, the analysis content of the quality defect of the image is displayed.
[0232] In some examples, the display unit 103 is used for:
[0233] When it is detected that the image has a quality defect, an operation control is displayed, where the operation control is used to provide a correction prompt for the image with a quality defect when triggered;
[0234] In response to the operation control being triggered, the correction prompt is displayed in the area where the hand-drawn trajectory is located and / or the image generation area.
[0235] In some examples, the correction prompt includes: correction suggestions and / or correction lines; the display unit 103 is used for:
[0236] Floatingly display the correction suggestions in the image generation area; and / or,
[0237] Display the correction line on the image and / or the hand-drawn trajectory as the creation input.
[0238] In some examples, the first processing unit 102 is further configured to:
[0239] According to the creation input and the image content that has been generated in the image generation area currently, predict the resource requirements for the content to be generated for generating the image within a preset future time period, and preload the resources corresponding to the resource requirements from the resource library.
[0240] In some examples, the device further includes a second processing unit, and the second processing unit is configured to:
[0241] When it is detected that the image has quality defects, locate the defect positions in the image, and generate corresponding knowledge guidance in the image generation area.
[0242] In some examples, the second processing unit is configured to:
[0243] Query target knowledge resources related to the quality defects in the knowledge graph according to the quality defects of the image;
[0244] Schedule each of the target knowledge resources using a priority algorithm;
[0245] Generate corresponding knowledge guidance in the image generation area according to each of the scheduled target knowledge resources.
[0246] In some examples, the second processing unit is configured to:
[0247] Perform multimodal content assembly on each of the scheduled target knowledge resources to be transformed into a guidance scheme for the user to reference in the image generation area; and / or,
[0248] Anchor the knowledge elements extracted from the scheduled target knowledge resources to the defect positions of the image, and floatingly display the corresponding knowledge guidance on the image.
[0249] In some examples, the device further includes a third processing unit, and the third processing unit is configured to:
[0250] In response to receiving an operation instruction for the corrected image, perform an operation corresponding to the operation instruction on the corrected image;
[0251] Wherein, the operation instruction includes at least one of image rotation, image scaling, image translation, image parameter adjustment, and image comparison analysis.
[0252] In some examples, the device further includes a fourth processing unit, and the fourth processing unit is configured to:
[0253] When it is detected that the user draws painting content associated with the user's historical incorrect painting, a painting prompt is output to guide the user to draw correctly; wherein, the painting prompt is determined based on the historical incorrect painting.
[0254] In some examples, the creation input includes: the hand-drawn trajectory content input by the user in the image drawing area; the first processing unit 102 is configured to:
[0255] Generate an image in the image generation area based on the text content and / or image content included in the hand-drawn trajectory content; wherein, the image is an enhanced image including the text content and / or the image content.
[0256] In some examples, the creation input includes: the hand-drawn trajectory content input by the user in the image drawing area; the first processing unit 102 is further configured to:
[0257] In response to receiving a modification operation on the target content in the hand-drawn trajectory, synchronously adjust the image generated in the image generation area.
[0258] The image processing apparatus provided in this embodiment belongs to the same application concept as the image processing method provided in the foregoing embodiments of the present application, and can execute the image processing method provided in any foregoing embodiment of the present application, and has corresponding functional modules and beneficial effects for executing the image processing method. For the technical details not described in detail in this embodiment, reference may be made to the specific processing content of the image processing method provided in the foregoing embodiments of the present application, and details are not described herein again.
[0259] The functions implemented by the above-mentioned acquisition unit 101, first processing unit 102, display unit 103, correction unit 104 and other units can be implemented by the same or different processors respectively, and the embodiments of the present application do not make any limitations.
[0260] It should be understood that each unit in the above image processing device can be implemented in the form of a processor invoking software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor invokes the instructions stored in the memory to implement any of the above image processing methods or the functions of each unit of the device. The processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory inside the device or a memory outside the device. Alternatively, the units in the image processing device can be implemented in the form of hardware circuits, and the functions of some or all of the units can be implemented through the design of the hardware circuits. The hardware circuits can be understood as one or more processors. For example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are implemented through the design of the logical relationships of the components in the circuit. Again, for example, in another implementation, the hardware circuit can be implemented through a PLD. Taking an FPGA as an example, it can include a large number of logic gate circuits, and the connection relationships between the logic gate circuits are configured through a configuration file, so as to implement the functions of some or all of the above units. All units of the above device can be all implemented in the form of a processor invoking software, or all implemented in the form of hardware circuits, or some implemented in the form of a processor invoking software, and the remaining part implemented in the form of hardware circuits.
[0261] In the embodiments of the present application, the processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a GPU, or a DSP, etc. In another implementation, the processor can implement certain functions through the logical relationships of hardware circuits, and the logical relationships of the hardware circuits are fixed or can be reconstructed. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as a kind of ASIC, such as an NPU, a TPU, a DPU, etc.
[0262] It can be seen that each unit in the above image processing device can be one or more processors (or processing circuits) configured to implement the above image processing method. For example: a CPU, a GPU, an NPU, a TPU, a DPU, a microprocessor, a DSP, an ASIC, an FPGA, or a combination of at least two of these processor forms.
[0263] In addition, the units in the above image processing device can be fully or partially integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above image processing methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0264] The embodiment of the present application further provides a control device, which includes a processor and an interface circuit. The processor in the control device is connected to a handwriting function component through the interface circuit of the control device.
[0265] The handwriting function component is used to sense the creative input provided by the user. For example, the handwriting function component is a functional component that can sense the user's handwriting operation and generate handwriting content corresponding to the user's handwriting operation as the creative input, such as a professional handwriting board, a touch module integrated with a display screen, etc.
[0266] The above-mentioned interface circuit can be any interface circuit that can realize the data communication function, for example, it can be a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIE circuit, etc.
[0267] The processor in the control device is also a circuit with signal processing capability, which is used to execute any one of the image processing methods described in the above embodiments. The specific implementation of the processor can refer to the above processor implementation, and the embodiments of this application are not strictly limited.
[0268] When the control device is applied to a handwriting device, the handwriting function component of the control device can be the handwriting function component mentioned above in the handwriting device, such as the touch function component mentioned above in a touch screen handwriting device. At the same time, the processor of the control device can be a CPU or GPU, etc. provided in the handwriting device, and the interface circuit of the control device can be an interface circuit between the handwriting function component of the handwriting device and a processor such as a CPU or GPU.
[0269] Exemplary Electronic Devices
[0270] The present application embodiment provides a handwriting device, see Figure 9 As shown, the handwriting device includes a handwriting function component 300 and a processor 310 connected to the handwriting function component.
[0271] The handwriting function component 300 is used to sense the creative input provided by the user;
[0272] The processor 310 is used to execute any image processing method provided by any of the above embodiments.
[0273] The above-mentioned handwriting function component 300 can be any function component that supports users to write by hand. For example, it can be a touchpad, a touch screen, etc. For the specific handwriting trajectory acquisition and handwriting content generation and processing process, reference can be made to the handwriting function implementation process in the conventional technology.
[0274] For the specific processing process of the above-mentioned processor 310, reference can be made to the introduction in the above method embodiment. For the specific implementation manner of the processor 310, reference can also be made to the introduction in the above embodiment.
[0275] This handwriting device can specifically be a terminal device with handwriting function, such as a handwriting tablet, a handheld terminal with handwriting function, a wearable terminal, a computer with handwriting function, a smart terminal, etc.
[0276] Another embodiment of this application also proposes an electronic device. Refer to Figure 10 As shown, this device includes:
[0277] A memory 200 and a processor 210;
[0278] Wherein, the memory 200 is connected to the processor 210 and is used for storing programs;
[0279] The processor 210 is used for implementing the image processing method provided in any of the above embodiments by running the programs stored in the memory 200.
[0280] Specifically, the above-mentioned electronic device may further include: a bus, a communication interface 220, an input device 230, and an output device 240.
[0281] The processor 210, the memory 200, the communication interface 220, the input device 230, and the output device 240 are interconnected through the bus. Among them:
[0282] The bus may include a path for transmitting information between various components of the computer system.
[0283] The processor 210 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention solution. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0284] The processor 210 may include a main processor and may also include a baseband chip, a modem, etc.
[0285] The program for implementing the technical solution of the present invention is stored in the memory 200, and the operating system and other key services may also be stored. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.
[0286] The input device 230 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.
[0287] The output device 240 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.
[0288] The communication interface 220 may include devices of any transceiver type for communicating with other devices or communication networks, such as an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.
[0289] The processor 210 executes the program stored in the memory 200 and calls other devices, and can be used to implement each step of any one of the image processing methods provided in the above embodiments of the present application.
[0290] An embodiment of the present application also provides a chip, which includes a processor and a data interface. The processor reads and runs the program stored on the memory through the data interface to execute the image processing method introduced in any of the above embodiments. The specific processing process and its beneficial effects can be referred to the embodiment introduction of the above image processing method.
[0291] Exemplary computer program product and storage medium
[0292] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the image processing method according to various embodiments of the present application described in any of the above embodiments of this specification.
[0293] The computer program product can be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, executed as an independent software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0294] In addition, an embodiment of the present application can also be a storage medium on which a computer program is stored. The computer program is executed by a processor to perform the steps in the image processing method of the intelligent door lock according to various embodiments of the present application described in any of the above embodiments of this specification. Specifically, the following steps can be implemented:
[0295] Obtain the creative input provided by the user;
[0296] During the process of generating an image in the image generation area based on the creative input, synchronously detect whether the generated image has quality defects;
[0297] In the case where it is detected that the image has quality defects, display a correction prompt for the image with quality defects;
[0298] In response to receiving a correction instruction input by the user based on the correction prompt, correct the image displayed in the image generation area.
[0299] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0300] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0301] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0302] In each embodiment of the present application, the modules and sub-modules in the device and the terminal can be combined, divided, and deleted according to actual needs.
[0303] In several embodiments provided in the present application, it should be understood that the disclosed terminal, device, and method can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or module can be in electrical, mechanical, or other forms.
[0304] The module or sub-module described as a separated component may or may not be physically separated. The component as a module or sub-module may or may not be a physical module or sub-module, that is, it can be located in one place, or it can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0305] In addition, each functional module or sub-module in each embodiment of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above integrated module or sub-module can be implemented in the form of hardware or in the form of a software functional module or sub-module.
[0306] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0307] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented directly in hardware, in a software unit executed by a processor, or in a combination thereof. The software unit may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0308] Finally, it should also be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0309] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image processing method, characterized in that, The method includes: Obtaining the creative input provided by the user; During the process of generating an image in the image generation area based on the creative input, synchronously detecting whether there are quality defects in the generated image; When it is detected that the image has quality defects, displaying a correction prompt for the image with quality defects; In response to receiving a correction instruction input by the user based on the correction prompt, correcting the image displayed in the image generation area.
2. The method according to claim 1, wherein The creative input includes: the natural language description input by the user and / or the hand-drawn trajectory input by the user in the image drawing area; generating an image in the image generation area based on the creative input includes: When the creative input includes a hand-drawn trajectory and a natural language description, extracting spatio-temporal features from the hand-drawn trajectory and extracting text semantics from the natural language description; Aligning and fusing the spatio-temporal features and the text semantics, and determining the user's creative intention based on the fusion result; Generating an image in the image generation area based on the user's creative intention.
3. The method according to claim 1, wherein The step of synchronously detecting whether there are quality defects in the generated image during the process of generating an image in the image generation area based on the creative input includes: During the process of generating an image in the image generation area based on the creative input, synchronously detecting whether there are quality defects in the generated image according to at least one preset detection dimension; wherein, the preset detection dimension includes at least one of the following dimensions: spatial rationality, physical correctness, aesthetic coordination, and semantic consistency.
4. The method according to claim 3, wherein The step of synchronously detecting whether there are quality defects in the generated image during the process of generating an image in the image generation area based on the creative input includes: During the process of generating an image in the image generation area based on the creative input, performing a detection operation on the generated image to determine whether there are quality defects in the generated image; Wherein, the detection operation includes: Extracting the content features of the image; According to the category of the content features, calling a target agent from multiple agents to detect the content features, obtaining the detection result of the content features and the corresponding confidence level; wherein, different agents or agent combinations correspond to different categories of content features, and different agents are used to implement image detection functions under different preset detection dimensions; When the confidence level corresponding to the detection result of the content features reaches the confidence threshold and the detection result indicates that the content features are abnormal features, it is determined that the image has quality defects.
5. The method according to claim 4, characterized in that The image generation area is divided into multiple sub-areas, and each sub-area is configured with a detection unit; The step of performing a detection operation on the generated image to determine whether there are quality defects in the generated image includes: Performing the detection operation on the image generated by the sub-area to which the detection unit of any one of the sub-areas belongs through the detection unit of the sub-area; When the detection unit of any one of the sub-areas determines that the image generated by the sub-area to which the detection unit belongs has quality defects, it is determined that the generated image has quality defects.
6. The method according to claim 1, wherein The method further includes: When it is detected that the image has a quality defect, a floating window is displayed in the image generation area, where the floating window is used to display the quality defect description information of the image; In response to an interaction operation on the floating window, the analysis content of the quality defect of the image is displayed.
7. The method according to claim 1, characterized in that, The step of, when it is detected that the image has a quality defect, displaying a correction prompt for the image with the quality defect includes: When it is detected that the image has a quality defect, an operation control is displayed, where the operation control is used to provide a correction prompt for the image with the quality defect when triggered; In response to the operation control being triggered, the correction prompt is displayed in the area where the hand-drawn trajectory is located and / or the image generation area.
8. The method according to claim 7, wherein The correction prompt includes: correction suggestions and / or correction lines; The step of displaying the correction prompt in the area where the hand-drawn trajectory is located and / or the image generation area includes: Floatingly displaying the correction suggestions in the image generation area; and / or, Displaying the correction lines on the image and / or the hand-drawn trajectory as the creative input.
9. The method according to claim 1, wherein The method further includes: Based on the creative input and the image content that has been generated in the image generation area currently, predicting the resource requirements for the content to be generated for generating the image within a preset future time period, and preloading the resources corresponding to the resource requirements from the resource library.
10. The method according to claim 1, wherein The method further includes: When it is detected that the image has a quality defect, locating the defect position in the image, and generating corresponding knowledge guidance in the image generation area.
11. The method according to claim 10, wherein The step of generating corresponding knowledge guidance in the image generation area includes: Querying target knowledge resources related to the quality defect in the knowledge graph according to the quality defect of the image; Scheduling each of the target knowledge resources using a priority algorithm; Generating corresponding knowledge guidance in the image generation area according to each of the scheduled target knowledge resources.
12. The method according to claim 11, wherein The step of generating corresponding knowledge guidance in the image generation area according to each of the scheduled target knowledge resources includes: Performing multi-modal content assembly on each of the scheduled target knowledge resources to be transformed into a guidance scheme for the user to refer to in the image generation area; and / or, Anchoring the knowledge elements extracted from the scheduled target knowledge resources to the defect position of the image, and floatingly displaying the corresponding knowledge guidance on the image.
13. The method according to claim 1, characterized in that, The method further includes: In response to receiving an operation instruction for the corrected image, performing an operation corresponding to the operation instruction on the corrected image; Wherein, the operation instruction includes at least one of image rotation, image scaling, image translation, image parameter adjustment, and image comparison analysis.
14. The method according to claim 1, wherein The method further includes: When it is detected that the user draws painting content associated with the user's historical incorrect paintings, outputting a painting prompt to guide the user to draw correctly; wherein the painting prompt is determined based on the historical incorrect paintings.
15. The method according to claim 1, wherein The creative input includes: the hand-drawn trajectory content input by the user in the image drawing area; Generating an image in the image generation area based on the creation input includes: Generating an image in the image generation area based on the text content and / or image content included in the hand-drawn trajectory content; wherein, the image is an enhanced image including the text content and / or the image content.
16. The method according to any one of claims 1 to 15, characterized in that, The creation input includes: hand-drawn trajectory content input by the user in the image drawing area; The method further includes: In response to receiving a modification operation for target content in the hand-drawn trajectory, synchronously adjusting the image generated in the image generation area.
17. An image processing apparatus, characterized in that, The device includes: An acquisition unit for acquiring the creation input provided by the user; A first processing unit for synchronously detecting whether there are quality defects in the generated image during the process of generating an image in the image generation area based on the creation input; A display unit for displaying a correction prompt for the image with quality defects when it is detected that the image has quality defects; A correction unit for correcting the image displayed in the image generation area in response to receiving a correction instruction input by the user based on the correction prompt.
18. A handwriting device, characterized in that, Includes: A handwriting function component and a processor connected to the handwriting function component; The handwriting function component is used to sense the creation input provided by the user; The processor is used to execute the image processing method according to any one of claims 1 to 16.
19. A control device, characterized in that, Includes a processor and an interface circuit, and the processor is connected to the handwriting function component through the interface circuit; The handwriting function component is used to sense the creation input provided by the user; The processor is used to execute the image processing method according to any one of claims 1 to 16.
20. An electronic device, characterized in that, Includes a memory and a processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the image processing method according to any one of claims 1 to 16 by running the programs in the memory.