Image processing method, device, equipment and storage medium
By identifying the causes of abnormal areas in the generated images and explaining them, the problem of discrepancies between user intent and generated images is solved, and the efficiency and accuracy of image generation are improved.
Patent Information
- Application Number
- CN202510941837.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-09
AI Technical Summary
There are significant differences in the accuracy and control of user-entered prompts, which causes the generated images or videos to be inconsistent with the user's intentions, resulting in inefficient image generation.
By performing image recognition on the generated images, abnormal image areas are determined, and the natural language generation model is used to explain the causes of the abnormalities and provide modification suggestions to improve image generation efficiency.
It achieves precise positioning of abnormal image areas and explanation of causes, improves user information acquisition efficiency, and enhances the accuracy and efficiency of image generation.
Smart Images

Figure CN120451514B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to an image processing method, apparatus, device, and storage medium. Background Art
[0002] With the advancement of computer technology, various applications now offer image or video generation capabilities. Users input a prompt (a word) to generate an image or video that matches the prompt. However, due to limitations in user language expression and domain understanding, prompts entered by different users can vary significantly in accuracy and control. This can cause the generated images or videos to deviate from the user's intended image information, such as object relationships, spatial layout, and lighting logic. Unaware of the causes of these issues, users are forced to blindly and frequently modify the prompt in order to obtain the desired image or video, resulting in inefficient image generation. Summary of the Invention
[0003] The present disclosure provides an image processing method, apparatus, device, and storage medium. This solution improves the efficiency of obtaining information about a target object, thereby enabling the target object to modify its object intention information to obtain an image that better meets its expectations, thereby improving image generation efficiency.
[0004] According to one aspect of an embodiment of the present disclosure, there is provided an image processing method, the method comprising:
[0005] In response to a target object inputting object intention information, generating a first image based on the object intention information, wherein the object intention information is used to constrain conditions that the generated image needs to meet;
[0006] performing image recognition on the first image to obtain image features of the first image;
[0007] determining, based on the object intention information and the image features, at least one abnormal image region in the first image, wherein image information contained in the abnormal image region is inconsistent with the object intention information;
[0008] For any abnormal image region, causal explanation information of the abnormal image region is determined based on the object intention information and the image features, and the causal explanation information is used to describe the reason why the abnormal image region is abnormal through natural language.
[0009] According to another aspect of the present disclosure, there is provided an image processing apparatus, the apparatus comprising:
[0010] an image generating unit configured to generate a first image based on the object intention information input by the target object, wherein the object intention information is used to constrain conditions that the generated image needs to satisfy;
[0011] an image recognition unit, configured to perform image recognition on the first image to obtain image features of the first image;
[0012] a region determining unit configured to determine, based on the object intention information and the image features, at least one abnormal image region in the first image, wherein image information contained in the abnormal image region is inconsistent with the object intention information;
[0013] The information determination unit is configured to determine, for any abnormal image area, causal explanation information of the abnormal image area based on the object intention information and the image features, wherein the causal explanation information is used to describe the reason why the abnormal image area is abnormal through natural language.
[0014] In some implementations, the region determination unit is configured to parse the object intention information to obtain multiple semantic information; construct a semantic structure tree based on the multiple semantic information, the semantic structure tree includes a root node and multiple child nodes, the multiple child nodes are distributed in different layers of the semantic structure tree, the root node is used to indicate the overall features of the image to be generated, and the multiple child nodes are used to indicate local features of the image to be generated from different dimensions; based on the semantic structure tree and the image features, determine at least one abnormal image region in the first image.
[0015] In some embodiments, the area determination unit is further configured to obtain, for any image area in the first image, local features corresponding to the image area from the image features; determine the similarity between the local features of the image area and the semantic features corresponding to each node in the semantic structure tree; if the similarity between the image area and the semantic features corresponding to at least one node is less than a similarity threshold, determine that the image area is an abnormal image area.
[0016] In some embodiments, the local feature includes features of multiple visual attributes; the area determination unit is further configured to determine, for any node in the semantic structure tree, the target visual attribute corresponding to the node in the local feature; based on the feature length of the target visual attribute, encode the information contained in the node to obtain the semantic feature corresponding to the node, and the feature length of the semantic feature is consistent with the feature length of the target visual attribute.
[0017] In some embodiments, the image recognition unit is configured to recognize the first image from the perspective of multiple visual attributes to obtain multiple image feature matrices, each image feature matrix corresponding to a visual attribute; and fuse the multiple image feature matrices to obtain image features of the first image.
[0018] In some embodiments, the information determination unit is configured to construct, for any abnormal image area, an abnormal information triplet based on the object intention information and the image features, the abnormal information triplet including an object intention segment, a target local feature and an abnormality type, the object intention segment being used to indicate a condition that the abnormal image area does not meet, and the target local feature being a local feature corresponding to the abnormal image area in the image features; the abnormal information triplet is processed based on a first natural language generation model to obtain the causal explanation information of the abnormal image area.
[0019] In some embodiments, the information determination unit is further configured to process the abnormal information triples based on a second natural language generation model and display modification suggestion information, where the modification suggestion information is used to indicate how to modify the object intention information.
[0020] In some embodiments, the information determination unit is further configured to process the abnormal information triples based on a second natural language generation model, displaying at least one suggestion intention information and at least one second image, each suggestion intention information corresponding to a second image generated based on the suggestion intention information.
[0021] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, the electronic device including:
[0022] one or more processors;
[0023] a memory for storing program codes executable by the processor;
[0024] The processor is configured to execute the program code to implement the above-mentioned image processing method.
[0025] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the above-mentioned image processing method.
[0026] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned image processing method when executed by a processor.
[0027] The disclosed embodiment provides an image processing solution that, by performing image recognition on a first image generated based on object intent information, can determine the image features of the first image, and then, by comparing and analyzing the image features and the object intent information, can determine abnormal image regions in the first image that are inconsistent with the object intent information, thereby achieving regional-level abnormality location and deviation identification, and improving the accuracy of subsequent explanations of the causes of abnormalities. Finally, the cause of the abnormality in the abnormal image region is explained through natural language description, so that the target object can intuitively understand the cause of the difference between the first image and the object intent information, thereby improving the target object's information acquisition efficiency, and further enabling the target object to modify the object intent information to obtain an image that is more in line with expectations, thereby improving image generation efficiency.
[0028] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0030] Figure 1 The figure is a schematic diagram showing an implementation environment of an image processing method according to an exemplary embodiment.
[0031] Figure 2 The figure is a flowchart of an image processing method according to an exemplary embodiment.
[0032] Figure 3 is a flowchart of another image processing method according to an exemplary embodiment.
[0033] Figure 4 is a schematic diagram of a semantic structure tree provided according to an exemplary embodiment.
[0034] Figure 5 is a schematic diagram of a model architecture provided according to an exemplary embodiment.
[0035] Figure 6 is a block diagram of an image processing apparatus according to an exemplary embodiment.
[0036] Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0037] In order to enable ordinary people in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0038] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0039] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the images involved in this disclosure were obtained with full authorization.
[0040] Figure 1 FIG1 is a schematic diagram showing an implementation environment of an image processing method according to an exemplary embodiment. Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102. The terminal 101 can be connected to the server 102 via a wireless network or a wired network.
[0041] Terminal 101 can be at least one of a smartphone, a smartwatch, a desktop computer, a laptop, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), and a portable computer. Terminal 101 can have an application installed and running that generates images or videos based on user input. This application is associated with server 102, which provides backend services to terminal 101.
[0042] Terminal 101 may generally refer to one of multiple terminals. This embodiment uses terminal 101 as an example. Those skilled in the art will appreciate that the number of terminals may be greater or lesser. For example, there may be a few terminals, or dozens, hundreds, or even more. This embodiment does not limit the number or device type of terminals.
[0043] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may also include other functional servers to provide more comprehensive and diversified services. In some embodiments, the server 102 undertakes the main computing work, and the terminal 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the terminal 101 undertakes the main computing work; or, a distributed computing architecture is used between the server 102 and the terminal 101 for collaborative computing. The server 102 can be connected to the terminal 101 and other terminals via a wireless network or a wired network. Optionally, the number of the above servers may be more or less, and the embodiments of the present disclosure do not limit this.
[0044] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment. Figure 2 As shown, the method is executed by an electronic device and includes the following steps.
[0045] In step S201 , in response to a target object inputting object intention information, a first image is generated based on the object intention information.
[0046] In an embodiment of the present disclosure, the target object may be an account currently logged into the electronic device, which is used to represent the user using the electronic device. That is, the target object refers to the user who provides the object intent information. The object intent information is information input by the user that describes the specific requirements or conditions of the image that the user wants to generate. That is, the object intent information is used to constrain the conditions that the generated image must meet. The object intent information can be text, keywords, sketches or other forms that can express the user's expectations for the image. The present disclosure does not limit the specific form of the object intent information.
[0047] Electronic devices can use image generation models, such as diffusion models and GANs (Generative Adversarial Networks), to generate an image that satisfies the object's intent information. For ease of description, this is referred to as the first image. It should be noted that the first image can be a single image or a frame from a video. If the first image is a video frame, the image generation model can generate multiple images based on the object's intent information.
[0048] In step S202, image recognition is performed on the first image to obtain image features of the first image.
[0049] In the embodiment of the present disclosure, performing image recognition on the first image refers to analyzing the first image using computer vision technology to understand what content the image contains, what objects it contains, what visual attributes it has, etc.
[0050] Optionally, image recognition includes technologies such as object detection and recognition, style classification, scene classification, attribute recognition, and image segmentation. Object detection and recognition is used to identify objects in the first image and determine their positions within the image. Style classification is used to determine the image style of the first image. Scene classification is used to determine the scene of the first image. Attribute recognition is used to identify visual attributes of the first image, such as object relationships, lighting conditions, subject motion, camera movement, and text content. Image segmentation involves dividing the first image into multiple image regions.
[0051] In step S203, based on the object intention information and the image features, at least one abnormal image region in the first image is determined, where image information contained in the abnormal image region is inconsistent with the object intention information.
[0052] In the disclosed embodiments, since image features are features identified from the first image, and the first image is generated based on the object's intended information, theoretically, the image information in the first image should be consistent with the object's intended information. Image features are the actual features of the first image, and image features can be used to determine which image regions in the first image contain image information inconsistent with the object's intended information. For ease of description, inconsistent image regions are referred to as abnormal image regions. In other words, abnormal image regions are localized image regions that do not meet the intent or requirements of the target object.
[0053] It should be noted that there may be one or more abnormal image regions. The abnormality types of different abnormal image regions may be the same or different. For example, a composition abnormality may involve two or three image regions, while an object color abnormality may only involve the image region containing the object. This disclosure does not limit this.
[0054] In step S204 , for any abnormal image region, causal explanation information of the abnormal image region is determined based on the object intention information and the image features.
[0055] In the embodiment of the present disclosure, taking any abnormal image area as an example, the first image information that the target object expects to be included in the abnormal image area, that is, the intention of the target object, can be determined based on the object intention information. The second image information actually contained in the abnormal image area can be determined based on the image features. Through the difference between the second image information and the first image information of the abnormal image area, the specific cause of the abnormality in the abnormal image area is inferred, and then the above cause is converted into fluent natural language through natural language generation technology to obtain causal explanation information. That is, the causal explanation information is used to describe the cause of the abnormality in the abnormal image area through natural language.
[0056] The disclosed embodiment provides an image processing solution that, by performing image recognition on a first image generated based on object intent information, can determine the image features of the first image, and then, by comparing and analyzing the image features and the object intent information, can determine abnormal image regions in the first image that are inconsistent with the object intent information, thereby achieving regional-level abnormality location and deviation identification, and improving the accuracy of subsequent explanations of the causes of abnormalities. Finally, the cause of the abnormality in the abnormal image region is explained through natural language description, so that the target object can intuitively understand the cause of the difference between the first image and the object intent information, thereby improving the target object's information acquisition efficiency, and further enabling the target object to modify the object intent information to obtain an image that is more in line with expectations, thereby improving image generation efficiency.
[0057] In some embodiments, determining at least one abnormal image region in the first image based on the object intention information and the image features includes:
[0058] Parse the object intention information to obtain multiple semantic information;
[0059] Constructing a semantic structure tree based on multiple semantic information. The semantic structure tree includes a root node and multiple child nodes. The multiple child nodes are distributed at different layers of the semantic structure tree. The root node is used to indicate the overall features of the image to be generated, and the multiple child nodes are used to indicate the local features of the image to be generated from different dimensions.
[0060] At least one abnormal image region in the first image is determined based on the semantic structure tree and the image features.
[0061] By constructing a multi-level semantic structure tree based on the semantic information in the object intent information and then aligning the semantic structure tree with the image features, consistency detection of intent information from multiple dimensions is achieved, so that abnormal image areas can be accurately located.
[0062] In some embodiments, determining at least one abnormal image region in the first image based on the semantic structure tree and the image features includes:
[0063] For any image region in the first image, obtaining a local feature corresponding to the image region from the image features;
[0064] Determine the similarity between the local features of the image region and the semantic features corresponding to each node in the semantic structure tree;
[0065] If the similarity between the image region and the semantic feature corresponding to at least one node is less than a similarity threshold, the image region is determined to be an abnormal image region.
[0066] By performing multi-level similarity matching between image features and semantic structure trees for each image region, combined with a threshold judgment mechanism, it is possible to accurately identify image regions where local features deviate from semantic requirements, achieve fine-grained anomaly localization, and significantly improve the robustness of detection and the efficiency of verification of intent compliance.
[0067] In some embodiments, the local features include features of multiple visual attributes;
[0068] Determine the similarity between the local features of the image region and the semantic features corresponding to each node in the semantic structure tree, including:
[0069] For any node in the semantic structure tree, determine the target visual attribute corresponding to the node in the local feature;
[0070] Based on the feature length of the target visual attribute, the information contained in the node is encoded to obtain the semantic feature corresponding to the node. The feature length of the semantic feature is consistent with the feature length of the target visual attribute.
[0071] By dynamically aligning the dimensions of semantic features and visual attributes, accurate cross-modal comparison can be achieved. This not only ensures that the length of semantic encoding is consistent with the target visual attribute, avoiding dimensionality reduction loss, but also focuses on node-related visual attributes, significantly improving the accuracy of anomaly detection.
[0072] In some embodiments, performing image recognition on the first image to obtain image features of the first image includes:
[0073] Recognize the first image from the perspective of multiple visual attributes to obtain multiple image feature matrices, each image feature matrix corresponding to a visual attribute;
[0074] Multiple image feature matrices are fused to obtain image features of the first image.
[0075] By extracting feature matrices from multiple visual attributes (such as color and texture) and fusing them into a unified feature, the comprehensiveness and discriminability of the features are enhanced. Attribute extraction allows for targeted capture of diverse visual information, avoiding the limitations of a single perspective. Matrix fusion integrates multidimensional information, improving the robustness and expressiveness of features and facilitating the accuracy of subsequent image tasks.
[0076] In some embodiments, for any abnormal image region, determining causal explanation information of the abnormal image region based on object intention information and image features includes:
[0077] For any abnormal image region, an abnormal information triplet is constructed based on the object intent information and image features. The abnormal information triplet includes the object intent segment, the target local feature, and the abnormality type. The object intent segment is used to indicate the conditions that the abnormal image region does not meet, and the target local feature is the local feature corresponding to the abnormal image region in the image features.
[0078] The abnormal information triples are processed based on the first natural language generation model to obtain causal explanation information of the abnormal image area.
[0079] By constructing anomaly information triples (object intent fragment - target local features - anomaly type) and leveraging natural language generation technology, we achieve interpretable output of anomaly detection results. Specifically, structured triples clearly associate intent requirements, actual features, and anomaly types, ensuring clear causal logic in the explanation. Cross-modal alignment of semantic intent (text) and visual features (data) enhances the objectivity of the explanation. Finally, by converting technical detection results into natural language descriptions, the results are significantly more understandable, making them suitable for non-expert users to quickly locate the root cause of the problem and improving the efficiency of human-computer interaction.
[0080] In some embodiments, the method further comprises:
[0081] The abnormal information triples are processed based on the second natural language generation model, and modification suggestion information is displayed. The modification suggestion information is used to indicate how to modify the object intention information.
[0082] A second natural language generation model analyzes abnormal triples and generates targeted modification suggestions, accurately guiding intent optimization, such as adjusting description granularity or adding constraints. Using natural language descriptions reduces the cost of human-computer communication and significantly improves human-computer interaction efficiency.
[0083] In some embodiments, the method further comprises:
[0084] The abnormal information triples are processed based on the second natural language generation model to display at least one suggestion intention information and at least one second image, where each suggestion intention information corresponds to a second image generated based on the suggestion intention information.
[0085] By intelligently generating pairings between suggested intents and example images, we achieve the effect of simultaneously displaying text suggestions and corresponding images, allowing the target audience to intuitively compare and make decisions. Furthermore, the target audience can infer semantic expressions from example images, lowering the threshold for description. Furthermore, the target audience can try to adjust the intent and verify the results, significantly improving image generation efficiency.
[0086] above Figure 2 The figure shows a flow chart of an image processing method of the present disclosure. The image processing solution provided by the present disclosure is further explained below. Figure 3 is a flowchart of another image processing method according to an exemplary embodiment. Figure 3 The method is executed by an electronic device and includes the following steps.
[0087] In step S301 , in response to a target object inputting object intention information, a first image is generated based on the object intention information.
[0088] In the embodiment of the present disclosure, the target object is the subject that initiates the image generation request. Optionally, the target object can be the account currently logged in to the electronic device, and the target object is used to represent the user using the electronic device. The target object can express the image generation requirements in the form of structured or natural language, that is, the object intent information. For example, the object intent information is: a Samoyed wearing a red hat is running in the snow, and the overall style is anime style. The object intent information includes the constraints on the image to be generated, that is, what objects need to be included in the image, what attributes the objects have, the scene in which the objects are located, the spatial relationship, the overall style of the image, etc. In other words, the object intent information is used to constrain the conditions that need to be met for the generated image.
[0089] Among them, the first image can be generated by an image generation model, such as a diffusion model, GAN, etc., and the embodiment of the present disclosure does not limit the image generation model. It should be noted that the image generation model takes the object intention information as a conditional input, and generates an image based on the content contained in the object intention information, rather than randomly generating or generating based on a preset template. Optionally, taking the image generation model as an example of a diffusion model, the diffusion model encodes the object intention information through an encoder and converts the object intention information into a high-dimensional semantic vector. Then, in each denoising process, the diffusion model constrains the evolution direction of the image through the above-mentioned high-dimensional semantic vector to ensure that the generated image includes the above-mentioned image elements such as the red hat, Samoyed and snow.
[0090] In step S302, the object intention information is parsed to obtain multiple semantic information.
[0091] In the disclosed embodiments, object intent information typically includes requirements across multiple dimensions. By parsing the object intent information, structured or quantified semantic information can be obtained. Optionally, the semantic information includes explicit or implicit information such as composition goal, style, object relationships, lighting conditions, subject motion, camera shot, and text content.
[0092] The composition objective refers to global information such as the overall layout of the image, the subject's position, and the viewing angle. For example, "sitting by the window" indicates the subject's position is centered or slightly to the left. "Looking at the wind chimes on the eaves" indicates the subject's gaze direction.
[0093] Style refers to the visual expression, artistic genre, media texture, era characteristics, cultural atmosphere, or the iconic techniques of a specific creator / work that the user expects to generate.
[0094] For example, if the object intent information clearly refers to art genres such as cyberpunk and impressionism, or if the object intent information includes descriptions such as distinct brushstrokes and colorful light and shadow, it indicates impressionism. If the object intent information includes descriptions such as steam, technology, and the future, it indicates cyberpunk. Another example is if the object intent information clearly refers to the texture of media such as watercolor, oil painting, and pencil drawing. Or if the object intent information includes descriptions such as blurred color edges, it indicates watercolor. If the object intent information includes thick brushstrokes and a sheen, it indicates oil painting. Another example is if the object intent information clearly refers to keywords that are characteristic of an era and culture, such as ancient Greek murals and Dunhuang murals. Or if the object intent information includes tranquil landscapes and white space, it indicates Chinese ink painting. Similarly, object intent information can also include photographic terms such as long exposure, wide-angle lens, and Wong Kar-wai's color palette.
[0095] Object relationships refer to the spatial, logical, or interactive relationships between different objects in an image. For example, a dog pulling a sled reflects the positional relationship between the dog and the sled. A cat wearing a red hat reflects the clothing relationship between the hat and the cat.
[0096] Lighting and shadow conditions refer to the lighting effects, distribution of light and dark, and direction of the light source in an image. For example, warm light sources such as the warm-toned evening sun or wet bluestone indicate a reflective floor.
[0097] The subject's motion describes the action or posture of an object in the image. For example, "dog sledding" depicts a dog pulling a sled. "Watching the rain" depicts a person standing still and gazing in a certain direction.
[0098] Camera movement refers to the imitation of photography or film shooting techniques. For example, the cinematic effect and telephoto effect are the effects of using a specific shooting method.
[0099] Text content refers to the text elements contained in an image. For example, the object intent information describes that an image contains a signboard of a stationery store, and the text on the signboard is the above text content.
[0100] Optionally, the electronic device can parse the object intent information using a pre-trained large language model and output structured multiple semantic information. For example: {"composition": "central symmetry", "style": "cyberpunk", "subject": {"dog": {"location": "snow", "color": "metallic gray"}}}}. The disclosed embodiments do not limit the form of the multiple semantic information.
[0101] In step S303, a semantic structure tree is constructed based on multiple semantic information.
[0102] In the embodiment of the present disclosure, a semantic structure tree is an attribute data structure. By constructing a semantic structure tree, multiple semantic information included in the object intent information can be hierarchically organized. The semantic structure tree includes a root node and multiple child nodes. The root node is an equal-level node of the semantic structure tree, which is used to represent the global constraints of the image, such as the overall style and core objectives. The multiple child nodes are distributed at different levels of the semantic structure tree, and the multiple child nodes are used to indicate the local features of the image to be generated from different dimensions.
[0103] Optionally, the semantic information representing the global style or core goal is used as the root node, the main object is used as the first-level child node, the object attributes are used as the second-level child nodes, the spatial relationship is used as the associated node, and the dynamic effect is used as the leaf node.
[0104] For example, see Figure 4As shown, let's take the generation of a cyberpunk-themed street scene as an example. The root node is the global feature. The first-level child nodes include composition, style, subject, light and shadow, and dynamics. The second-level child nodes of composition include perspective as an upward-looking lens and layout as central symmetry. The second-level child nodes of style include artistic genre as cyberpunk, medium as 3D rendering, and glitch art. Glitch art refers to special effects that distort data, such as intermittently flashing mosaic effects. The second-level child node of subject includes the core object as a cyborg. The environment is a neon street. The second-level child node of light and shadow includes the main light source as a holographic billboard and the reflected light source as a wet opposite. The second-level child node of dynamics includes moving, suspended traffic. The cyborg also includes leaf nodes to further define the cyborg, such as its attributes as having a mechanical prosthetic eye, its action as standing with an umbrella, and its standing position as the center of the neon street. The neon street also includes leaf nodes to further define the neon street, such as including X-speaking signs and holographic projections to embody the cyberpunk theme. Similarly, the suspended traffic flow is not static. To reflect this, the leaf nodes are used to limit the suspended traffic flow to have a red light trailing to indicate that the suspended traffic flow is moving at high speed.
[0105] By constructing a semantic structure tree, image elements can be hierarchically organized, significantly improving image controllability and interpretability. This structure tree allows for associating anomalous image regions with subnodes within the structure tree, enabling rapid identification of the source of deviations. Subnodes within the structure tree support partial replacement, eliminating the need to completely rewrite object intent information and replacing only a portion of the prompt word, improving efficiency.
[0106] In step S304, image recognition is performed on the first image to obtain image features of the first image.
[0107] In an embodiment of the present disclosure, the electronic device may perform image recognition on the first image using a pre-trained model to obtain multi-level features and thereby obtain image features. The pre-trained model may be a convolutional neural network or a ViT (Vision Transformer) model. Optionally, the image features include low-level features such as pixel-level texture features, mid-level features such as local structural features, high-level features such as object semantic features, and global features such as overall style features.
[0108] In some embodiments, image features include features identified from the perspective of multiple visual attributes. Accordingly, the electronic device identifies the first image from the perspective of multiple visual attributes to obtain multiple image feature matrices, each corresponding to a visual attribute. The electronic device fuses the multiple image feature matrices to obtain image features for the first image. Feature matrices are extracted separately from multiple visual attributes (such as color and texture) and then fused into a unified feature, enhancing the comprehensiveness and discriminability of the features. Attribute extraction can capture different visual information in a targeted manner, avoiding the limitations of a single perspective; matrix fusion integrates multidimensional information, improving the robustness and expressiveness of features, and facilitating the accuracy of subsequent image tasks.
[0109] For example, the electronic device calls a pre-trained model such as DinoV2 or VAE to parse the first image from four orthogonal dimensions. Among them, the electronic device can capture the position and category of the object in the first image from the object distribution dimension. The electronic device can encode the properties of the object surface from the texture detail dimension, such as the reflectivity of metal. The electronic device can infer the light source parameters from the light and shadow direction dimension, such as the azimuth angle of the main light source is 120°. The electronic device can also extract visual guide lines from the composition structure dimension. Then, the electronic device outputs image feature matrices with different resolutions, and each image feature matrix corresponds to a dimension, that is, a visual attribute. Then, multiple image feature matrices are spatially aligned by uniform downsampling. Finally, the electronic device splices the above multiple image feature matrices in the channel dimension to obtain image features.
[0110] In step S305 , for any image region in the first image, local features corresponding to the image region are obtained from the image features.
[0111] In an embodiment of the present disclosure, the electronic device divides the first image into multiple image regions. Optionally, the electronic device divides the first image into multiple image regions according to the bounding boxes of the detected targets through target detection, with each image region corresponding to a bounding box. Alternatively, the electronic device divides the first image into pixel-level regions through semantic segmentation. Alternatively, the electronic device divides the first image into multiple regular image regions using a sliding window. The present embodiment of the disclosure does not limit the division method.
[0112] It should be noted that the local features corresponding to the image region refer to the features that can be identified in the image region. Optionally, the electronic device can crop the activation map corresponding to the image region from the feature map of the convolution layer to obtain the local features. Alternatively, the electronic device can perform image recognition on each image region in the first image separately to obtain the local features corresponding to each image region. By splicing the local features corresponding to each image region, the image features of the first image can be obtained.
[0113] In step S306, the similarity between the local features of the image region and the semantic features corresponding to each node in the semantic structure tree is determined.
[0114] In an embodiment of the present disclosure, an electronic device may encode the semantic information corresponding to each node in a semantic structure tree to obtain semantic features for each node. Optionally, for any node in the semantic structure tree, the electronic device may convert the semantic information of the node into a semantic feature vector using a language model. Optionally, when extracting the semantic features of a node, this structural information may be incorporated into the feature vector of the node based on the node's position and relationship in the semantic structure tree (e.g., the node's parent and child nodes).
[0115] The electronic device can use a cross-comparison method to compare the local features of each image region with the semantic features of each node one by one to calculate a similarity score between them. A higher similarity score indicates greater similarity, and a lower similarity score indicates lower similarity. Alternatively, similarity can be calculated using cosine similarity or dot product. This step is illustrated using an image region as an example.
[0116] In some embodiments, the local features include features of multiple visual attributes. Accordingly, the similarity between the local features of the image area and the semantic features corresponding to each node in the semantic structure tree is determined, including: for any node in the semantic structure tree, the electronic device determines the target visual attribute corresponding to the node in the local features. Then, based on the feature length of the target visual attribute, the electronic device encodes the information contained in the node to obtain the semantic feature corresponding to the node, and the feature length of the semantic feature is consistent with the feature length of the target visual attribute. By dynamically aligning the dimensions of the semantic features and the visual attributes, accurate cross-modal comparison is achieved, which can ensure that the semantic encoding is consistent with the length of the target visual attribute, avoiding dimensionality reduction loss, and focusing on the visual attributes related to the node, significantly improving the accuracy of anomaly detection.
[0117] For example, the local features of an image region are extracted from the image region and contain feature vectors of multi-dimensional visual attributes, such as color distribution, texture, shape contours, and depth features. Taking a node as an example, if the node represents the sky, the corresponding target visual attribute can be texture. Accordingly, the electronic device encodes the semantic information of the node through a semantic encoder to obtain a semantic feature with the same feature length as the texture feature. The electronic device can then calculate the cosine similarity between the texture feature and the semantic feature to determine the similarity between the texture feature and the semantic feature.
[0118] In step S307 , if the similarity between the image region and the semantic feature corresponding to at least one node is less than a similarity threshold, the image region is determined to be an abnormal image region.
[0119] In an embodiment of the present disclosure, the electronic device can determine whether an image region is an abnormal image region based on a pre-set similarity threshold. If the similarity between the image region and any node is lower than the similarity threshold, the image region is considered abnormal, meaning that the visual content of the region does not fully match at least one semantic concept that it is expected to conform to in the semantic tree. Optionally, the electronic device can record at least one node whose similarity to the abnormal image region is lower than the similarity threshold, to facilitate explanation of the cause of the abnormality.
[0120] This rigorous and sensitive judgment method prioritizes detecting any possible semantic inconsistencies, making it particularly suitable for high-precision anomaly detection scenarios.
[0121] Optionally, the electronic device may further detect physical inconsistencies, such as light source direction conflicts, depth of field inversion, etc., to assist in filtering possible false alarms.
[0122] In step S308, for any abnormal image region, an abnormal information triplet is constructed based on the object intention information and the image features.
[0123] In the disclosed embodiment, for any abnormal image region, the electronic device can construct an abnormal information triplet for the abnormal image region, thereby generating interpretable and actionable information for the abnormal image region through the abnormal information triplet. Optionally, the abnormal information triplet includes an object intent fragment, a target local feature, and an abnormality type. The format of the abnormal information triplet can be <object intent fragment, target local feature, abnormality type>. For example, <"ground support", [0.21, -1.32, ..., 0.78], "gravity anomaly">.
[0124] The object intent segment indicates the conditions that the abnormal image region does not meet, and the target local features are the local features of the image features corresponding to the abnormal image region. The object intent segment can be used to guide the modification of object intent information, such as enhancing the description of the ground projection. The target local features can be used to assist in locating the feature dimensions that require modification. The anomaly type can be used to determine the repair strategy.
[0125] In step S309 , the abnormal information triples are processed based on the first natural language generation model to obtain causal explanation information of the abnormal image region.
[0126] In the disclosed embodiment, the first natural language generation model extracts semantic constraints from the object intent fragment. The first natural language generation model can restore visual facts through the target local features. For example, through feature vector decoding, it is concluded that there are no shadow pixels at the bottom of the stone. The first natural language generation model activates the domain knowledge base through the anomaly type. For example, when gravity is abnormal, the first natural language generation model can associate the physical knowledge "unsupported objects should fall." The first natural language generation model maps the above-mentioned target local features to the semantic space, performs reasoning, and generates causal explanation information based on the result of reasoning.
[0127] For example, if the object intent segment includes the constraint of ground projection, the expected effect is that the bottom of the stone should have a shadow. If the target's local features indicate a lack of shadow, this indicates a violation of the physical rule that objects cast shadows when illuminated by a light source. This results in the stone having no shadow at the bottom, appearing to be suspended against gravity. In this case, the first natural language generation model can output an explanation: no projection features matching the light source direction were detected in this stone area, violating the physical rule that illuminated objects must cast shadows, and therefore determining a gravity anomaly.
[0128] It should be noted that the first natural language generation model is a generative AI model specially designed to convert structured data into natural language interpretations that conform to causal logic. Optionally, the first natural language generation model includes a cross-modal embedding module for aligning visual feature vectors with text semantics. The first natural language generation model also includes an indication prompt injection module for loading preset rule libraries according to the exception type, such as physical rule libraries, material rule libraries, etc., to ensure the credibility of the reasoning results. In addition, rapid adaptation to new scenarios can be achieved by adding and deleting knowledge bases. Optionally, the first natural language generation model can be a causal LLM (Large Language Model), etc., which is not limited in this disclosure.
[0129] In step S310, the abnormal information triple is processed based on the second natural language generation model, and modification suggestion information is displayed, where the modification suggestion information is used to indicate how to modify the object intention information.
[0130] In an embodiment of the present disclosure, the electronic device can convert the abnormality diagnosis result, i.e., the abnormality information triple, into an executable modification suggestion based on a second natural language generation model. Optionally, the second natural language generation model can map the abnormality type to a preset modification rule library, determine how the abnormality type should be modified, and then infer a text description that meets the conditions based on the target local features. Finally, the modification suggestion information is output in the form of a natural language description.
[0131] In some embodiments, the electronic device can also display a corresponding image for the suggestion intention information. Accordingly, the electronic device processes the abnormal information triples based on the second natural language generation model, and displays at least one suggestion intention information and at least one second image. Each suggestion intention information corresponds to a second image generated based on the suggestion intention information. By intelligently generating a pairing scheme for suggestion intentions and example images, the effect of synchronously displaying text suggestions and corresponding images is achieved, allowing the target object to make intuitive comparisons and decisions. In addition, the target object can also reversely infer semantic expressions through example images, lowering the description threshold of the target object. Moreover, the target object can try to adjust the intention and verify the effect, which significantly improves the efficiency of image generation.
[0132] It should be noted that while the embodiments of this disclosure illustrate the case where the first and second natural language generation models are different models, a third natural language generation model can also be used to simultaneously implement the functions of the first and second natural language generation models. That is, after inputting the abnormal information triples into the third natural language generation model, the third natural language generation model can output both causal explanation information for the abnormal image region and modification suggestion information. This embodiment of this disclosure does not impose any restrictions on this.
[0133] It should be noted that in order to make the image processing solution provided by the present disclosure easier to understand, see Figure 5 As shown, Figure 5 This is a schematic diagram of a model architecture provided according to an embodiment of the present disclosure. Figure 5 As shown, the architecture of this solution includes an image generation model 501, a semantic extraction module 502, an image recognition module 503, an anomaly determination module 504, a triple construction module 505, a first natural language generation model 506, and a second natural language generation model 507. The image generation model 501 takes object intent information as input and outputs a first image. The semantic extraction module 502 takes object intent information as input and outputs a semantic structure tree. Optionally, the semantic extraction module 502 is integrated into the image generation model 501. The image recognition module 503 takes the first image as input and outputs image features. Optionally, the image recognition module 503 is integrated into the image generation model 501. The anomaly determination module 504 takes the semantic structure tree and image features as input and outputs an abnormal image region. The triple construction module 505 takes the abnormal image region, object intent information, and image features as input and outputs an abnormal information triple. The first natural language generation model 506 takes the abnormal information triple as input and outputs causal explanation information. The input of the second natural language generation model 507 is the abnormal information triple, and the output is the modification suggestion information. Optionally, the abnormality determination module 504 is also integrated into the image generation model 501.
[0134] The disclosed embodiment provides an image processing solution that, by performing image recognition on a first image generated based on object intent information, can determine the image features of the first image, and then, by comparing and analyzing the image features and the object intent information, can determine abnormal image regions in the first image that are inconsistent with the object intent information, thereby achieving regional-level abnormality location and deviation identification, and improving the accuracy of subsequent explanations of the causes of abnormalities. Finally, the cause of the abnormality in the abnormal image region is explained through natural language description, so that the target object can intuitively understand the cause of the difference between the first image and the object intent information, thereby improving the target object's information acquisition efficiency, and further enabling the target object to modify the object intent information to obtain an image that is more in line with expectations, thereby improving image generation efficiency.
[0135] Figure 6 FIG. 1 is a block diagram of an image processing apparatus according to an exemplary embodiment. Figure 6 As shown, the apparatus includes: an image generating unit 601 , an image recognizing unit 602 , a region determining unit 603 and an information determining unit 604 .
[0136] An image generating unit 601 is configured to generate a first image based on the object intention information input by the target object, wherein the object intention information is used to constrain conditions that the generated image needs to meet;
[0137] The image recognition unit 602 is configured to perform image recognition on the first image to obtain image features of the first image;
[0138] A region determining unit 603 is configured to determine at least one abnormal image region in the first image based on the object intention information and the image features, where image information included in the abnormal image region is inconsistent with the object intention information;
[0139] The information determination unit 604 is configured to determine the causal explanation information of any abnormal image area based on the object intention information and image features, and the causal explanation information is used to describe the reason why the abnormal image area is abnormal through natural language.
[0140] In some implementations, the region determination unit 603 is configured to parse the object intention information to obtain multiple semantic information; construct a semantic structure tree based on the multiple semantic information, the semantic structure tree includes a root node and multiple child nodes, the multiple child nodes are distributed in different layers of the semantic structure tree, the root node is used to indicate the overall features of the image to be generated, and the multiple child nodes are used to indicate the local features of the image to be generated from different dimensions; based on the semantic structure tree and the image features, determine at least one abnormal image region in the first image.
[0141] In some embodiments, the region determination unit 603 is further configured to obtain, for any image region in the first image, local features corresponding to the image region from the image features; determine the similarity between the local features of the image region and the semantic features corresponding to each node in the semantic structure tree; if the similarity between the image region and the semantic features corresponding to at least one node is less than a similarity threshold, determine that the image region is an abnormal image region.
[0142] In some embodiments, the local feature includes features of multiple visual attributes; the area determination unit 603 is also configured to determine the target visual attribute corresponding to the node in the local feature for any node in the semantic structure tree; based on the feature length of the target visual attribute, the information contained in the node is encoded to obtain the semantic feature corresponding to the node, and the feature length of the semantic feature is consistent with the feature length of the target visual attribute.
[0143] In some embodiments, the image recognition unit 602 is configured to recognize the first image from the perspective of multiple visual attributes to obtain multiple image feature matrices, each image feature matrix corresponding to a visual attribute; and fuse the multiple image feature matrices to obtain the image features of the first image.
[0144] In some embodiments, the information determination unit 604 is configured to construct an abnormal information triple for any abnormal image area based on the object intention information and the image features. The abnormal information triple includes an object intention segment, a target local feature, and an abnormality type. The object intention segment is used to indicate the conditions that the abnormal image area does not meet, and the target local feature is the local feature corresponding to the abnormal image area in the image feature; the abnormal information triple is processed based on the first natural language generation model to obtain causal explanation information of the abnormal image area.
[0145] In some embodiments, the information determination unit 604 is further configured to process the abnormal information triples based on the second natural language generation model and display modification suggestion information, where the modification suggestion information is used to indicate how to modify the object intention information.
[0146] In some embodiments, the information determination unit 604 is further configured to process the abnormal information triples based on the second natural language generation model, displaying at least one suggestion intention information and at least one second image, each suggestion intention information corresponding to a second image generated based on the suggestion intention information.
[0147] The disclosed embodiment provides an image processing device that, by performing image recognition on a first image generated based on object intent information, can determine the image features of the first image, and then, by comparing and analyzing the image features with the object intent information, can determine abnormal image regions in the first image that are inconsistent with the object intent information, thereby achieving regional-level abnormality location and deviation identification, and improving the accuracy of subsequent explanations of the causes of abnormalities. Finally, the cause of the abnormality in the abnormal image region is explained through natural language description, so that the target object can intuitively understand the cause of the difference between the first image and the object intent information, thereby improving the target object's information acquisition efficiency, and further enabling the target object to modify the object intent information to obtain an image that is more in line with expectations, thereby improving image generation efficiency.
[0148] It should be noted that the image processing device provided in the above embodiment is merely an example of the division of the above functional units. In actual applications, the above functions can be assigned to different functional units as needed, that is, the internal structure of the electronic device can be divided into different functional units to complete all or part of the functions described above. In addition, the image processing device provided in the above embodiment and the image processing method embodiment are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0149] Regarding the image processing apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.
[0150] In the embodiments of the present disclosure, the electronic device may be a terminal or a server. When the electronic device is a terminal, the terminal serves as the execution subject to implement the technical solutions provided in the embodiments of the present disclosure; when the electronic device is a server, the server serves as the execution subject to implement the technical solutions provided in the embodiments of the present disclosure; or, alternatively, the technical solutions provided in the present disclosure may be implemented through interaction between the terminal and the server. The embodiments of the present disclosure are not limited in this regard.
[0151] Figure 7 7 is a block diagram of an electronic device according to an exemplary embodiment. Generally, the electronic device 700 includes a processor 701 and a memory 702 .
[0152] Processor 701 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 701 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 701 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 701 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content displayed on the display screen. In some embodiments, processor 701 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0153] The memory 702 may include one or more computer-readable storage media, which may be non-transitory. The memory 702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 702 is used to store at least one program code, which is executed by the processor 701 to implement the image processing method provided in the method embodiment of the present disclosure.
[0154] In some embodiments, electronic device 700 may optionally include a peripheral device interface 703 and at least one peripheral device. Processor 701, memory 702, and peripheral device interface 703 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 703 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 704, a display screen 705, a camera assembly 706, an audio circuit 707, and a power supply 708.
[0155] The peripheral device interface 703 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 701 and the memory 702. In some embodiments, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board. In other embodiments, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0156] The RF circuit 704 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 704 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 704 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 704 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 704 can communicate with other electronic devices via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 704 may also include circuitry related to Near Field Communication (NFC), although this disclosure does not limit this.
[0157] Display screen 705 is used to display a user interface (UI). This UI can include graphics, text, icons, videos, or any combination thereof. If display screen 705 is a touchscreen display, it can also detect touch signals on or above the surface of display screen 705. These touch signals can be input as control signals to processor 701 for processing. Display screen 705 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 705, located on the front panel of electronic device 700. In other embodiments, there can be at least two display screens 705, located on different surfaces of electronic device 700 or in a foldable design. In still other embodiments, display screen 705 can be a flexible display, located on a curved or foldable surface of electronic device 700. Display screen 705 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 705 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0158] The camera assembly 706 is used to capture images or videos. Optionally, the camera assembly 706 includes a front camera and a rear camera. Typically, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 706 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0159] The audio circuit 707 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals to be input into the processor 701 for processing, or to be input into the radio frequency circuit 704 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there can be multiple microphones, which are respectively arranged in different parts of the electronic device 700. The microphone can also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 701 or the radio frequency circuit 704 into sound waves. The speaker can be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signals into sound waves audible to humans, but also convert the electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 707 may also include a headphone jack.
[0160] Power supply 708 is used to power the various components of electronic device 700. Power supply 708 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 708 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0161] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation on the electronic device 700, and the electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0162] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 702 including instructions. The instructions can be executed by a processor 701 of an electronic device 700 to perform the above-described image processing method. Alternatively, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0163] A computer program product includes a computer program, which implements the above-mentioned image processing method when executed by a processor.
[0164] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0165] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: In response to a target object inputting object intention information, generating a first image based on the object intention information, wherein the object intention information is used to constrain conditions that the generated image needs to meet; performing image recognition on the first image to obtain image features of the first image; Parsing the object intention information to obtain multiple semantic information; Constructing a semantic structure tree based on the multiple semantic information, the semantic structure tree including a root node and multiple child nodes, the multiple child nodes being distributed at different layers of the semantic structure tree, the root node being used to indicate overall features of the image to be generated, and the multiple child nodes being used to indicate local features of the image to be generated from different dimensions; determining, based on the semantic structure tree and the image features, at least one abnormal image region in the first image, where image information contained in the abnormal image region is inconsistent with the object intent information; For any abnormal image region, causal explanation information of the abnormal image region is determined based on the object intention information and the image features, and the causal explanation information is used to describe the reason why the abnormal image region is abnormal through natural language.
2. The image processing method according to claim 1, wherein: The determining, based on the semantic structure tree and the image features, at least one abnormal image region in the first image includes: For any image region in the first image, obtaining a local feature corresponding to the image region from the image features; Determining the similarity between the local features of the image region and the semantic features corresponding to each node in the semantic structure tree; If the similarity between the image region and the semantic feature corresponding to at least one node is less than a similarity threshold, the image region is determined to be an abnormal image region.
3. The image processing method according to claim 2, wherein: The local features include features of multiple visual attributes; Determining the similarity between the local feature of the image region and the semantic feature corresponding to each node in the semantic structure tree includes: For any node in the semantic structure tree, determining a target visual attribute corresponding to the node in the local feature; Based on the feature length of the target visual attribute, the information contained in the node is encoded to obtain a semantic feature corresponding to the node, and the feature length of the semantic feature is consistent with the feature length of the target visual attribute.
4. The image processing method according to any one of claims 1 to 3, characterized in that: The performing image recognition on the first image to obtain image features of the first image includes: Recognize the first image from the perspective of multiple visual attributes to obtain multiple image feature matrices, each image feature matrix corresponding to a visual attribute; The multiple image feature matrices are fused to obtain image features of the first image.
5. The image processing method according to any one of claims 1 to 3, characterized in that: The step of determining, for any abnormal image region, causal explanation information of the abnormal image region based on the object intention information and the image features includes: For any abnormal image region, construct an abnormal information triplet based on the object intent information and the image features, wherein the abnormal information triplet includes an object intent segment, a target local feature, and an abnormality type. The object intent segment is used to indicate a condition that the abnormal image region does not meet, and the target local feature is a local feature corresponding to the abnormal image region in the image features; The abnormal information triples are processed based on a first natural language generation model to obtain the causal explanation information of the abnormal image area.
6. The image processing method according to claim 5, characterized in that The method further comprises: The abnormal information triple is processed based on a second natural language generation model, and modification suggestion information is displayed, where the modification suggestion information is used to indicate how to modify the object intention information.
7. The image processing method according to claim 5, characterized in that: The method further comprises: The abnormal information triple is processed based on the second natural language generation model to display at least one suggestion intention information and at least one second image, each suggestion intention information corresponds to a second image generated based on the suggestion intention information.
8. An image processing device, characterized in that: The device comprises: an image generating unit configured to generate a first image based on the object intention information input by the target object, wherein the object intention information is used to constrain conditions that the generated image needs to satisfy; an image recognition unit, configured to perform image recognition on the first image to obtain image features of the first image; A region determination unit is configured to parse the object intent information to obtain multiple semantic information; construct a semantic structure tree based on the multiple semantic information, the semantic structure tree including a root node and multiple child nodes, the multiple child nodes being distributed at different layers of the semantic structure tree, the root node being used to indicate the overall features of the image to be generated, and the multiple child nodes being used to indicate local features of the image to be generated from different dimensions; and determine at least one abnormal image region in the first image based on the semantic structure tree and the image features, wherein image information contained in the abnormal image region is inconsistent with the object intent information; The information determination unit is configured to determine, for any abnormal image area, causal explanation information of the abnormal image area based on the object intention information and the image features, wherein the causal explanation information is used to describe the reason why the abnormal image area is abnormal through natural language.
9. An electronic device, characterized in that: The electronic device comprises: one or more processors; a memory for storing program code executable by the processor; The processor is configured to execute the program code to implement the image processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the image processing method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the image processing method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image-text cross-modal retrieval method and system based on text tree local matching
CN114048282A
Abnormal recognition method and device, storage medium and electronic equipment
CN117788918A