Image processing method and related equipment
By combining the cultural graph model with a large multimodal model, image verification and optimization are performed, which solves the problem of large fluctuations and strong randomness in the quality of images generated by the cultural graph model, and improves the quality of image generation and user experience.
Patent Information
- Application Number
- CN202510820330.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
The image quality generated by the existing text-based image model fluctuates greatly and is highly random, making it difficult to accurately generate images that match user prompt information, resulting in a poor user experience.
By obtaining prompt information, the initial image is generated using the text-based image model, and verification is performed based on the multimodal large model, including verification of feature information such as the subject, environment, style, and emotion. A layer-by-layer progressive analysis is performed to determine the target image. If the verification passes, it is determined as the final image.
The quality of images output by the cultural graph model and user experience have been improved, the efficiency and accuracy of image generation have been improved, and the relevance of images to user prompt information has been ensured.
Smart Images

Figure CN120707702A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to an image processing method, an image processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In related technologies, the Wensheng graph model can generate images based on user prompts. For simple prompts, the generated images are high-quality and highly relevant to the prompts. However, the images generated by the Wensheng graph model suffer from large quality fluctuations and high randomness.
[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the Invention
[0004] The purpose of the present disclosure is to provide an image processing method and related equipment, which can at least to some extent overcome the problems of large quality fluctuations and strong randomness of images generated by Chinese raw image models in related technologies.
[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by practice of the present disclosure.
[0006] According to one aspect of the present disclosure, an image processing method is provided, comprising: obtaining prompt information; generating an initial image based on a text graph model and the prompt information; processing the prompt information based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; and verifying the initial image using the multimodal large model based on the feature information of each verification type, and determining the initial image as a target image if the verification passes.
[0007] In one embodiment of the present disclosure, the initial image is verified by the multimodal large model based on the feature information of each verification type. If the verification passes, the initial image is determined as the target image, including: if the initial image is clear; the target in the initial image is not deformed, distorted, and does not violate the laws of nature; and / or the initial image matches the feature information of at least one verification type in the prompt information, then the verification passes.
[0008] In one embodiment of the present disclosure, the verification type includes at least one of subject, environment, style, and emotion.
[0009] In one embodiment of the present disclosure, the method further includes: if the verification fails, storing the initial image and the verification result in an image acquisition list; if the number of iterations does not reach a preset threshold, adjusting the prompt information according to the verification result, generating an image and verifying the image; if the verification fails, updating the image acquisition list; if the verification passes, obtaining the target image; if the number of iterations reaches a preset threshold, determining the target image in the image acquisition list.
[0010] In one embodiment of the present disclosure, the verification result includes the number of verification rounds passed; wherein, determining the target image in the image acquisition list includes: calculating a correlation score based on the correlation between the image in the image acquisition list and the corresponding prompt information; and determining the target image based on the correlation score and the number of verification rounds passed.
[0011] In one embodiment of the present disclosure, determining the target image based on the correlation score and the number of times the image has passed the verification rounds includes: screening a preset number of images with the highest correlation scores from the image acquisition list; determining the image with the largest number of times the image has passed the verification rounds as the target image, or when the preset number of images have the same number of times the image has passed the verification rounds, determining the image with the highest correlation score as the target image.
[0012] According to another aspect of the present disclosure, an image processing device is provided, comprising: an acquisition module for acquiring prompt information; a generation module for generating an initial image based on a text graph model and the prompt information; an extraction module for performing feature extraction on the prompt information based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; and a verification module for verifying the initial image using the multimodal large model based on the feature information of the at least one verification type, and determining the initial image as a target image if the verification passes.
[0013] According to another aspect of the present disclosure, an electronic device is provided, including: a processor; and a memory for storing executable instructions of the processor; the processor is configured to perform the above-mentioned image processing method by executing the executable instructions.
[0014] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the image processing method described above is implemented.
[0015] According to another aspect of the present disclosure, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the image processing method described above is implemented.
[0016] In an embodiment of the present disclosure, prompt information is obtained; an initial image is generated based on a Wensheng graph model and the prompt information; the prompt information is processed based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; based on the feature information of each verification type, the initial image is verified using the multimodal large model; if the verification passes, the initial image is determined as the target image. The present disclosure can improve the quality of the image output by the Wensheng graph model and enhance the user experience.
[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification, are used to explain the principles of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0019] Figure 1 A schematic diagram illustrating an exemplary application system architecture of an image processing method in an embodiment of the present disclosure is shown.
[0020] Figure 2 A flow chart of an image processing method in an embodiment of the present disclosure is shown.
[0021] Figure 3 A flowchart of another image processing method in an embodiment of the present disclosure is shown.
[0022] Figure 4 A flow chart of a method for determining a target image in an embodiment of the present disclosure is shown.
[0023] Figure 5 A flow chart of another method for determining a target image in an embodiment of the present disclosure is shown.
[0024] Figure 6 A flowchart of an example 1 of an image processing method in an embodiment of the present disclosure is shown.
[0025] Figure 7 A flowchart of Example 2 of an image processing method in an embodiment of the present disclosure is shown.
[0026] Figure 8 A schematic diagram showing an initial image obtained in an embodiment of the present disclosure.
[0027] Figure 9 A schematic diagram showing another initial image obtained in an embodiment of the present disclosure.
[0028] Figure 10 A schematic structural diagram of an image processing device in an embodiment of the present disclosure is shown.
[0029] Figure 11 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0030] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0031] In addition, the accompanying drawings are merely schematic illustrations of the present disclosure and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0032] Figure 1 FIG. 1 shows an exemplary application system architecture diagram to which the image processing method in the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture may include a terminal device 101 , a network 102 and a server 103 .
[0033] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103 , and can be a wired network or a wireless network.
[0034] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is typically the Internet, but can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hypertext Markup Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged over the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPSec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above-mentioned data communication technologies.
[0035] The terminal device 101 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, wearable devices, augmented reality devices, virtual reality devices, etc.
[0036] Optionally, the client of the application installed in different terminal devices 101 is the same, or the client of the same type of application based on different operating systems. The specific form of the application client can also vary based on the terminal platform. For example, the application client can be a mobile phone client, a PC client, etc. The application installed in the terminal device 101 can obtain the user's prompt information Prompt, send the prompt information Prompt to the server 103, and generate an image based on the prompt information Prompt.
[0037] Server 103 can be a server that provides various services, such as a background management server that supports devices operated by users using terminal device 101. The background management server can obtain prompt information Prompt, generate an initial image based on the text-image model and prompt information Prompt, process the prompt information Prompt based on the multimodal large model, obtain feature information of at least one verification type in the prompt information, verify the initial image based on the feature information of each verification type using the multimodal large model, and if the verification passes, determine the initial image as the target image and feedback the target image to terminal device 101.
[0038] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0039] Those skilled in the art will know that Figure 1 The number of terminal devices, networks, and servers in the embodiment is merely illustrative, and any number of terminal devices, networks, and servers may be provided based on actual needs. This embodiment of the present disclosure does not limit this.
[0040] After the Wensheng Graph model generates an image based on the user's prompt, it can be evaluated. During the evaluation process, it was found that the quality of the images generated by the Wensheng Graph model fluctuated greatly and was highly random. The problems included the following:
[0041] 1. The character's movements, expressions, and limbs may be deformed, with distorted expressions and elongated limbs.
[0042] 2. The number of objects in the generated image is difficult to control. For example, if you are asked to generate 7 cats, the generated image may contain 5 cats.
[0043] 3. The generated images may violate the laws of nature, for example, tasks under sunlight have no effect.
[0044] The above problem makes it impossible to accurately generate an image that matches the prompt, which will greatly reduce the user's perception and subjective evaluation of the text image or text image model, resulting in a poor user experience.
[0045] In order to at least solve the above technical problems, the embodiments of the present disclosure provide an image generation and optimization image processing method based on Chain of Thought (CoT), which improves the task performance of the text-based graph model only through the Chain of Thought method without the need to retrain the direct data volume.
[0046] Specifically, the image processing method provided by the embodiment of the present disclosure obtains prompt information; generates an initial image based on the Wensheng graph model and the prompt information; processes the prompt information based on the multimodal large model to obtain feature information of at least one verification type in the prompt information; based on the feature information of each verification type, the initial image is verified through the multimodal large model, and if the verification passes, the initial image is determined as the target image. The present disclosure can improve the quality of the image output by the Wensheng graph model and improve the user experience. Based on the layer-by-layer progressive analysis of CoT, low-quality images can be filtered out early to improve image generation efficiency.
[0047] In the above system architecture, the image processing method provided by the embodiment of the present disclosure can be executed by any electronic device with computing and processing capabilities. In some embodiments, the image processing method provided by the embodiment of the present disclosure can be executed by a server of the above system architecture.
[0048] The image processing method in this exemplary embodiment will be described in more detail below with reference to the accompanying drawings and examples.
[0049] Figure 2 FIG. 1 shows a flow chart of an image processing method according to an embodiment of the present disclosure. Figure 2 As shown, in one embodiment, the image processing method disclosed herein mainly includes the following steps:
[0050] S202: Obtain prompt information.
[0051] In one embodiment, the prompt information is used to prompt the context of input information for the document graph model and input parameter information of the document graph model.
[0052] The prompt information may be collected through a terminal device used by the user. For example, the prompt information may be text information input by the user, or text information converted from voice information input by the user.
[0053] For example, the prompt message is "Please generate an image of a scene of two kittens catching butterflies in a park."
[0054] S204: Generate an initial image based on the text-based graph model and the prompt information.
[0055] In one embodiment, the text-generated image model is a model that generates a style image based on prompt information in text format.
[0056] In one embodiment, the culture graph model may be a diffusion model, an autoregressive model, a generative adversarial network model, etc. The present disclosure does not specifically limit the type of the culture graph model.
[0057] In one embodiment, multiple prompt word texts describing image features may be extracted from the prompt information, the multiple prompt word texts may be input into a text-based image model, and an initial image may be output.
[0058] S206: Process the prompt information based on the multimodal large model to obtain feature information of at least one verification type in the prompt information.
[0059] In one embodiment, the multimodal large model can disassemble the prompt information and extract feature information of at least one verification type from the prompt information.
[0060] In one embodiment, the verification type includes at least one of: subject, environment, style, and emotion. The subject may include a single subject or multiple subjects. The characteristic information of a single subject may include the number, position, posture, expression, etc. of the subjects. The characteristic information of multiple subjects may include the relationship between the subjects and the description of detailed features. The environment is used to describe the specific environmental requirements of the subject. It should be considered whether the description of the subject has environmental restrictions. The characteristic information of the environment may include natural environment, social environment, indoor scene, no scene requirements, or no environmental requirements. The style is used to describe whether the subject and the environment have style requirements. It should be considered whether the description of the subject has style requirements. For example, the characteristic information of the style may include realism, cartoon, photograph, watercolor, oil painting, landscape painting, or no style requirements. The emotion is used to describe the emotion displayed by the subject in the image. Hidden semantics may be considered. The characteristic information of the emotion may include warmth, horror, joy, sadness, no clear emotional requirements, etc.
[0061] The multimodal large model can also be a language large model.
[0062] S208 : Based on the feature information of each verification type, the initial image is verified using the multimodal large model. If the verification passes, the initial image is determined as the target image.
[0063] In one embodiment, the initial image is verified using a multimodal large model to determine the correlation between the initial image and the prompt information, and an image that meets the user's needs is output.
[0064] In one embodiment, the above S208 verifies the initial image through the multimodal large model based on the feature information of each verification type. If the verification passes, the initial image is determined as the target image, including: if the initial image is clear; the target in the initial image is not deformed, not distorted, and does not violate the laws of nature; and / or the initial image matches the feature information of at least one verification type in the prompt information, then the verification passes.
[0065] In one embodiment, if the initial image is not clear, the object in the initial image is deformed, distorted, violates the laws of nature, and / or the initial image does not match the feature information of at least one verification type in the prompt information, the verification fails.
[0066] In one embodiment, the initial image is input into the multimodal large model to analyze whether it is clear. If not, the image is regenerated. If it is clear, the initial image is input into the multimodal large model to determine whether the target is deformed or distorted or violates the laws of nature. If there is a problem, the specific problem description is output, and prompt information is added to regenerate the image. If there is no problem, the initial image and each type of feature information are verified one by one through the multimodal large model. Items without a problem are directly skipped. If any item is not satisfied, the specific problem is output, prompt information is added, and the image is regenerated. If there is no problem, the initial image is determined as the target image, and the target image is output.
[0067] For example, object deformation or distortion includes distorted human faces, unreasonable facial features, deformed limbs, inconsistent or distorted animal limbs, and violations of natural laws, such as objects casting no shadows in the sun.
[0068] It should be noted that the verification process for dimensions such as image clarity, target deformation, and feature information can be performed in parallel or serially, or in a combination of parallel and serial methods. This disclosure does not make specific limitations on this.
[0069] In an embodiment of the present disclosure, prompt information is obtained; an initial image is generated based on the Wensheng graph model and the prompt information; the prompt information is processed based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; based on the feature information of each verification type, the initial image is verified using the multimodal large model; if the verification passes, the initial image is determined as the target image. The present disclosure can improve the quality of the image output by the Wensheng graph model and enhance the user experience. Based on the layer-by-layer progressive analysis of CoT, low-quality images can be filtered out early, thereby improving image generation efficiency.
[0070] Figure 3 FIG. 1 shows another flow chart of an image processing method according to an embodiment of the present disclosure. Figure 3 As shown, in one embodiment, the method further includes:
[0071] S302: If the verification fails, the initial image and the verification result are stored in an image acquisition list;
[0072] S304: If the number of iterations does not reach the preset threshold, the prompt information is adjusted according to the verification result, an image is generated and verified, and if the verification fails, the image list is updated; if the verification passes, the target image is obtained;
[0073] S306: If the number of iterations reaches a preset threshold, a target image is determined in the image acquisition list.
[0074] In S302, the verification results may include the verification results of the initial image for each verification type and the specific reasons for failure of the verification. The form of the image in the list is not specifically limited.
[0075] Table 1 Image acquisition list
[0076]
[0077] As shown in Table 1, the image acquisition list may include image ID, different verification dimensions, verification results, and specific reasons for failure of verification. Among them, for the image with ID1, it passed the verification in terms of clarity, non-distortion, style, and emotion, which is represented by Y, and failed the verification in terms of subject and environment, which is represented by N. The reason for failure of subject verification is the wrong number of subjects, and the reason for failure of environment verification is the wrong environment.
[0078] In S304, the preset number threshold can be determined according to actual needs. For example, the preset number threshold is 10, 15, etc., that is, when the verification result fails, it can be iterated 10 times to generate 10 pictures. During the iteration process, as long as the verification result passes, the image can be output.
[0079] Adjusting the prompt information according to the verification results may include adding prompt words for the verification results that fail the corresponding verification dimensions, modifying the prompt words for the verification results that fail the corresponding verification dimensions, and / or deleting the prompt words for the verification results that fail the corresponding verification dimensions. This disclosure does not make specific limitations on this.
[0080] For example, when the prompt information entered by the user is too complex, the cultural graph model cannot fully generate all the content, which may cause the cultural graph large model image generation optimization evaluation system based on CoT to fail. To ensure efficiency, the maximum number of iterations can be set. The larger the maximum number of iterations, the better.
[0081] In S306 , when the number of iterations reaches a preset threshold, the image that best matches the prompt information may be selected from the image acquisition list and used as the target image.
[0082] In the disclosed embodiment, an initial image is generated based on a prompt input by the user. The prompt is then fed into a multimodal macromodel to determine the generated content to be included in the prompt. The multimodal macromodel then performs multiple evaluations on the generated image. If it passes the evaluation, the generated image is output; otherwise, a reason for failure is given and additional information is added to the prompt. Multiple iterations are performed to output a generated image that passes the evaluation or is optimal, improving the relevance of the generated image to the prompt and the visual appeal of the generated image.
[0083] Figure 4 A flow chart of a method for determining a target image in an embodiment of the present disclosure is shown. Figure 4 As shown, in one embodiment, the verification result includes the number of verification rounds passed; wherein, determining the target image in the image acquisition list in the above S306 includes:
[0084] S402, calculating a correlation score based on the correlation between the images in the image list and the corresponding prompt information;
[0085] S404: Determine the target image based on the correlation score and the number of verification rounds passed.
[0086] For an image in the image list, the number of verification rounds passed can be determined by summing the dimensions passed in the verification results. For example, in Table 1, the number of verification rounds passed for the image ID1 is 4.
[0087] In one embodiment, the correlation between the images in the image acquisition list and the corresponding prompt information can be determined by calculating the correlation score based on feature embedding similarity, such as cosine similarity, dot product or Euclidean distance, etc.; it can also be determined by cross-modal alignment based on the attention mechanism; it can also be determined by probability evaluation based on the generative model, etc., and the present disclosure does not make specific limitations on this.
[0088] In one embodiment, by comprehensively considering the correlation scores and determining the target image through the number of verification rounds, the optimal generated image can be determined, effectively improving the correlation between the generated image and the prompt information and the visual appreciation of the image.
[0089] Figure 5 FIG. 1 shows another flow chart of a method for determining a target image in an embodiment of the present disclosure. Figure 5 As shown, in one embodiment, the above S404 determines the target image based on the correlation score and the number of verification rounds, including:
[0090] S502: Filter a preset number of images ranked top in correlation scores from the image acquisition list;
[0091] S504: Determine the image that has passed the most verification rounds among the preset number of images as the target image, or when the preset number of images have passed the same number of verification rounds, determine the image with the highest correlation score among the preset number of images as the target image.
[0092] In one embodiment, the above-mentioned preset number can be determined according to actual needs. For example, the preset number can be 3, 4, 5, etc., and the preset number is less than the preset iteration round threshold.
[0093] After obtaining the correlation scores of the images in the image list, the images in the image list are sorted in descending order according to the correlation scores, and a preset number of images with the highest correlation scores are selected.
[0094] In addition, the images in the image acquisition list may be sorted in ascending order, and a preset number of images with the later ranked correlation scores are selected, that is, a preset number of images with larger correlation scores are selected. The present disclosure does not specifically limit the ordering method of the images in the image acquisition list.
[0095] In S504 , by comparing the number of times a preset number of images have passed the verification rounds, the image that has passed the most verification rounds is selected as the target image.
[0096] In one embodiment, when a predetermined number of images pass the same number of verification rounds, the image with the highest correlation score is selected as the target image.
[0097] In the disclosed embodiment, the target image is determined by combining the correlation score and the number of verification rounds passed, and the image with the best correlation with the prompt information can be selected from the list of images that failed the verification, thereby ensuring the efficiency of the Vincent graph large model image generation optimization evaluation process, effectively improving the quality of the Vincent graph large model output image and its correlation with the prompt information, and enhancing the user experience. Based on the layer-by-layer progressive analysis of CoT, low-quality images can be filtered out early.
[0098] In order to deepen the understanding of the image processing method disclosed in the present invention, Figures 6 to 9 For detailed description. Figure 6 As shown, an image processing method provided by an embodiment of the present disclosure includes:
[0099] S601, get text Prompt;
[0100] S602, generating an initial image through a large model of Wensheng graph;
[0101] S603, decomposing the multimodal large model Prompt to obtain feature information of at least one verification type;
[0102] S604: Verify the image is clear using the multimodal large model. If the image is clear, the verification passes, and S605 is executed. If the image is not clear, the verification fails, and S608 is executed.
[0103] S605. Verify the image using the multimodal large model to see if the target is deformed, distorted, or violates natural laws. If not, the verification passes and the process goes to S606. If so, the verification fails and the process goes to S608.
[0104] S606: Verify whether the multimodal large model contains the decomposed features, where the decomposed features are feature information of at least one verification type obtained in S603. If they match, the verification passes, and S607 is executed; if they do not match, the verification fails, and S608 is executed.
[0105] S608: Generate the prompt to be added and re-execute S601;
[0106] S607: Output the target image.
[0107] It should be noted that in S604, S605, and S606, if the verification fails, the specific reason for the failure of the verification under the corresponding verification dimension is also required.
[0108] Considering the generation efficiency, if the iteration count exceeds ten, the system will be cancelled. The pedestrian detection result and the pedestrian frame image are obtained.
[0109] If the text image prompt entered by the user is too complex, the text image model may not be able to fully generate all the content, which may cause the CoT-based text image model image generation evaluation to fail verification. To ensure efficiency, it is necessary to preset a maximum number of iterations. This can be set. In theory, in non-urgent situations, the larger the number, the better. In this example, the preset iteration threshold of 10 is used as an example for illustration.
[0110] like Figure 7 As shown, an image processing method provided by an embodiment of the present disclosure, in Figure 6 Based on the example, S609 and S610 are added, as follows:
[0111] S609, when more than ten iterations are completed, the ten images are obtained into a list and verification results are obtained;
[0112] S610: Determine the image that best matches the text Prompt as the target image.
[0113] In S610, the images in the image acquisition list can be input into the multimodal large model one by one, the correlation between the image and the prompt information Prompt can be judged, and scored from 0 to 100. The top three images with the highest correlation are selected, and the number of verification rounds passed by the top three images in the image acquisition list is checked. The image with the largest number of verification rounds is output as the target image. If the number of verification rounds passed is the same, the image with the highest correlation is output as the target image.
[0114] The method of decomposing Prompt in a multimodal large model is as follows:
[0115] Input the text prompt of the text image into the multimodal large model or language large model, and require the multimodal large model or language large model to decompose the prompt into the following contents: subject, environment, style, and emotion. The specific implementation is:
[0116] Subject: Please analyze the text-image prompt "Text Prompt". What is the subject? If it is a single subject, please describe its detailed characteristics including quantity, position, posture, expression, etc.; if there are multiple subjects, please describe the relationship between the subjects and describe the detailed characteristics of each subject.
[0117] Environment: Please analyze the "Text Prompt" prompt to see if there are any specific environment requirements (you should consider whether the subject description has any environmental restrictions), such as natural environment, social environment (ancient, modern, future, etc.), or indoor scenes. If there are no scene requirements, please reply "No environment requirements".
[0118] Style: Please analyze the "Text Prompt" text-image prompt to see if there are any style requirements (you should consider whether the subject description has any style restrictions), such as realism, cartoon, photo, watercolor, oil painting, landscape painting, etc. If there are no style requirements, please reply "No style requirements".
[0119] Emotion: Please analyze the text prompt "Text Prompt" and what emotions it contains (including hidden semantics), such as warmth, horror, joy, sadness, etc. If there is no clear emotional requirement, please reply "No emotion required".
[0120] For the example prompt "Please generate an image of two kittens chasing butterflies in a park," we can use the multimodal large model to decompose the prompt to obtain feature information in four aspects: subject, environment, style, and emotion, as shown below:
[0121] Dimension 1: Subject
[0122] Subject: Two kittens
[0123] Quantity: 2;
[0124] Positional relationship: There is no specific information in the tooltip, so you can assume that they are in the center of the picture or slightly off-center, close to each other and focusing on the same point (butterfly);
[0125] Posture: Both kittens may be in a jumping position, with their front paws extended trying to catch a butterfly, showing vitality and dynamism;
[0126] Expression: Although the prompt message does not explicitly express contempt, one can infer an expression of curiosity and excitement.
[0127] Dimension 2: Environment
[0128] Specific Environment Requirements: Park - This setting clearly defines a natural environment, suggesting an outdoor scene with elements such as trees, grass, and flowers. Considering the activity of catching butterflies, the scene should include some vegetation as background or embellishment to enhance authenticity and context.
[0129] Dimension 3: Style
[0130] Style Requirements: No specific style requirements - Although not specified, based on the content description, a more lively and vivid style may be preferred. If you want to be more specific, you can choose realistic, cartoon, or other styles as needed. The default here is "No style requirements."
[0131] Dimension 4: Emotion
[0132] Emotion: Cheerful - The scene depicting two kittens chasing butterflies in the park usually gives people a light-hearted and cheerful feeling, conveying a carefree and fun-filled emotional atmosphere.
[0133] The multimodal large model evaluates images as follows:
[0134] When the text graph model is generated as Figure 8 When the image shown is Figure 8 The image shown is fed into a large multimodal model and asked;
[0135] Analyze whether the image is clear. --Verification passed.
[0136] Determine whether the target is deformed or distorted, such as a distorted face, unreasonable facial features, or deformed limbs; an animal has the wrong number of limbs or distorted limbs, or violates natural laws, and have the large model output specific problems. --Verification passed.
[0137] Analyzing the main content, the two kittens: Number: 2; Position and Relationship: Since there's no specific description, we can assume they're in the center or slightly off-center, close together, and focused on the same point (the butterfly). Posture: Both kittens are likely leaping, their front paws extended, attempting to catch the butterfly, displaying energy and dynamism. Expression: While not explicitly described, we can infer curiosity and excitement. --Verification passed.
[0138] Analyze the environment section, parks. --Verification passed.
[0139] Analysis style section, no clear style. --Skip.
[0140] Analyzing the emotional part, cheerful. --Verification passed.
[0141] Output the image.
[0142] When the text graph model is generated as Figure 9 When the image shown is Figure 9 The image shown is fed into a large multimodal model and asked;
[0143] Analyze whether the image is clear. --Verification passed.
[0144] Determine whether the target is deformed or distorted, such as a distorted face, unreasonable facial features, or deformed limbs; an animal has the wrong number of limbs or distorted limbs, or violates natural laws, and have the large model output specific problems. --Verification passed.
[0145] Analyze the main content, two kittens: Number: 2; Position and Relationship: Since there's no specific description, we can assume they're in the center or slightly off-center, close together, and focused on the same point (the butterfly). Posture: Both kittens are likely leaping, their front paws extended, attempting to catch a butterfly, displaying energy and dynamism. Expression: Although not explicitly described, we can infer curiosity and excitement. --Verification failed. Output the specific reason for verification failure: There are three kittens in the image. Modify the prompt in the text image to "Generate a scene of two kittens catching butterflies in a park. Note that there are only two kittens."
[0146] This paper uses a CoT-based optimized evaluation method for image generation of a cultural graph model, evaluates the generated images using a large multimodal model, and proposes prompt optimization. The resulting image conforms to the initial prompt description, which can improve the quality of the cultural graph model output image and its relevance to the prompt information, thereby enhancing the user experience. The layer-by-layer progressive analysis based on CoT can filter low-quality images early to improve efficiency.
[0147] Based on the same inventive concept, the present disclosure also provides an image processing device, as described in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0148] Figure 10 FIG. 1 is a schematic diagram showing the structure of an image processing device according to an embodiment of the present disclosure. Figure 10 As shown, in one embodiment, the device includes: an acquisition module 1010, used to obtain prompt information; a generation module 1020, used to generate an initial image based on the text image model and the prompt information; an extraction module 1030, used to extract features of the prompt information based on the multimodal large model, and obtain feature information of at least one verification type in the prompt information; a verification module 1040, used to verify the initial image through the multimodal large model based on the feature information of at least one verification type, and if the verification passes, the initial image is determined as the target image.
[0149] It should be noted that the acquisition module 1010, extraction module 1020, extraction module 1030, and verification module 1040 described above correspond to S202 to S208 in the method embodiment. The examples and application scenarios implemented by these modules and corresponding steps are the same, but are not limited to the contents disclosed in the method embodiment described above. It should be noted that the above modules, as part of the apparatus, can be executed in a computer system, such as a set of computer-executable instructions.
[0150] In one embodiment, the verification module 1040 is used to determine that the verification is passed if the initial image is clear; the target in the initial image is not deformed, not distorted, and does not violate the laws of nature; and the initial image matches the feature information of at least one verification type in the prompt information.
[0151] It should be noted that the verification type includes at least one of: subject, environment, style, and emotion.
[0152] In one embodiment, the device also includes an iterative processing module not shown in the accompanying drawings, and the iterative processing module is used to store the initial image and the verification result in the image acquisition list if the verification fails; if the number of iterations does not reach the preset number threshold, the prompt information is adjusted according to the verification result, the image is generated and the image is verified, if the verification fails, the image acquisition list is updated, if the verification passes, the target image is obtained; if the number of iterations reaches the preset number threshold, the target image is determined in the image acquisition list.
[0153] In one embodiment, the verification result includes the number of verification rounds passed; the iterative processing module is used to calculate the correlation score based on the correlation between the images in the image acquisition list and the corresponding prompt information; and determine the target image based on the correlation score and the number of verification rounds passed.
[0154] In one embodiment, the iterative processing module is used to filter a preset number of images with the highest correlation scores from the image acquisition list; determine the image that has passed the largest number of verification rounds among the preset number of images as the target image, or when the number of verification rounds passed by the preset number of images is the same, determine the image with the highest correlation score among the preset number of images as the target image.
[0155] In an embodiment of the present disclosure, prompt information is obtained; an initial image is generated based on a text graph model and the prompt information; the prompt information is processed based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; based on the feature information of each verification type, the initial image is verified using the multimodal large model; if the verification passes, the initial image is determined as the target image. The present disclosure can improve the quality of images output by the text graph model and enhance user experience. Based on the layer-by-layer progressive analysis of CoT, low-quality images can be filtered out early, thereby improving image generation efficiency.
[0156] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods, or program products. Therefore, various aspects of the present disclosure may be implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits," "modules," or "systems."
[0157] Refer to the following Figure 11 1100 according to this embodiment of the present disclosure will be described. Figure 11 The electronic device 1100 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present disclosure.
[0158] like Figure 11 As shown, electronic device 1100 is implemented as a general-purpose computing device. Components of electronic device 1100 may include, but are not limited to, the aforementioned at least one processing unit 1110, the aforementioned at least one storage unit 1120, and a bus 1130 connecting various system components (including storage unit 1120 and processing unit 1110).
[0159] The storage unit stores program code, which can be executed by the processing unit 1110, so that the processing unit 1110 performs the steps described in the "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 1110 can perform the following steps of the above-mentioned method embodiment: obtaining prompt information; generating an initial image based on the text graph model and the prompt information; processing the prompt information based on the multimodal large model to obtain feature information of at least one verification type in the prompt information; based on the feature information of each verification type, verifying the initial image using the multimodal large model, and if the verification passes, determining the initial image as the target image.
[0160] The storage unit 1120 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 11201 and / or a cache memory unit 11202 , and may further include a read-only memory unit (ROM) 11203 .
[0161] The storage unit 1120 may also include a program / utility 11204 having a set (at least one) of program modules 11205, such program modules 11205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0162] The bus 1130 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0163] Electronic device 1100 may also communicate with one or more external devices 1140 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1100, and / or any device that enables electronic device 1100 to communicate with one or more other computing devices (e.g., a router, modem, etc.). Such communication may occur via input / output (I / O) interface 1150. Furthermore, electronic device 1100 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via network adapter 1160. As shown, network adapter 1160 communicates with other modules of electronic device 1100 via bus 1130. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with electronic device 1100, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0164] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0165] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer program product, which includes: a computer program, which implements the above image processing method when executed by a processor.
[0166] In an exemplary embodiment of the present disclosure, a computer-readable storage medium is also provided. The computer-readable storage medium may be a readable signal medium or a readable storage medium. The computer-readable storage medium stores a program product capable of implementing the above-mentioned method of the present disclosure. In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.
[0167] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0168] In the present disclosure, a computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0169] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0170] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0171] It should be noted that although several modules or units of the device for action execution are mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.
[0172] Furthermore, although the steps of the method of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the steps must be performed in this particular order, or that all steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0173] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0174] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
Claims
1. An image processing method, characterized in that: include: Get prompt information; Generate an initial image based on the text graph model and the prompt information; Processing the prompt information based on the multimodal large model to obtain feature information of at least one verification type in the prompt information; Based on the feature information of each verification type, the initial image is verified using the multimodal large model. If the verification passes, the initial image is determined as the target image.
2. The image processing method according to claim 1, wherein: The verifying the initial image using the multimodal large model based on the feature information of each verification type, and determining the initial image as the target image if the verification passes, includes: If the initial image is clear; the target in the initial image is not deformed, distorted, or violates the laws of nature; and / or the initial image matches the feature information of at least one verification type in the prompt information, the verification passes.
3. The image processing method according to claim 1, wherein: The verification type includes at least one of subject, environment, style, and emotion.
4. The image processing method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the verification fails, the initial image and the verification result are stored in an image acquisition list; If the number of iterations does not reach a preset threshold, the prompt information is adjusted according to the verification result, an image is generated and the image is verified. If the verification fails, the image list is updated. If the verification passes, the target image is obtained. If the number of iterations reaches a preset threshold, the target image is determined in the image acquisition list.
5. The image processing method according to claim 4, characterized in that The verification result includes the number of verification rounds passed; Wherein, determining the target image in the image acquisition list includes: Calculating a correlation score based on the correlation between the images in the image list and the corresponding prompt information; The target image is determined according to the correlation score and the number of verification rounds passed.
6. The image processing method according to claim 5, characterized in that Determining the target image according to the correlation score and the number of verification rounds passed includes: Filtering a preset number of images ranked top in correlation scores from the image acquisition list; The image that has passed the verification round the most times among the preset number of images is determined as the target image, or When the preset number of images pass the verification rounds the same number of times, the image with the highest correlation score among the preset number of images is determined as the target image.
7. An image processing device, characterized in that: include: Acquisition module, used to obtain prompt information; A generation module, configured to generate an initial image based on the text graph model and the prompt information; an extraction module, configured to perform feature extraction on the prompt information based on a multimodal large model to obtain feature information of at least one verification type in the prompt information; A verification module is used to verify the initial image through the multimodal large model based on the feature information of the at least one verification type, and if the verification passes, determine the initial image as the target image.
8. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to perform the image processing method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 6 is implemented.
10. A computer program product having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 6 is implemented.