A computer-implemented process, producing a computer program and a computer system for generating and validating images.

The method automates GAI image validation using text description matching, neuro-aesthetic criteria, and heatmaps to reduce user intervention and energy consumption, improving the efficiency of GAI image generation.

FR3165984A1Inactive Publication Date: 2026-03-06ACCENTURE GLOBAL SOLUTIONS LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
FR2024009304
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-31
Publication Date
2026-03-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing generative artificial intelligence (GAI) image generation methods require extensive user intervention and energy consumption due to subjective validation processes, leading to high dissatisfaction and inefficient image generation loops.

Method used

A computer-implemented method and system that utilizes text description matching, neuro-aesthetic criteria, and heatmaps to validate GAI images, reducing user interaction and energy consumption by automating the validation process.

Benefits of technology

Enhances the probability of user acceptance of GAI images, minimizing resubmission loops and energy consumption by providing objective validation criteria.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention relates to a method, a system, and computer-readable storage media for image generation and validation. Information describing the characteristics of a desired image is received and enhanced into a text prompt. The enhanced text prompt is used to generate a generative artificial intelligence (GAI) image, and a GAI text description of the GAI image is generated. Furthermore, validations are performed to determine whether the generated GAI image is valid based on a comparison of the enhanced prompt with the GAI text description, a list of predetermined neuro-aesthetic criteria, and a heat map. If the generated GAI image is valid, it is used for further processing. If the generated GAI image is invalid, the text prompt enhancement or GAI image generation process is repeated. Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Computer-implemented method, computer program and computer system for generating and validating images Scope of the invention

[0001] Various embodiments described herein generally relate to a computer-implemented method, a computer system, and a computer program product for the generation and validation of generative artificial intelligence (GAI) images. Background of the invention

[0002] Humanity is entering a new era of creativity—an era in which anyone can generate digital content. Artificial intelligence is finding applications in various use cases within the context of digital content generation. In the field of AI, generative AI (GAI) has found effective application in text-to-image generation, where it is used to generate images from text prompts with zero examples in natural language, with the aim of creating realistic and diverse images. Summary

[0003] Implementations of this disclosure generally relate to the generation and validation / evaluation of generative artificial intelligence (GAI) images with reduced user intervention and energy consumption. The images are validated using text description / text prompt matching, neuro-aesthetic criteria, and heat maps.

[0004] In general, innovative aspects of the object described in this specification provide a computer-implemented method for generating and validating images. The method includes the initial reception of information describing characteristics of a desired image and the enhancement of this information into a text prompt. The method includes the initial submission of the enhanced text prompt to a generative artificial intelligence (GAI) image generator. The method includes the second reception of a generated GAI image corresponding to the text prompt from the GAI image generator. The method includes the third reception of a GAI text description of the generated GAI image from a GAI image description engine. The method includes the initial determination of whether the GAI text description sufficiently matches the enhanced text prompt with respect to a first predetermined threshold.In response to the fact that the first determination discovers a non-match in a first predetermined variance with respect to the first predetermined threshold, the . The process involves performing the initial submission. In response to the fact that the initial determination reveals a mismatch in a second predetermined variance relative to the first predetermined threshold, the process involves improving and configuring the information based on the enhanced text prompt and the problems identified with the generated GAI image. The second predetermined variance is greater than the first predetermined variance. The process involves determining whether the generated GAI image sufficiently matches a list of predetermined neuro-aesthetic criteria relative to a second predetermined threshold.In response to the fact that the second determination reveals a mismatch below the second predetermined threshold, the method involves enhancing and configuring information based on the enhanced text prompt and elements from the list of predetermined neuro-aesthetic criteria not found in the generated GAI image. The method includes the fourth reception of a heatmap of the generated GAI image. In response to the rejection of the heatmap, the method includes the first reception of additional information.In response to at least one combination of the fact that the first determination finds that the GAI textual description sufficiently matches the enhanced textual prompt, the fact that the second determination finds that the generated GAI image sufficiently matches the list of predetermined neuro-aesthetic criteria, and the acceptance of the generated heat map, the process involves the transmission of the generated GAI image for further use and / or further processing.

[0005] This disclosure further relates to a system for implementing the method described herein. This disclosure also relates to computer-readable storage media coupled to one or more processors and on which instructions are stored which, when executed by the processor(s), cause the processor(s) to perform operations in accordance with the method described herein.

[0006] It is understood that processes according to this disclosure may include any combination of the aspects and features described herein. In other words, the process according to this disclosure is not limited to the combinations of aspects and features specifically described herein, but also includes any combination of the aspects and features provided.

[0007] Details of one or more embodiments of this disclosure are set forth in the accompanying drawings and in the following description. Other features and advantages of this disclosure will become apparent from the description, the drawings, and the claims. Brief description of the figures

[0008] Various embodiments conforming to this disclosure will be described with reference to the drawings, in which:

[0009] [Fig. 1] represents an example environment that can be used to run implementations of this disclosure.

[0010] [Fig.2] represents an example of the architecture of a computer system for generating and validating images according to embodiments of the present disclosure.

[0011] [Fig.3] is a schematic diagram of an image generation and validation engine of the image generation and validation computer system according to embodiments of the present disclosure.

[0012] [Fig.4] is a block diagram that presents an example of a first validator for validating a generative artificial intelligence (GAI) image on the basis of a matching of text prompts in accordance with the embodiments of this disclosure.

[0013] [Fig.5] is a block diagram that presents an example of a second validator for validating the GAI image on the basis of a list of neuro-aesthetic criteria in accordance with embodiments of this disclosure.

[0014] [Fig.6] is a block diagram that presents an example of a third validator for validating the GAI image on the basis of a heat map in accordance with the embodiments of this disclosure.

[0015] [Fig.7] is a flowchart that presents an example of a computer-implemented method for generating and validating GAI images in accordance with the implementations of this disclosure.

[0016] [Fig.8A], [Fig.8B], [Fig.9A], [Fig.9B], [Fig.lOA], [Fig.lOB], [Fig.llA] and [Fig.llB] represent the generation of acceptable GAI images in accordance with the embodiments of this disclosure.

[0017] [Fig. 12] illustrates a computer system that can be used to implement the generation and validation of GAI images in accordance with the implementations of this disclosure.

[0018] Reference numbers and similar designations in the different drawings indicate similar elements. Detailed description

[0019] In the following description, various embodiments will be illustrated by way of example and not limitation in the figures of the accompanying drawings. References to different embodiments in this disclosure do not necessarily refer to the same embodiment, and such references mean at least one. Although specific implementations and other details are discussed, it is It is understood that this is for illustrative purposes only. A person skilled in the art will recognize that other components and configurations can be used without departing from the scope of the claimed object.

[0020] Any reference to an “example” (e.g., “for example”, “an example of”, “by way of example” or similar) shall be regarded as non-limiting examples, whether expressly stated or not.

[0021] The terms used in this specification generally have their ordinary meaning in the field, in the context of the disclosure, and in the specific context in which each term is used. Alternative language and synonyms may be used for one or more of the terms discussed here, and no particular emphasis should be placed on whether a term is specified or discussed here. Synonyms for some terms are provided. Mention of one or more synonyms does not preclude the use of other synonyms. The use of examples anywhere in this specification, including examples of any terms discussed here, is for illustrative purposes only and is not intended to further limit the scope and meaning of the disclosure or of any term cited as an example. Likewise, the disclosure is not limited to the various embodiments given in this specification.

[0022] Without limiting the scope of this disclosure, examples of instruments, apparatus, processes, and their associated results according to embodiments of this disclosure are given below. Note that headings or subheadings may be used in the examples for the convenience of the reader, which should in no way limit the scope of the disclosure. Unless otherwise defined, technical and scientific terms used in this document have the meaning commonly understood by a person skilled in the art to whom this disclosure relates. In case of conflict, this document, including the definitions, shall prevail.

[0023] The term "including", when used, means "comprising, but not necessarily limited to"; it specifically indicates inclusion or open membership in the combination, group, series and analogs thus described.

[0024] “First”, “second”, etc., are designations used to distinguish between components or blocks of otherwise similar names, but do not imply any sequence or numerical limitation.

[0025] “And / or” for two possibilities means either of the stated possibilities or The two (“A and / or B” covers A alone, B alone, or A and B together), and when present with three or more stated possibilities, means any individual possibility alone, or all possibilities together, or a combination of possibilities less than all possibilities. The language in the format “at least one of A... and N” where A to N are possibilities, means "and / or" for the indicated possibilities (e.g., at least one A, at least one N, at least one A and at least one N, etc.).

[0026] It should also be noted that in certain variants of embodiments, the functions / actions shown may take place outside the order shown in the figures. For example, two steps described or shown successively may in fact be executed substantially simultaneously or sometimes in reverse order, depending on the functionalities or actions involved.

[0027] Specific details are provided in the following description to enable a thorough understanding of the embodiments. However, it will be understood by those skilled in the art that embodiments can be practiced without these specific details. For example, systems may be represented in the form of block diagrams so as not to unnecessarily obscure the embodiments. In other cases, well-known processes, structures, and techniques may be presented without superfluous detail so as not to obscure examples of embodiments.

[0028] The specification and drawings are to be considered by way of illustration only and without limitation. It will be evident, however, that various modifications and transformations may be made without departing from the scope and extent of the invention as stated in the claims.

[0029] Generative artificial intelligence (GAI) has become a popular and interactive paradigm for image generation. GAI includes text-image models, which are used to generate realistic and diverse images (hereinafter referred to as GAI images) based on text prompts.

[0030] An example of GAI-based image generation allows a user to request GAI's text-image models to generate GAI images. As an example, a user provides input describing what they need in a desired GAI image. The input is enhanced into a text prompt via preprocessing. The enhanced text prompt is submitted to the text-image models for GAI image generation. The GAI image is presented to the user. The user must either accept the generated GAI image or enter new information to generate a new GAI image if they are not satisfied with the generated image. The new information includes modifications to the inputs initially provided by the user. The new information can be submitted again to the text-image models to regenerate the new GAI image.This process of resubmitting and regenerating images continues until the user either accepts the GAI image or abandons it. Therefore, GAI image generation operates on a take-or-leave-it basis and requires a high degree of iteration and experimentation to obtain a satisfactory GAI image.

[0031] Traditional image generation methodologies present several technical problems. Validating the quality and authenticity of the GAI image is a highly visual and subjective process, requiring a level of discernment and subjective judgment inherent in the user. For example, what is visually appealing to one user may not appeal to another. The user can look at the image and manipulate the inputs / text prompts based on their judgment until they obtain a satisfactory GAI image. However, exemplary GAI image generation has no specific way of taking into account the user's subjective preferences / tastes. Therefore, the user must subjectively identify their preferences regarding the image and manually try to modify the image according to their preferences by changing the inputs / text prompts.

[0032] Furthermore, since GAI image validation is a highly subjective process, the user may not even know exactly what is wrong with the GAI image, other than disliking the generated GAI image. Consequently, the user may not be able to provide sufficient input in a subsequent loop to request the text-image models to produce a satisfactory GAI image, resulting in user dissatisfaction. By analogy, the image may be worth 1000 words, but the GAI image generation example does not provide a mechanism for the user to find the appropriate words needed to generate the image they are trying to describe. Moreover, the nuanced nature of aesthetic preferences, contextual understanding, and ethical considerations introduce complexities into the design of comprehensive GAI image validation.Therefore, finding a balance between the creative potential of GAI imagery and the need for manual effort to discern GAI imagery raises critical questions about the reliability, interpretability, and ethical implications of GAI image validation.

[0033] Consequently, with exemplary GAI image generation, the probability of user dissatisfaction with the GAI images is high, leading to a greater number of resubmission-revision loops and requiring extensive user interaction. Furthermore, each resubmission loop has its own energy requirements. Therefore, the demand placed on the text-image models for generating GAI images consumes a considerable amount of energy and processing capacity. Moreover, the continuous resubmission loops for revised GAI images represent a collective energy consumption.

[0034] As a result, the implementations of this disclosure allow for efficient validation of GAI images by increasing the probability of acceptance of GAI images and reducing the overall energy consumption required to achieve acceptance of GAI images.

[0035] [Fig. 1] represents an example environment 100 that can be used to perform implementations of this disclosure. In some examples, the example environment 100 handles image generation and validation.

[0036] As shown in [Fig. 1], the example environment 100 comprises one or more computing devices 102, one or more computing systems 104, and a network 106. The computing device 102 and the computing system 104 can communicate with each other using the network 106. In some examples, the network 106 may include a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. In some examples, the network 106 can be accessed via a wired and / or wireless communication link.

[0037] In certain examples, the computing device 102 is used by a respective user 108 to connect to and interact with computing platforms running image-generating applications. Examples of computing devices 102 may include a desktop computer, a smartphone, a laptop, a tablet, a voice-activated device, and / or the like. It is anticipated that implementations of this disclosure may be made with any suitable type of computing device. Examples of computing platforms may include content distribution platforms, multimedia platforms, and / or the like.In some examples, the computer device 102 can display one or more graphical user interfaces (GUIs) that allow the user 108 to interact with the computer platform running the image generation applications. Interaction with the computer platform may include providing information to generate one or more images. This information may describe characteristics of the image to be generated. In some examples, the information may be provided in the form of text prompts for image generation.

[0038] In some examples, the computer system 104 may be implemented as an on-premises system operated by a company or a third party engaged in cross-platform interactions and image generation management. In some examples, the computer system 104 may be implemented as an off-premises system (e.g., cloud-based or on-demand) operated by a company or a third party on behalf of a company. In some examples, the computer system 104 may be implemented in a cloud environment. For simplicity, the computer system 104 shown in [Fig. 1] may be a cloud environment intended to represent various types of servers, including a web server, an application server, a proxy server, a network server, a group of servers, and / or the like.

[0039] In some examples, the computer system 104 hosts the image generation applications, which can be run on the computer platforms (with which the user 108 of the computer device 102 can interact for image generation). The image generation applications can provide image generation functions or services.

[0040] According to embodiments of this disclosure, the computer system 104 enables the generation of images based on information received from the computer device 102 and the validation / evaluation of the generated images. The computer system 104 is described in detail in [Fig. 2].

[0041] [Fig. 2] represents an example of the architecture of the computer system 104 for generating and validating images according to embodiments of the present disclosure. As shown in [Fig. 2], the computer system 104 can be configured to communicate with a generative artificial intelligence (GAI) image generator 202, a GAI image description engine 204, a heatmap generator 206, and a data store 208.

[0042] The GAI image generator 202 generates images based on text prompts. Hereafter, images generated using the GAI image generator 202 are referred to as GAI images. The GAI image generator 202 may include one or more GAI models 202a-202n / text-image models, which may be instructed to generate / create GAI images based on text prompts. The GAI models 202a-202n may be trained using deep learning techniques. Deep learning techniques may enable the GAI models 202a-202n to learn patterns and features from a large amount of training data to generate the GAI images. In some examples, the GAI 202a-202n patterns can be classified as, for example, variable autoencoders (VAEs) or generative adversarial networks (GANs), which are known and not described here.

[0043] Although implementations of this disclosure are described in more detail herein with non-limiting reference to the 202 GAI image generator comprising the GAI 202a-202n models for the generation of GAI images, it is envisaged that implementations of this disclosure may also be achieved using any suitable machine learning (ML) or artificial intelligence (AI) model.

[0044] The GAI image description engine 204 generates GAI text descriptions for the GAI images generated by the GAI image generator 202. The GAI text descriptions can be generated in a natural language format. The GAI text descriptions describe various characteristics of the generated GAI images. In some examples, the characteristics of the GAI image may Describe characteristics of objects present in the GAI image, details concerning the objects' surrounding environment, and visual characteristics of the GAI image. Visual characteristics can relate to aspects such as appearance, style, presentation, perspective, and / or analogs. Examples of visual characteristics may include, but are not limited to, GAI image type, color, intensity characteristics, texture patterns, image layout characteristics (shape, structure, or analog), neuro-aesthetic elements, and / or analogs. In some examples, the GAI image description engine may employ various models, for example, foundational / large language models (LLMs) (e.g., GPT Vision), computer vision models, machine learning models, AI models, and / or analogs, to generate the GAI text descriptions.Such models are already known and are not described further.

[0045] The heatmap generator 206 generates heatmaps of the GAI images generated by the GAI image generator 202. A heatmap of the GAI image can provide clear indications of areas / regions of the GAI image that are of interest to the user or that require focused visual attention. The areas / regions can be linked to a visual appeal of the GAI image. Therefore, the user can focus on these visually attractive areas / regions in the GAI image. A non-limiting methodology for generating the heatmap uses a CRISP engine, as described in US patent 10957086B1.

[0046] The data store 208 can serve as a repository for storing various data necessary for the validation of GAI images. The data store 208 can include a list of neuro-aesthetic criteria 210 defined for the generation of specific GAI images, a set of image layout rules 212, first and second thresholds 214-216 defined for the validation of GAI images, a set of feedback parameters 218, one or more sets of external parameters 220 and / or similar.

[0047] The neuro-aesthetic criteria list 210 defines multiple elements / neuro-aesthetic features to be evaluated in the generated GAI images. The elements defined by the neuro-aesthetic criteria list are related to the visual characteristics / appearances of the GAI images that the user generally finds pleasing. These elements can be evaluated to find an emotional response likely received from the user to the GAI images. The multiple elements can be defined based on the user's aesthetic preferences, ethical and / or similar considerations. The user's aesthetic preferences may be collected and used only on the basis of explicit consent obtained from the user. Furthermore, the user's aesthetic preferences may be stored and deleted in accordance with regulations and the user's prior consent. By Therefore, the implementations of this disclosure only work on the small piece of data to which the user has consented, and do not work on a full brain scan data set. Ethical considerations may indicate laws and / or rules and / or regulations applicable to the generation of GAI images.

[0048] In some embodiments, the multiple elements of the neuro-aesthetic criteria list 210 may indicate colors, shapes, objects, and / or analogs. For example, the neuro-aesthetic criteria list may indicate evaluating elements such as: colors are pale, shapes are round, trees are present, and / or analogs in the generated GAI image.

[0049] The neuro-aesthetic criteria list 210 can define any number of elements based on the variance of the GAI images that the computer system 104 can generate according to the implementation of the present invention. For example, the neuro-aesthetic criteria 210 can define 10 to 20 elements to be presented in the GAI image. Furthermore, each of the elements defined by the neuro-aesthetic criteria list 210 can be assigned a weight that indicates the priority / importance of the respective element. Consequently, the elements can be evaluated in the GAI image according to their weight. For example, if a "color" element is assigned a greater weight than a "tree" element (an example of the object), then it is understood that colors are more important than trees for validation.

[0050] The image layout rule set 212 can be used for validating GAI images. The image layout rule set 212 specifies geometric characteristics of the objects to be evaluated in the GAI images. In some examples, the image layout rule set 212 may specify the symmetry, proportion, size, and / or analogy of the objects. The image layout rule set 212 can be defined and modified dynamically according to the generation of specific GAI images.

[0051] The first and second thresholds 214-216 can be used in the validation of GAI images (described in detail below). In some examples, the user may have the option to configure and refine the first and second thresholds 214-216 based on the generation of specific GAI images. In some examples, the computer system 104 can dynamically determine and refine the first and second thresholds 214-216 based on any of the GAI images already generated (for example, all previously generated GAI images) that have been successfully validated and accepted by the user.

[0052] The feedback parameter set 218 may include parameters to be considered for the generation of GAI images. The parameters may be collected and stored based on validations of previous GAI images. The parameters may These parameters can indicate problems in previous GAI images that caused validations of those images to fail. They can specify contextual information, visual characteristics, unbiased data for biased data, and / or similar data to be considered when generating specific GAI images.

[0053] The external parameter set 220 may include parameters to be considered for the generation and validation of GAI images. The parameters may include a set of branding rules, refinement parameters, guidelines, and / or the like. The branding rule set may be considered for the generation of GAI images related to any product and may relate to product branding. The branding rule set may include considerations such as color(s), product content alignment, and / or the like. Therefore, the branding rule set may support the generation of GAI images associated with specific emotions and the corresponding product-related attributes. The guidelines may specify rules that prevent the creation of specific content in the GAI images.Content to be prevented in GAI images may include offensive content, content that violates laws, rules, and regulations, content containing sensitive or protected data, or similar content. The guidelines may also specify rules for removing biased data from GAI images. Image generation and validation are described in detail below, along with the components of the computer system 104.

[0054] Still referring to [Fig.2], the computer system 104 comprises one or more processors 222 and a memory 224. The computer system 104 may also include other components such as communication interfaces, input / output (I / O) devices, etc. (not shown in [Fig.2]).

[0055] In some examples, the processor 222 may include, but is not limited to, microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuits, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or any device that manipulates data or signals according to operational instructions. Among other possibilities, the processor 222 may be programmed to cooperate with computer-readable instructions stored in the memory 224 (also called computer-readable media) to perform operations according to the present invention. The memory 224 may be a non-transient or non-volatile medium, such as a magnetic disk or semiconductor non-volatile memory, or a volatile medium such as random access memory (RAM), and / or the like.

[0056] The computer system 104 further includes an image generation and validation engine 226, as shown in [Fig. 2]. The image generation and validation engine 226 can be stored in memory 224 and provided as a downloadable library containing computer-readable instructions. The image generation and validation engine 226 can be run on the processor 222 for the generation and validation of GAI images.

[0057] The image generation and validation engine 226 includes an interface module 228, a prompt enhancer 230, an image and description generator 232 and an image evaluator 234 (also called an image filter).

[0058] The interface module 228 can represent one or more front-end / front-end components / interfaces of the image generation application. The image generation application can be run on the computing platform with which the computing device / user 102 can interact to provide information (also called a first-intent prompt, initial text prompt, or analog) describing features of a desired GAI image. In some examples, the information can be received through various modalities, including, but not limited to, input into a chatbot, information provided via a GUI, and / or the like.In some examples, the characteristics of the desired GAI image may indicate: an object (or objects) to be present in the desired GAI image and / or an environment in which the objects must be present and / or visual characteristics of the desired GAI image and / or analogous features. Therefore, the characteristics provide context for the generation of the GAI image.

[0059] Once the information is received, the prompt enhancer 230 improves the received information into a text prompt. In some embodiments, the improvement of the received information in the text prompt may involve adding supplementary information to the received information. Therefore, the enhanced text prompt may include both the received information and the supplementary information. The supplementary information may include: additional context for generating the GAI image and / or specific GAI keywords and / or a type of GAI image to be generated and / or the visual features to be presented / enhanced in the GAI image and / or the like. In some examples, the context may indicate industry / business considerations, demographic considerations, the visual appearance of objects and / or the like.In some examples, the type of GAI image to be generated may indicate an image based on acrylic paint, an image based on oil paint, a digitally manipulated image, and / or analog. In some examples, the visual characteristics may indicate color, brightness, contrast, intensity characteristics, etc. The textures, layouts, and / or analogs of the GAI image to be generated. In some other examples, the additional information may indicate unbiased data for biased data present in the information received for generating the desired image.

[0060] In some examples, the prompt enhancer 230 can use various models, such as foundation / LLM models, ML models, AI models, and / or similar models, to enhance the information received in the text prompt. The prompt enhancer 230 can input the received information into one of the models, which is trained to enhance the information received in the text prompt. In some examples, the prompt enhancer 230 can input the received information, along with the feedback parameter set 218 and the external parameter set 220, into one of the models. The feedback parameter set 218 and the external parameter set 220 are accessible from the data store 208. The feedback parameter set 218 can indicate problems identified during the generation of the previous similar GAI image.The external parameter set 220 may include the branding rule set, refinement parameters, guidelines, and / or similar elements. In response to the input information and / or the feedback parameter set 218, the prompt enhancer 230 can receive the model's enhanced text prompt. Therefore, the enhanced text prompt can be derived based on additional information / criteria determined to be appealing to the user or to elicit a desired user response within the context of user safety.

[0061] Based on the enhanced text prompt, the image and description generator 232 enables the generation of the GAI image. The image and description generator 232 submits the enhanced text prompt to the GAI image generator 202. In some examples, the image and description generator 232 may submit the enhanced text prompt along with the feedback parameter set 218 and the external parameter set 220 (accessible from data memory 208) to the GAI image generator 202 for the generation of the GAI image. In response to this submission, the image and description generator 232 receives the generated GAI image from the GAI image generator 202.

[0062] The image and description generator 232 also allows for the generation of a GAI textual description for the generated GAI image. The image and description generator 232 submits the generated GAI image corresponding to the enhanced text prompt to the GAI image description engine 204. In response to the submission, the image and description generator 232 receives the GAI textual description of the GAI image generated by the GAI image description engine 204.

[0063] The image and description generator 232 provides the generated GAI image and the associated GAI textual description to the image evaluator 234 to automatically validate / evaluate the generated GAI image.

[0064] According to embodiments of the present invention, the image evaluator 234 validates the GAI image based on multiple criteria such as the text description / text prompt correspondence, the neuro-aesthetic criteria list 210, and heatmaps. Accordingly, the image evaluator 234 comprises a first validator 236, a second validator 238, and a third validator 240. The first, second, and third validators 236-240 can operate in a daisy chain. For example, the GAI image successfully validated by the first validator 236 can be sent to the second validator 238. Furthermore, the GAI image successfully validated by the second validator 238 can be sent to the third validator 240 for further validation. All these validators 236-240 are operated sequentially until one of the validators fails.It should be noted that the first validator 236, the second validator 238 and the third validator 240 can be used in any other order, but allowing the third validator 240 to operate at a later stage may tend to minimize user interactions and corresponding user reviews.

[0065] Furthermore, it can be understood that the validation of the GAI image according to the present invention can be performed by implementing user-specified validators with the first, second, and third validators 236-240, for example, responsible AI-based validators, brand compliance-based validators, and / or the like. The user-specified validators may be able to operate anywhere in the GAI image validation chain.

[0066] The first validator 236 validates the GAI image by matching the GAI text description corresponding to the generated GAI image with the enhanced text prompt (which was used to generate the GAI image). If the GAI text description matches the enhanced text prompt, the first validator 236 identifies that the GAI image is valid with respect to the information received for generating the GAI image. If the GAI text description does not match the enhanced text prompt, the first validator 236 identifies that the GAI image is not valid with respect to the information received for generating the GAI image.

[0067] For the validation of the GAI image, the first validator 236 performs a semantic comparison of the GAI textual description and the enhanced textual prompt. Once the semantic comparison is performed, the first validator 236 generates a score (for example, an arbitrary value) for a result of the semantic comparison using, for example, a method based on cosine similarity, which is already known and not described further. In addition, the first validator 236 evaluates The score is relative to the first threshold 214, which is accessed from data store 208. The first threshold 214 can be a similarity value (e.g., percentile) that must be met by a result of the semantic comparison. Alternatively, the first threshold 214 can be predetermined or refined automatically or by the user. The first threshold 214 may have predetermined first and second variances, which help identify whether or not there are problems with generating the GAI image or improving the information in the text prompt. In some examples, the second predetermined variance may be greater than the first predetermined variance.

[0068] If the score falls within the first predetermined variance of the first threshold 214, the first validator 236 finds a first type of mismatch between the GAI text description and the enhanced text prompt. This first type of mismatch can identify problems in the generation of the GAI image. Therefore, once the first type of mismatch is found, the image and description generator 232 can initiate the regeneration of a new GAI image based on the enhanced text prompt and the problems identified in the GAI image. By way of non-limiting example, the problems may indicate that one or more objects are missing from the generated GAI image, that the generated GAI image lacks visual features, that the generated GAI image contains biased and / or analogous data.

[0069] If the score falls within the second predetermined variance of the first threshold 214 (for example, if the score is above the first predetermined variance), the first validator 236 finds a second type of mismatch between the GAI text description and the enhanced text prompt. This second type of mismatch can identify information enhancement problems in the text prompt. Therefore, once the second type of mismatch is found, the prompt enhancer 230 can perform enhancement and configuration of the received information for GAI image generation. The prompt enhancer 230 can perform information enhancement and configuration based on the enhanced text prompt and the identified problems in the generated GAI image. Based on the enhanced text prompt, the image and description generator 232 initiates the regeneration of the GAI image.

[0070] If the score is above the first threshold, the first validator 236 determines that the GAI text description corresponds to the enhanced text prompt, thus successfully validating the generated GAI image. An example of validating the generated GAI image using the first validator 236 is described in detail in connection with [Fig. 4]. When the first validator 236 successfully validates the GAI image, the second validator 238 may be able to operate for subsequent validation.

[0071] The second validator 238 validates the GAI image generated from the list of neuro-aesthetic criteria 210. The list of predetermined neuro-aesthetic criteria 210 for validating GAI images is accessible from the data store 208. The list of neuro-aesthetic criteria 210 can indicate the multiple elements to be evaluated in the generated GAI image. Examples of elements may include luminance, color, faces, bodies, and landscapes associated with objects, emotional and / or analogous aspects.

[0072] For the validation of the generated GAI image, the second validator 238 submits a query to the GAI image description engine 204. The query includes a request to identify a number of elements (from the list of neuroaesthetic criteria 210) present in the generated GAI image. For the submitted query, the second validator 238 receives a response from the GAI image description engine 204. From this response, the second validator 238 identifies the number of elements present in the generated GAI image. The second validator 238 determines whether the number of elements present in the generated GAI image satisfies the second threshold 216 (accessible from the data store 208). The second threshold 216 can indicate a maximum number of elements (from the list of neuroaesthetic criteria 210) present in the generated GAI image.

[0073] If the number of elements present in the generated GAI image does not satisfy the second threshold 216, the second validator 238 determines that the generated GAI image does not contain enough elements / required elements from the neuro-aesthetic criteria list 210. Thus, a mismatch between the elements of the generated GAI image and the elements of the neuro-aesthetic criteria list 210 can be identified. Once the mismatch is identified, the prompt enhancer 230 can perform an enhancement and configuration of the information received for the generation of the GAI image in the text prompt. The prompt enhancer 230 can perform the enhancement and configuration of the information based on the enhanced text prompt and the elements from the neuro-aesthetic criteria list 210 that were not found in the generated GAI image.Based on the enhanced text prompt, the image and description generator 232 initiates the regeneration of the GAI image.

[0074] If the number of elements in the generated GAI image satisfies the second threshold, the second validator 238 determines that the generated GAI image is valid with respect to the information received for generating the GAI image. An example of validating the generated GAI image using the second validator 238 is described in detail with reference to [Fig. 5]. When the second validator 238 successfully validates the GAI image, the third validator 240 may be able to operate for further validation.

[0075] The third validator 240 validates the generated GAI image based on the heatmap. The third validator 240 submits the generated GAI image to the heatmap generator 206 and receives the heatmap for the generated GAI image from the heatmap generator 206. The heatmap can indicate areas / regions of the GAI image that may be of interest to the user. In some embodiments, the third validator 240 evaluates the heatmap to determine how well it conforms to the predefined image layout rule set 212 for generating the GAI image. The image layout rule set 212 can refer to generic image layout rules, for example, symmetry, object proportions, object shapes, and / or the like. The third validator 240 can use one or more of the following: ML models, AI models, and / or the like, to evaluate the heatmap.

[0076] In addition, the heatmap and / or the results of the heatmap evaluation can be provided to the user for acceptance or rejection. If the heatmap is accepted by the user, the third validator 240 determines that the GAI image is valid with respect to the information received for generating the GAI image. If the heatmap is rejected by the user, the third validator 240 determines that the generated GAI image is not valid with respect to the information received for generating the GAI image. The user may then be prompted to re-enter information for generating the new GAI image. In some examples, the user may enter new information, resulting in a new text prompt. In some examples, the user may modify the previously entered information by adding further details.The additional details may reflect changes to be made to the new GAI image.

[0077] When the first, second, and third validators 236-240 determine that the generated GAI image is valid with respect to the information received for GAI image generation, the interface module 228 transmits the generated GAI image to the computer device 102 / to the user for further use and / or processing. The generated GAI image is an optimized, debiased, and high-quality image. With the proposed validation, the probability of user acceptance of the GAI image is high. Consequently, the resubmission of the text prompt / information for GAI image regeneration and the review of the regenerated GAI image is reduced, which further reduces the overall energy consumption of the computer system 104 during GAI image generation and validation.

[0078] [Fig.3] is a block diagram of the 226 generation and validation engine image for the generation and validation of images according to embodiments of the present invention.

[0079] The prompt enhancer 230 receives the initial text prompt 302 from the computing device 102, and optionally the feedback parameter set 218 and a first set of external parameters 304 from the data store 208. The initial text prompt 302 may describe a context for generating the GAI image. The first set of feedback parameters 218 may indicate problems identified during the validation of previous GAI images. In some examples, the problems may relate to visual features of the GAI images or features of objects present in the GAI images. In other examples, the problems may be due to biased data in the text prompt used for generating the GAI images. In still other examples, the problems may relate to areas / regions of the GAI images that may be of interest to users.The first set of external parameters 304 may be part of the set of external parameters 220 stored in the data store 208. The set of external parameters 220 may indicate the markup rule set, the refinement parameters, and / or similar.

[0080] Based on the initial text prompt 302, the feedback parameter set 218, and the first external parameter set 304, the prompt enhancer 230 formulates a pre-prompt. The pre-prompt may indicate additional information / a list of criteria (e.g., additional contextual information, visual features, or the like) to be added to the initial text prompt 302. In some examples, if the initial text prompt 302 contains biased data, the additional information / list of criteria may indicate to remove the biased data. An example of a pre-prompt might be, "You are a prompt engineer who understands prompts; here is a prompt for image generation."<invite textuelle initiale> Rewrite it incorporating the following criteria:<liste de critères / informations supplémentaires> "

[0081] The prompt enhancer 230 submits the pre-prompt to an LLM (for example, GPT-4 vision) for processing and receives the enhanced text prompt 306 from the LLM, based on the processing of the pre-prompt.

[0082] Consider an example scenario in which the initial text prompt 302 might indicate "a working lunch with corporate executives." For such an initial text prompt 302, the enhanced text prompt 306 might be provided as "A vivid photograph depicting corporate executives enjoying a productive working lunch in a modern corporate boardroom with abundant natural light, showcasing a harmonious blend of business attire, neutral tones, an energetic atmosphere, and balanced composition." In an example here, the additional context as "a modern corporate boardroom with abundant natural light, showcasing a harmonious mix of professional attire”, and required visual characteristics such as “neutral tones, an energetic atmosphere and a balanced composition” are added as additional information / criteria list to the initial text prompt.

[0083] The prompt enhancer 230 provides the enhanced text prompt 306 to the image and description generator 232. The image and description generator 232 can also obtain the feedback parameter set 218, and a second external parameter set 308 from the data store 208. The second external parameter set 308 can be obtained from the external parameter set 220 in the data store 208. The second external parameter set 308 can include the markup rule set, guidelines, and / or similar elements. The image and description generator 232 uses the GAI image generator 202 to generate the GAI image 310, based on the enhanced text prompt 306, the feedback parameter set 218, and the second external parameter set 308.

[0084] The image and description generator 232 also uses the GAI image description engine 204 to generate the GAI text description 312 for the generated GAI image 310. The GAI text description 312 can describe the features of the generated GAI image 310. In one example, for the generated GAI image 310 (for example, generated using the enhanced text prompt 306), the GAI text description 312 can be generated as "A photograph depicting business executives taking part in a productive working lunch."

[0085] The image and description generator 232 provides the first validator 236 with the generated GAI image 310 and the corresponding GAI text description 312 for the generated GAI image 310. The first validator 236 can also obtain the enhanced text prompt 306, the initial text prompt 302, and the second set of external parameters 308. Based on the GAI text description 312, the enhanced text prompt 306, the initial text prompt 302, and the second set of external parameters 308, the first validator 236 determines whether the generated GAI image 310 is valid ("OK") or not ("KO") with respect to the initial text prompt 302.

[0086] To determine whether the generated GAI image 310 is valid, the first validator 236 performs a semantic comparison of the GAI text description 312 and the enhanced text prompt 306. The first validator 236 further determines the score for a result of the semantic comparison. The score can be determined based on the evaluation of the GAI text description 312 and the enhanced text prompt 306, taking into account the initial text prompt 302 and the second set of external parameters 308. The first validator 236 compares the score to the first threshold 214.

[0087] If the score falls within the first threshold 214, the first validator 236 determines the mismatch between the enhanced text prompt 306 and the GAI text description 312. Consequently, the first validator 236 determines that the generated GAI image 310 is invalid ("KO") with respect to the initial text prompt 302. In such a scenario, the first validator 236 can identify problems in the generated GAI image 310 or problems with the enhanced text prompt 306. In some examples, problems in the generated GAI image 310 may indicate a set of missing criteria in the generated GAI image, such as contextual information, visual features, and / or analogous features. Problems in enhanced text prompt 306 may indicate missing details in the visual features of the generated GAI 310 image or the presence of biased data, or the like.When problems are identified with the generated GAI image 310 or the enhanced text prompt 306, the first validator 236 sends the identified problems as a set of feedback parameters with a rejection signal 314 to the GAI image generator 202 or the prompt enhancer 230 for the regeneration of a new image or the regeneration of a new text prompt, which is described in detail in connection with [Fig.4].

[0088] In an example here, consider that the GAI 312 text description and the enhanced text prompt 306 include "A vivid photograph depicting business executives taking part in a productive working lunch, in a modern corporate boardroom with abundant natural light, showcasing a harmonious mix of business attire, neutral tones, an energetic atmosphere and a balanced composition" and "A photograph depicting business executives taking part in a productive working lunch", respectively.In such a consideration, the first validator 236 determines that the score of the semantic comparison of the GAI text description 312 and the enhanced text prompt 306 is below the first threshold 214, because criteria such as additional contextual information and visual features included in the enhanced text prompt 306 are not present in the GAI text description 312 corresponding to the generated GAI image 310. In such a scenario, the first validator 236 provides the feedback parameter set along with the rejection signal 314 to the image and description generator 232 for the regeneration of a new GAI image. The feedback parameter set can indicate the criteria missing in the generated GAI image 310.

[0089] If the score is greater than the first threshold 214, the first validator 236 determines that the generated GAI image is valid ("OK") with respect to the initial text prompt 302. Then, the first validator 236 can send an acceptance signal 316 to the second validator 238.

[0090] Upon receiving the acceptance signal 316 from the first validator 236, the second validator 238 initiates the validation of the generated GAI image 310. The second validator 238 receives the generated GAI image 310, the second set of external parameters 308, and the list of neuro-aesthetic criteria 210 (which indicates the predetermined elements that must be present in the GAI image). The second validator 238 evaluates the generated GAI image 310 and the list of neuro-aesthetic criteria 210 using the GAI image description engine 204 and determines the number of elements from the list of neuro-aesthetic criteria present in the generated GAI image 310. The second validator 238 compares the number of elements to the second threshold 216.

[0091] If the number of elements does not satisfy the second threshold 216, the second validator 238 determines that the generated GAI image 310 does not contain the predetermined elements on the neuro-aesthetic criteria list 210. Consequently, the second validator 238 determines that the generated GAI image 310 is invalid ("KO") with respect to the initial text prompt. Upon determination, the second validator 238 sends a rejection signal 318 to the prompt enhancer 230 for a further enhancement of the initial text prompt 302. The newly enhanced text prompt 306 is used for the regeneration of the new GAI image. The rejection signal 318 includes the number of missing elements in the generated GAI image as a set of feedback parameters for the further enhancement of the initial text prompt 302.

[0092] In an example here, consider that the generated GAI 310 image is a vivid photograph depicting a businessperson having a productive working lunch in a modern corporate boardroom with abundant natural light, showcasing a harmonious mix of business attire, neutral tones, an energetic atmosphere, and a balanced composition. In such a scenario, the second validator 238 compares the number of items from the list of neuro-aesthetic criteria present in the generated image to the second threshold 216 and identifies, for example, that the number of items is less than the second threshold 216. Consequently, the second validator 238 determines that the generated GAI 310 image is not valid with respect to the initial text prompt 302.Furthermore, the second validator 238 identifies that objects such as "vegetables" specified in the list of neuro-aesthetic criteria 210 for lunch are absent from the generated GAI image 310. Upon identification, the second validator 238 sends the rejection signal 318 to the image and description generator 232 for the regeneration of the new GAI image. The rejection signal 318 can indicate which "vegetables" objects to display in the new GAI image.

[0093] If the number of elements satisfies the second threshold 216, the second validator 238 determines that the generated GAI image is valid with respect to the initial text prompt 302. In addition, the second validator 238 sends an acceptance signal 320 ("OK") to the second validator 238. The acceptance signal indicates successful validation of the generated GAI image 310.

[0094] Upon receiving the acceptance signal 320, the third validator 240 initiates the validation of the GAI image generated by means of the heatmap. The third validator 240 obtains the generated GAI image 310 and submits it to the heatmap generator 206. The heatmap generator 206 uses the CRISP engine to generate the heatmap for the generated GAI image 310. The third validator 240 provides the heatmap 322 to the computing device 102 via the interface module 228 for acceptance or rejection by the user. In response to the provision of the heatmap, the third validator 240 receives a response 324 from the computing device 102 via the interface module 228. The response 324 indicates either acceptance or rejection of the heatmap by the user.

[0095] If the thermal card is rejected, the third validator 240 sends an indication 326 to the computing device 102 via the interface module 228, which allows the user to enter a new initial text prompt / information. The new initial text prompt may contain new information for generating the new GAI image or a modification of the initial text prompt / information 302 initially provided for generating the new GAI image.

[0096] If the thermal map is accepted, the third validator 240 determines that the generated GAI image is valid and sends an indication 328 to the interface module 228 to provide the generated GAI image to the computing device 102 for further use / processing.

[0097] Therefore, with the proposed validation / evaluation of the GAI image, the user can launch the image generation application, provide the initial text prompt to the image generation application, and obtain the GAI image without further intervention. Thus, the "launch and forget" optimization process can be followed for GAI image generation. Furthermore, multiple GAI images are generated in parallel without any time consumption or delay.

[0098] [Fig. 4] is a block diagram showing an example of a first validator 236 for validating the GAI image from text prompt matching according to embodiments of the present invention. The first validator 236 comprises an integration module 402, a similarity calculation module 404, and a comparison module 406.

[0099] The integration module 402 receives the GAI text description 312 corresponding to the generated GAI image 310 and the enhanced text prompt 306 used for generation of the GAI image. The integration module 402 converts the GAI text description 312 into a description vector 410. Similarly, the integration module 402 converts the enhanced text prompt 306 into a prompt vector 408. The GAI text description 312 and the enhanced text prompt 306 can be converted into a description vector 410 and a prompt vector 408, respectively, by way of non-limiting examples, using SIAMOIS-BERT networks, global vector representations, as is known in the art and is not discussed further here. The description vector 410 and the prompt vector 408 can be provided to the similarity calculation module 404.

[0100] The similarity calculation module 404 calculates the score 412 between the description vector 410 and the prompt vector 408, by way of non-limiting example, using a cosine similarity method not detailed here. If the score 412 is high, then the GAI text description 312 and the enhanced text prompt 306 are semantically identical. Thus, the mismatch between the GAI text description 312 and the enhanced text prompt 306 may be low. In some examples, the score 412 may be expressed as a percentile.

[0101] The comparison module 406 compares the score 412 to the first threshold 214. By way of non-limiting example, the first threshold 214 could be 90%. If the score is equal to or greater than the first threshold 214, the comparison module 406 determines that the generated GAI image is valid with respect to the initial text prompt 302. Once the generated GAI image has been successfully validated, the second validator 238 initiates the validation of the generated GAI image based on the list of neuro-aesthetic criteria.

[0102] If the score 412 is below the first threshold 214 within a small range defined by the first predetermined variance (for example, if the score is slightly below the first threshold), the comparison module 406 determines that the generated GAI image 310 is closer to the initial text prompt 302, indicating a sign of failure to generate the GAI image. By way of non-limiting example, the first variance could be 10% (for example, 80% to 90% are within the first threshold 214). The failure to generate the GAI image could be due to randomness in the generation of the GAI image, or to one or more criteria (described in the initial / enhanced text prompt) being missing from the generated GAI image 310, or to the quality of the generated GAI image 310, or the like. Therefore, the generated GAI 310 image is considered invalid with respect to the initial text prompt 302.Furthermore, the regeneration of a new GAI image is initiated without requiring an enhanced text prompt or any user intervention.

[0103] If the score 412 is less than the first threshold 214 within a larger range defined by the second predetermined variance (for example, if the score is less than the first threshold), the comparison module 406 determines that the GAI image 310 The generated GAI image is invalid compared to the initial text prompt 302 due to the corresponding enhanced text prompt 306. Therefore, the issues specified in the generated GAI image can be identified and used to enhance and define the initial text prompt 302 in the enhanced text prompt 306, without requiring any user intervention. The enhanced text prompt 306 can then be used to regenerate a new GAI image.

[0104] Consider an example scenario in which the enhanced text prompt 306, corresponding to the initial text prompt 302 and used for generating the GAI image, contains "generate a soft, furry, round, pink ball," and the GAI text description 312, corresponding to the generated GAI image 310, contains "the image shows a furry ball with smooth textures and attractive colors." In such a scenario, the generated score indicating the similarity between the enhanced text prompt 306 and the GAI text description 312 is 90%, which is equal to the first threshold 214. Therefore, the generated GAI image is considered valid with respect to the initial text prompt 302.

[0105] Let us take another example scenario, in which the enhanced text prompt 306, corresponding to the initial text prompt 302 and used for generating the GAI image, contains "generate a soft, furry, pink, round ball," and the GAI text description 312, corresponding to the generated GAI image 310, contains "the image shows a cat eating an apple." In such a scenario, the generated score indicating the similarity between the enhanced text prompt 306 and the GAI text description 312 is 19%, which is lower than the first threshold 214 of the larger range defined by the second predetermined threshold. Therefore, the generated GAI image 310 is considered invalid with respect to the initial text prompt 302. Then, the improvement and adjustment of the initial text prompt 302 into the improved text prompt 306 are regenerated by targeting a correct component for the regeneration of the GAI image.

[0106] [Fig. 5] is a block diagram showing an example of a second validator 238 for validating the GAI image based on the list of neuro-aesthetic criteria 210 in accordance with embodiments of the present invention. The second validator 238 comprises an element detection module 502, a score calculation module 504, and a comparison module 506.

[0107] The element detection module 502 receives the neuro-aesthetic criteria list 210 and the generated GAI image 310 for validation. The neuro-aesthetic criteria list 210 can specify a total number of elements (in terms of color, objects, shapes, textures, or analogs) that must be present in the generated GAI image. The element detection module 502 submits the neuro-aesthetic criteria list and the generated GAI image 310 to the GAI image description engine 204 and receives a a number of elements 508 from the list of neuro-aesthetic criteria present in the generated GAI 310 image.

[0108] Once the number of elements present in the generated GAI 310 image has been identified, the scoring module 504 calculates a neuro-aesthetic score 510 out of 100, based on the number of elements present in the generated GAI 310 image and the total number of elements that should be present in the generated GAI 310 image.

[0109] The comparison module 506 compares the neuro-aesthetic score 510 to the second threshold 216. If the neuro-aesthetic score 510 is equal to or greater than the second threshold 216, the comparison module 506 determines that the generated GAI image 310 is valid 512 with respect to the initial text prompt 302. Once the generated GAI image 310 has been determined to be valid 512, a subsequent validation step is initiated by the third validator 240.

[0110] If the neuro-aesthetic score 510 is below the second threshold 216, the comparison module 506 determines that the generated GAI image 310 is not valid 514 with respect to the initial text prompt 302. Once the generated GAI image 310 is determined to be invalid 514, the improvement and configuration of the initial text prompt 302 in the improved text prompt 306 are restarted, considering the initial text prompt 302 and the number of items on the neuro-aesthetic criteria list 210 missing in the generated GAI image 310. The improved text prompt can also be used to regenerate a new GAI image. Therefore, the new GAI image is regenerated without any user intervention.

[0111] Consider an example scenario in which the generated GAI image 310 shows a round, hairy ball with a smooth texture and attractive colors, and neuro-aesthetic criteria list 210 specifies that “the colors are pale,” “the objects are round,” and “the texture is smooth.” In such a scenario, the neuro-aesthetic score 510 of the generated GAI image is calculated at 90%, because the generated GAI image 310 does not have pale colors as specified by neuro-aesthetic criteria list 210. Furthermore, since the neuro-aesthetic score 510 is above the second threshold 216 (e.g., 80%), the generated GAI image 310 is determined to be valid with respect to the initial text prompt 302.

[0112] Consider another example scenario, in which the generated GAI 310 image shows a war scene with explicit violence. In such a scenario, the neuro-aesthetic score 510 of the generated GAI 310 image is calculated at 21%, since the list of neuro-aesthetic criteria does not specify explicit violence. Furthermore, because the neuro-aesthetic score 510 is below the second threshold 216 (e.g., 80%), the generated GAI 310 image is determined to be invalid with respect to the initial text prompt.

[0113] [Fig. 6] is a block diagram showing an example of a third validator 240 for validating the GAI image based on the heatmap according to embodiments of the present invention. The third validator 240 comprises a map generation module 602 and a validation module 604.

[0114] The map generation module 602 receives the generated GAI image 310. The map generation module 602 submits the generated GAI image 310 to the heatmap generator 206 and receives the heatmap 606 from the heatmap generator 206. In some examples, the map generation module 602 may also receive heatmap information from the heatmap generator 206. The heatmap information can be derived by performing an analysis on the heatmap 606 to identify how the heatmap conforms to the image layout rule set 212 (e.g., symmetry, object proportions, or analogy).

[0115] The validation module 604 allows the user to validate the heat map 606 and / or the heat map information. If the user validates and accepts the heat map 606 and / or the heat map information, the validation module 604 determines that the generated GAI image 310 is valid 608 with respect to the initial text prompt 302. In such a scenario, the generated GAI image 310 is provided to the computer device 102 / to the user for further processing or use.

[0116] If the user rejects the heatmap 606 and / or heatmap information, the validation module 604 determines that the generated GAI image 310 is not valid 610 with respect to the initial text prompt 302. In such a scenario, the user is allowed to re-enter a new initial text prompt / new information for the generation of a new GAI image.

[0117] [Fig. 7] is a flowchart that presents an example of a computer-implemented method 700 for generating and validating GAI images in accordance with the implementations of the present invention. In some embodiments, the method 700 can be executed by the processor 222 of the computer system 104, as described in relation to [Fig. 2] to 6.

[0118] In step 702, process 700 first involves receiving the initial text prompt / information describing the characteristics of the desired image. In some examples, the characteristics may indicate objects to be present in the desired image, the context associated with the objects, the visual characteristics of the desired image, and / or similar features.

[0119] In step 704, process 700 involves enhancing the information received in the text prompt. The information received can be enhanced in the text prompt by means of any of the following: foundation / LLM models, AI models, ML models, and / or analogous models. In some examples, the The information received can be improved in the text prompt by adding supplementary information / criteria to the received information and / or by considering the Feedback Parameter Set 218 (collected from previous GAI image generation). The supplementary information may include additional contextual information incorporating demographic / sectoral considerations, specific GAI-based and / or analogous keywords, required visual characteristics, unbiased data to be replaced with biased (if applicable in the initial text prompt) and / or analogous data. The Feedback Parameter Set 218 may indicate problems identified in previously generated GAI images.Problems can be identified due to missing criteria / contextual information in GAI images, inappropriate visual characteristics of the GAI image, missing elements in the list of neuro-aesthetic criteria in GAI and / or analogous images.

[0120] In step 706, method 700 includes a first submission of the enhanced text prompt to the GAI image generator 202. In some examples, the feedback parameter set 218 and the external parameter set 220 may be submitted with the enhanced text prompt to the GAI image generator 202. The external parameter set 220 may include the markup rule set, the refinement parameters, and / or the like. In response to the first submission, in step 708, method 700 includes a second reception, from the GAI image generator 202, of the generated GAI image corresponding to the enhanced text prompt.

[0121] In step 710, the process 700 includes the third reception of the GAI image description of the GAI image generated from the GAI image description engine 204. The GAI image description can be generated by the GAI image description engine 204 using foundation / LLM models, computer vision models, and / or the like. In some examples, the enhanced text prompt and optionally the feedback parameter set 218 and the external parameter set 220 can be submitted to the GAI image description engine 204 for the GAI image description corresponding to the generated GAI image. The external parameter set 220 can include the markup rule set, guidelines, and / or the like.

[0122] In step 712, process 700 includes the initial determination of whether the GAI text description sufficiently corresponds to the enhanced text prompt with respect to the first predetermined threshold 214. The initial determination involves performing a semantic comparison of the GAI text description and the enhanced text prompt and generating a score for the result of performing the semantic comparison. The score is evaluated with respect to the first threshold 214.

[0123] When it is determined that the score lies within the first predetermined variance with respect to the first threshold 214, the first determination involves the discovery of the non-match (a first type of non-match 712a) within the first predetermined variance. Upon discovery of such a non-match 712a, the process 700 involves returning to step 706 of the first submission of the enhanced text prompt to the GAI image generator 202 for the regeneration of a new GAI image.

[0124] When it is determined that the score lies within the second predetermined variance (higher than the first predetermined variance) with respect to the first threshold 214, the first determination involves finding the mismatch (a second type of mismatch 712b) within the second predetermined variance. Upon discovering such a 712b mismatch, the procedure 700 involves returning to step 704 of improving and configuring the information based on the improved text prompt and the problems identified with the generated GAI image.

[0125] In response to the fact that the first determination finds that the score is greater than or equal to the first threshold 214, step 714 is performed. In step 714, process 700 involves determining whether the generated GAI image sufficiently corresponds to the list of predetermined neuro-aesthetic criteria 210 with respect to the second threshold 216. The second determination involves requesting the GAI image description engine 204 to identify the number of elements from the list of neuro-aesthetic criteria 210 present in the generated GAI image. Based on the response to the request, the second determination involves determining whether the number of elements present in the generated GAI image satisfies the second threshold 216.

[0126] If the number of elements present in the generated GAI image does not satisfy the second threshold 216, the second determination includes the determination that the generated GAI image does not contain enough elements from the list of neuroaesthetic criteria 210 and, consequently, the discovery that the non-match is below the second threshold 216. In response to the fact that the second determination discovers that the non-match is below the second threshold 216, the process 700 includes returning to step 704 of improvement and configuration of information based on the improved text prompt and the elements of the list of neuroaesthetic criteria not found in the generated GAI image.

[0127] In response to the fact that the second determination finds that the match is greater than or equal to the second threshold 216, step 716 is performed. In step 716, process 700 includes the fourth reception of the thermal map of the generated GAI image. The thermal map can be generated using the CRISP engine.

[0128] In step 718, the process includes 700 activating the thermal map validation. The thermal map can be provided to the user for acceptance or rejection. In response to the rejection of the thermal map, the process 700 includes returning to step 702 for the initial receipt of additional information. In some examples, this information may include new information for generating the new GAI image. In other examples, the information may include a modification of the original information for generating the new GAI image.

[0129] In response to at least one combination of the fact that the first determination finds that the GAI text description sufficiently matches the enhanced text prompt, that the second determination finds that the generated GAI image sufficiently matches the neuro-aesthetic criteria list 210, and the acceptance of the generated heat map, in step 720, the process 700 involves transmitting the generated GAI image for further use and / or further processing. The transmitted GAI image complies with the predetermined guidelines and thresholds.

[0130] Embodiments of the present invention provide technical solutions to numerous technical problems that arise in the context of traditional methods for generating and validating / evaluating GAI images. With the proposed methodology, acceptable GAI images are obtained with fewer resubmission and revision loops. Consequently, the overall energy consumption required to generate acceptable GAI images is reduced.

[0131] Furthermore, the proposed validation (based on the use of text description / prompt matching (first determination) and the list of neuroaesthetic criteria (second determination)) allows for computer processing as a substitute for human subjective preferences. Therefore, a separate computerized process is implemented to at least partially automate the validation of the generated GAI images, which was previously performed manually through continuous resubmission and review loops. This process reduces the number of user interventions required until the generated GAI images are accepted, thereby reducing the time needed to generate acceptable GAI images.

[0132] In addition, the proposed validation of GAI images generated on the basis of the heatmap provides a more detailed form of information to the user to evaluate the GAI image and provide specific feedback to the modification of the heatmap rather than the GAI image itself.

[0133] [Fig.8A]-8B, 9A-9B, 10A-10B and 11A-11B represent the generation of images acceptable GAI in accordance with the embodiments of the present invention.

[0134] Consider an example scenario, such as that shown in [Fig. 8A], in which an initial text prompt "cosmic science fiction diorama of a quasar The initial prompt, "and jellyfish in a resin cube," is received for generating a GAI image. However, the probability of accepting the GAI image generated using the initial text prompt is very low. Therefore, the proposed methodology improves the initial text prompt by adding further criteria, including additional contextual information, an image type, and visual characteristics. For example, the improved text prompt includes "a visually stunning work of science fiction using mixed media techniques, such as acrylic painting and digital manipulation, to create a cosmic diorama featuring a vivid quasar and ethereal jellyfish suspended in a resin cube. Depict a futuristic space station, surrounded by an awe-inspiring nebula, with shimmering neon lights casting a mesmerizing glow."Use a combination of electric blues, intense purples, and neon greens to create a vivid and otherworldly color palette. Imbue the work with a sense of wonder and mystery, evoking an atmosphere of both excitement and intrigue. Arrange the composition to highlight the intricate details of the quasar's energized beams intertwining with the graceful tentacles of the mystical Medusa, capturing the viewer's imagination. With such an enhanced text prompt, the GAI image that may be acceptable to the user is generated, as shown in [Fig. 8B].

[0135] Consider another example scenario, as depicted in [Fig. 9A], in which an initial text message, "a working lunch with corporate executives," is received for the generation of a GAI image. However, the probability of accepting the GAI image generated using the initial text prompt is very low. Therefore, the proposed methodology enhances the initial text prompt by adding further criteria, including additional contextual information and visual features. For example, the enhanced text prompt includes "A vivid photograph depicting corporate executives enjoying a productive working lunch in a modern corporate boardroom with abundant natural light, showcasing a harmonious blend of business attire, neutral tones, an energetic atmosphere, and a balanced composition."With such an improved text prompt, the GAI image that may be acceptable to the user is generated, as shown in [Fig.9B].

[0136] Consider yet another example scenario, as depicted in [Fig. 1OA], in which an initial text prompt, "two elderly people holding hands on a bench," is received for generating a GAI image. However, the probability of accepting the GAI image generated using the initial text prompt is very low. Therefore, the proposed methodology improves the initial text prompt by removing the biased data "elderly people" and adding additional criteria, including further contextual information, a type image and visual characteristics. For example, the enhanced text prompt includes "A serene oil painting depicting two companions of different age groups holding hands, sitting on an aged wooden bench in the middle of a flower garden in the soft glow of a golden sunset." With such an enhanced text prompt, the GAI image that may be acceptable to the user is generated, as shown in [Fig. 1OB].

[0137] Consider yet another example scenario, as depicted in [Fig. 1 IA], in which an initial text prompt, “Hyper-realistic portrait photo of a man looking into the camera, smiling slightly, calm features, resembles an advertising executive,” is received for generating a GAI image. However, the probability of accepting the GAI image generated using the initial text prompt is very low. Therefore, the proposed methodology improves the initial text prompt by removing the biased “man” data and adding additional criteria, including further contextual information, image type, and visual characteristics.For example, the enhanced text prompt includes "Create a hyper-realistic portrait photo of a calm advertising executive in a modern office environment with natural lighting, highlighting subtle smiles, warm colors, a composed mood, and a strong central composition." With such an enhanced text prompt, the GAI image that may be acceptable to the user is generated, as shown in [Fig. 1 IB].

[0138] [Fig. 12] illustrates a 1200 computer system that can be used to put into The 700 process is implemented as described in relation to [Fig. 7]. Specifically, computing devices such as desktop computers, laptops, smartphones, tablets, and wearable devices can be used for generating and validating GAI images and may have the structure of the 1200 computing system. The 1200 computing system may include additional components not shown, and some of the described process components may be removed and / or modified. As another example, the 1200 computing system may be deployed on external cloud platforms such as the cloud, internal enterprise cloud computing clusters, organizational computing resources, and / or similar resources.

[0139] The computer system 1200 comprises one or more processors 1202, such as a central processing unit, an ASIC, or another type of processing circuit; input / output devices 1204, such as a monitor, mouse, keyboard, etc.; a network interface 1206, such as a local area network (LAN), an 802.11x wireless LAN, a 3G or 4G mobile WAN, or a WiMAX WAN; and a computer-readable medium 1208. Each of these components can be functionally coupled to a bus 1210. The computer-readable medium 1208 can be any suitable medium participating in the provision of instructions programmed to cooperate with the processor(s) 1202 to carry out the computer-implemented process 700. For example, the computer-readable medium 1208 may be a non-transient or non-volatile medium, such as a magnetic disk or solid-state non-volatile memory, or a volatile medium such as RAM. The instructions or modules stored on the computer-readable medium 1208 may include machine-readable instructions 1212 executed by the processor(s) 1202 that cause the processor(s) 1202 to carry out the process 700 and functions of the computer system 104.

[0140] The process 700 can be implemented as software stored on a non-transient, processor-readable medium and executed by the processors 1202. For example, the computer-readable medium 1208 can store an operating system 1214, such as MAC OS, MS WINDOWS, UNIX, or LINUX, and code for implementing the process 700. The operating system 1214 can be multi-user, multiprocessing, multitasking, multi-threaded, real-time, and similar. For example, while running, the operating system 1214 is executed and the code implementing the process 700 is executed by the processor(s) 1202.

[0141] The computer system 1200 may include a data storage 1216, which may include non-volatile data storage. The data storage 1216 stores all data used or generated by the process 700.

[0142] The network interface 1206 connects the computer system 1200 to internal systems, for example via a LAN. In addition, the network interface 1206 can connect the computer system 1200 to the Internet. For example, the computer system 1200 can connect to web browsers and other external applications and systems via the network interface 1206.

[0143] What has been described and illustrated herein is an example, along with some variations thereof. The terms, descriptions, and figures used in this document are given by way of illustration only and are not considered limitations. Numerous variations are possible in the spirit and scope of the subject matter, which is intended to be defined by the following claims and their equivalents.

[0144] The implementations and all functional operations described in this specification can be carried out in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures described in this specification and their structural equivalents, or in combinations of one or more of them. Implementations can be carried out in the form of one or more computer program products (i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by a data processing device or for controlling its operation). (ci) Computer-readable media can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition that carries out a machine-readable propagated signal, or a combination of one or more of these. The term computer system encompasses all data-processing devices, appliances, and machines, including, for example, a programmable processor, a computer, or multiple processors or computers. The device may include, in addition to the hardware, code that creates an execution environment for the computer program in question (for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or any suitable combination of one or more of these).A propagated signal is an artificially generated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal) that is generated to encode information for transmission to a suitable receiving device.

[0145] A computer program (also known as a program, software, software application, script, or code) may be written in any suitable form of programming language, including compiled or interpreted languages, and may be deployed in any suitable form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that contains other programs or data (for example, one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in several coordinated files (for example, files that store one or more modules, subroutines, or portions of code).A computer program can be deployed to run on one or more computers located on one site or distributed across multiple sites and interconnected by a communication network.

[0146] The processes and logic flows described in this specification can be executed by one or more programmable processors running one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be executed by, and the device can also be implemented as, special logic circuitry (for example, an FPGA (user-programmable pre-diffused array) or an ASIC (application-specific integrated circuit).

[0147] Processors suitable for running a computer program include, by way of example, both general-purpose and special-purpose microprocessors, and any one or more processors of any suitable type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory, or both. The components of a computer may include a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes, or is functionally coupled to, one or more mass storage devices for storing data (for example, magnetic or magneto-optical disks or optical disks). However, a computer need not possess such devices.Furthermore, a computer can be integrated into another device (e.g., a mobile phone, a personal digital assistant (PDA), a mobile audio player, a global positioning system (GPS) receiver). Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, storage media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard drives or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM discs. The processor and memory may be complemented by, or incorporated into, special logic circuitry.

[0148] To enable interaction with a user, implementations can be carried out on a computer having a display device (e.g., a CRT (cathode ray tube), an LCD (liquid crystal display)) to display information to the user, and a keyboard and a pointing device (e.g., a mouse, a trackball, a touchpad), by means of which the user can provide input to the computer. Other types of devices can also be used to ensure interaction with a user; for example, the feedback provided to the user can be any appropriate form of sensory feedback (e.g., visual feedback, auditory feedback, tactile feedback); and the user input can be received in any appropriate form, including acoustic, vocal, or tactile input.

[0149] Implementations can be carried out in a computer system that includes a background component (e.g., as a data server), a middleware component (e.g., an application server), and / or a front-end component (e.g., a client computer with a graphical user interface or a web browser, through which a user can interact with an implementation (implementation), or any suitable combination of one or more such back-end, middleware, or front-end components. The system components may be interconnected by any suitable form or means of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0150] A computer system may include clients and servers. A client and a server are generally located far apart and usually interact via a communication network. The client-server relationship is established through computer programs running on the respective computers and having a client-server relationship with each other.

[0151] Although this specification contains many specifics, these should not be interpreted as limitations on the scope of disclosure or what may be claimed, but rather as descriptions of features specific to particular implementations. Certain features described in this specification as part of separate implementations may also be implemented in combination in a single implementation. Conversely, various features described as part of a single implementation may also be implemented in several implementations separately or in any appropriate sub-combination.Furthermore, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features of a claimed combination may in some cases be detached from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0152] Similarly, if the operations are shown in the drawings in a particular order, this should not be understood as requiring that these operations be performed in the particular order shown or in a sequential order, or that all the illustrated operations be performed, to obtain desired results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the implementations described above should not be understood as requiring such separation in all implementations, and it should be understood that the program components and systems described can generally be integrated together in a single software product or packaged in several software products.

[0153] A number of implementations have been described. However, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. For example, different forms of the flows illustrated above may can be used, with steps reordered, added or deleted. Therefore, other implementations fall within the scope of the following claims.

Claims

Demands

1. A computer-implemented method comprising: first receiving information describing features of a desired image (702); improving the received information into a text prompt (704); first submitting the improved text prompt (706) to a generative artificial intelligence (GAI) image generator (202); second receiving a generated GAI image corresponding to the improved text prompt (708) from the GAI image generator; third receiving a GAI text description of the generated GAI image (710) from a GAI image description engine (204); first determining whether the GAI text description sufficiently matches the improved text prompt with respect to a first predetermined threshold (214);in response to the fact that the first determination discovers a non-match in a first predetermined variance with respect to the first predetermined threshold, the return to the first submission; in response to the fact that the first determination discovers a non-match in a second predetermined variance with respect to the first predetermined threshold (214), the return to the improvement and configuration of information based on the improved text prompt and problems identified with the generated GAI image, the second predetermined variance being greater than the first predetermined variance; the second determination (714) if the generated GAI image sufficiently matches a list of predetermined neuro-aesthetic criteria (210) with respect to a second predetermined threshold (216);in response to the fact that the second determination discovers a non-match below the second predetermined threshold, the return to the improvement and configuration of information based on the improved text prompt and elements of the list of predetermined neuroaesthetic criteria not found in the generated GAI image; the fourth reception of a heatmap of the generated GAI image (716); in response to rejection of the heatmap, return to the first reception for further information; and transmission of the generated GAI image (720) for further use and / or further processing in response to at least one combination of the first determination finding that the GAI text description sufficiently matches the enhanced text prompt, the second determination finding that the generated GAI image sufficiently matches the list of predetermined neuro-aesthetic criteria, and acceptance of the generated heatmap.

2. A method according to claim 1, wherein the first determination comprises: performing a semantic comparison of the GAI textual description and the enhanced textual prompt; scoring a performance result to generate a score; and evaluating the score against the first predetermined threshold.

3. A method according to claim 2, wherein, in response to the fact that the first determination finds that the non-match is in the first predetermined variance, the method further comprises: determining that the score is in the first variance with respect to the first predetermined threshold.

4. A method according to claim 2, wherein, in response to the fact that the first determination finds that the non-match is in the second predetermined variance, the method further comprises: determining that the score lies beyond the first variance relative to the first predetermined threshold.

5. A method according to claim 1, wherein the second determination of whether the GAI image sufficiently corresponds to the list of predetermined neuro-aesthetic criteria further comprises: requesting the GAI image description engine to identify a number of items from the list present in the generated GAI image; and the determination, based on a response to the request, whether the number of elements present in the generated GAI image satisfies the second predetermined threshold.

6. A method according to claim 5, wherein, in response to the fact that the second determination finds that the non-match is less than the second predetermined threshold, the method further comprises: determining whether the generated GAI image does not include enough elements from the list.

7. Method according to claim 1, wherein the fourth reception of the thermal map includes the processing of the GAI image generated with a CRISP engine.

8. A system comprising: a memory storing instructions; and a processor programmed to cooperate with the instructions to perform operations comprising: the first reception of information describing features of a desired image (702); the enhancement of the received information into a text prompt (704); the first submission, to a generative artificial intelligence (GAI) image generator (202), of the enhanced text prompt (706); the second reception, from the GAI image generator, of a generated GAI image corresponding to the text prompt (708); the third reception, from a GAI image description engine (204), of a GAI text description of the generated GAI image (710); the first determination (712) whether the GAI text description sufficiently matches the enhanced text prompt with respect to a first predetermined threshold (214);in response to the fact that the first determination discovers a non-match in a first predetermined variance with respect to the first predetermined threshold, the return to the first submission; in response to the fact that the first determination discovers a non-match in a second predetermined variance with respect to the first predetermined threshold (214), the return to the improvement and configuration of information based on the text prompt; improved and problems identified with the generated GAI image, the second predetermined variance being greater than the first predetermined variance; the second determination (714) whether the generated GAI image sufficiently matches a list of predetermined neuro-aesthetic criteria (210) against a second predetermined threshold (216); in response to the fact that the second determination discovers a non-match below the second predetermined threshold, the return to the improvement and configuration of information based on the improved text prompt and items from the list of predetermined neuro-aesthetic criteria not found in the generated GAI image; the fourth receipt of a heatmap of the generated GAI image (716); in response to the rejection of the heatmap, the return to the first receipt for more information;and the transmission of the generated GAI image (720) for further use and / or further processing in response to at least one combination of the first determination finding that the GAI text description sufficiently matches the enhanced text prompt, the second determination finding that the generated GAI image sufficiently matches the list of predetermined neuro-aesthetic criteria, and acceptance of the generated heat map.

9. System according to claim 8, wherein the first determination comprises: performing a semantic comparison of the GAI textual description and the enhanced textual prompt; scoring a performance outcome to generate a score; and evaluating the score against the first predetermined threshold.

10. System according to claim 9, wherein, in response to the fact that the first determination finds that the non-match is in the first predetermined variance, the system further comprises: the determination that the score is in the first variance with respect to the first predetermined threshold.

11. A system according to claim 9, wherein, in response to the fact that the first determination discovers that the non-match is in the second predetermined variance, the system further includes: the determination that the score lies beyond the first variance relative to the first predetermined threshold.

12. System according to claim 8, wherein the second determination whether the GAI image sufficiently corresponds to the list of predetermined neuro-aesthetic criteria further comprises: requesting the GAI image description engine to identify a number of items from the list present in the generated GAI image; and determining, from a response to the request, whether the number of items present in the generated GAI image satisfies the second predetermined threshold.

13. System according to claim 12, wherein, in response to the fact that the second determination finds that the non-match is less than the second predetermined threshold, the system further comprises: the determination that the generated GAI image does not include enough elements of the list.

14. System according to claim 8, wherein the fourth thermal map reception includes processing the GAI image generated with a CRISP engine.

15. Non-transient, computer-readable media storing instructions which, when executed by computer hardware in combination with software, perform operations comprising: first receiving information describing features of a desired image (702); improving the received information into a text prompt (704); first submitting the text prompt (706) to a generative artificial intelligence (GAI) image generator (202); second receiving a generated GAI image from the GAI image generator corresponding to the text prompt (708); third receiving a GAI image description engine (204) from a GAI image description engine of the generated GAI image (710);

16. the first determination (712) whether the GAI textual description corresponds sufficiently to the enhanced textual prompt with respect to a first predetermined threshold (214); in response to the fact that the first determination discovers a non-match in a first predetermined variance with respect to the first predetermined threshold, the return to the first submission; in response to the fact that the first determination discovers a non-match in a second predetermined variance with respect to the first predetermined threshold (214), the return to the improvement and configuration of information based on the improved text prompt and problems identified with the generated GAI image, the second predetermined variance being greater than the first predetermined variance; the second determination (714) whether the generated GAI image corresponds sufficiently to a list of predetermined neuro-aesthetic criteria (210) in relation to a second predetermined threshold (216); in response to the fact that the second determination discovers a non-match below the second predetermined threshold, the return to the improvement and configuration of information based on the improved text prompt and elements of the list of predetermined neuroaesthetic criteria not found in the generated GAI image; the fourth receipt of a heat map of the generated GAI image (716); In response to the rejection of the thermal card, return to the initial receipt for more information; and the transmission of the generated GAI image (720) for further use and / or further processing in response to at least one combination of the first determination finding that the GAI text description sufficiently matches the enhanced text prompt, the second determination finding that the generated GAI image sufficiently matches the list of predetermined neuro-aesthetic criteria, and acceptance of the generated heat map. Non-transient computer-readable media according to claim 15, wherein the first determination comprises: performing a semantic comparison of the GAI textual description and the enhanced textual prompt; the scoring of a performance result to generate a score; and the evaluation of the score against the first predetermined threshold.

17. Non-transient computer-readable media according to claim 16, wherein, in response to the fact that the first determination discovers a non-match in the first predetermined variance, the method further comprises: determining that the score lies in the first variance with respect to the first predetermined threshold.

18. Non-transient computer-readable media according to claim 16, wherein, in response to the fact that the first determination discovers a non-match in a second predetermined variance, the method further comprises: determining that the score lies beyond the first variance with respect to the first predetermined threshold.

19. Non-transient computer-readable media according to claim 15, wherein the second determination whether the GAI image sufficiently corresponds to the list of predetermined neuro-aesthetic criteria further comprises: requesting the GAI image description engine to identify a number of items from the list present in the generated GAI image; and determining, from a response to the request, whether the number of items present in the generated GAI image satisfies the second predetermined threshold.

20. Non-transient computer-readable media according to claim 19, wherein, in response to the fact that the second determination discovers a non-match below the second predetermined threshold, the method further comprises: the determination that the generated GAI image does not include enough elements of the list.

Citation Information

Patent Citations

  • Visual and digital content optimization

    US10957086B1