AI agent-based text drawing method, system and device
By analyzing and adjusting the attention range of image elements in the process of artificial graphics based on AI agents, the problem of the neglected auxiliary elements during the iteration process is solved, and the quality and aesthetics of the generated images are improved.
Patent Information
- Application Number
- CN202510625525.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the process of textual production based on AI agents, the auxiliary elements that are not input by users in the text description input by users may be ignored during the iteration process, resulting in the proportion of auxiliary elements in the generated image being too low, affecting the quality and aesthetics of the image.
By obtaining the input text and generating the original image, quality analysis is performed to determine the reference image. During the iteration, the attention of the auxiliary element is adjusted according to the attention range of the element, ensuring that the importance of the auxiliary element in the image is reasonably reflected.
It effectively improves the quality and beauty of the generated images, ensures that the proportion of auxiliary elements in the picture is appropriate, and improves the efficiency and effect of literary and artistic images.
Smart Images

Figure CN120147483A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a text-to-image method, system and device based on an AI agent. Background Art
[0002] An AI agent refers to an artificial intelligence system that can perceive the environment, make decisions and execute tasks. The text-to-image method based on an AI agent is a technology that uses an artificial intelligence system to convert a natural language text description into an image. The main process of text-to-image based on an AI agent is as follows: The user inputs a text description, the AI agent model analyzes the text content input by the user, extracts the image features to be generated from the text input by the user, and the generator uses the image features to create an image, and multiple rounds of iterative optimization are required to ensure that the generated image is as consistent as possible with the text description.
[0003] Currently, when a user inputs text for text-to-image, the output results often contain content that is not input by the user (referred to as auxiliary elements), which is usually used to improve and enrich the picture details in the image. However, in the iterative process, these non-user input contents may be ignored due to the main content corresponding to the text that the user mainly biases towards during adjustment, resulting in too low a proportion of non-user input contents after multiple iterations, resulting in insufficient other contents for improving the picture in the generated image, and failing to achieve the desired effect of the user, which requires the user to further adjust and feedback, reducing the efficiency of text-to-image. In some solutions, the method of setting a fixed weight threshold is used to limit the minimum proportion of non-user input contents. However, since the required proportion of non-user input contents in different pictures is not the same, setting a fixed threshold may affect the quality and beauty of the picture, and the generation effect is poor. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a text-to-image method, system and device based on an AI agent, and the specific technical solutions adopted are as follows: In the first aspect, an embodiment of the present application provides a text-to-image method based on an AI agent, including: Obtain an input text composed of multiple prompt words, and generate a plurality of original images according to the input text and the AI agent; Perform quality analysis on the original images, determine a reference image according to the quality analysis results and the plurality of original images, and use the reference image as a reference image for the next image generation, and perform several iterations to determine a plurality of original images and reference images with different generation times; According to the plurality of original images with different generation times, determine the attention degree of each element in each original image of each generation time, where the attention degree represents the importance degree of the element in the original image; At each iteration, according to the attention of each element in each original image for each generation number and the reference image for each generation number, determine the attention range of each auxiliary element for each generation number, and generate a target image according to the attention range.
[0005] In one implementation, the determining the attention of each element in each original image for each generation number according to a plurality of original images with different generation numbers includes: Process each original image with different generation numbers through a preset model to determine several separated element regions in each original image with different generation numbers, and determine the semantic vector of each prompt; Determine the first cosine similarity between the semantic vectors of different prompts, and determine the central coefficient of each prompt according to the first cosine similarity and the number of prompts; According to the central coefficient and several separated element regions in each original image with different generation numbers, determine the attention of each element in each original image for each generation number.
[0006] In one implementation, the determining the attention of each element in each original image for each generation number according to the central coefficient and several separated element regions in each original image with different generation numbers includes: According to each element region, determine the corresponding feature vector of each element, and respectively determine the second cosine similarity between each feature vector and the semantic vectors of each prompt, and use the element with the largest second cosine similarity as the dominant element to obtain the dominant element corresponding to each prompt and the auxiliary elements other than the dominant element in each original image with different generation numbers; Respectively modify the central coefficient corresponding to the auxiliary element in each original image with different generation numbers to a first preset value; Determine the area ratio of each element region in each original image with different generation numbers to the corresponding original image, and respectively determine the sum value of each central coefficient and a second preset value, and respectively determine the attention of each element in each original image for each generation number according to the first product of the sum value and the area ratio and a normalization function.
[0007] In one implementation, the determining the attention range of each auxiliary element for each generation number at each iteration according to the attention of each element in each original image for each generation number and the reference image for each generation number includes: At each iteration, according to the attention degrees of each element in each original image for each generation time and the reference image for each generation time, determine the trend coefficients of each element in the reference image for each generation time; According to the trend coefficients of each element in the reference image for each generation time, determine the occupancy coefficients of each auxiliary element for each generation time; According to the occupancy coefficients, determine the attention degree ranges of each auxiliary element for each generation time.
[0008] In one implementation manner, the determining the trend coefficients of each element in the reference image for each generation time according to the attention degrees of each element in each original image for each generation time and the reference image for each generation time includes: Respectively determine the corresponding coefficients between each element in each original image for any generation time and each element in the reference image of the previous generation time of this generation time, and according to the corresponding coefficients, determine the elements in the reference image that each element in each original image respectively corresponds to; Respectively determine the variances of the attention degrees of each element in all original images for each generation time, and respectively determine the first differences between the attention degrees of each element in each original image for any generation time and the attention degrees of the elements that the elements in the reference image of the previous generation time of this generation time respectively correspond to; Calculate the arithmetic mean according to the first differences and the number of original images, and respectively determine the trend coefficients of each element in the reference image for each generation time according to the ratio of the arithmetic mean to the variance and the proportional normalization function.
[0009] In one implementation manner, the determining the occupancy coefficients of each auxiliary element for each generation time according to the trend coefficients of each element in the reference image for each generation time includes: Determine the maximum second cosine similarity between the feature vectors corresponding to each auxiliary element in the reference image for each generation time and the semantic vectors of each prompt, and determine the dominant element corresponding to the prompt corresponding to the maximum second cosine similarity as the related element of this auxiliary element; Respectively determine the differences between the trend coefficients of each auxiliary element in the reference image for each generation time and the trend coefficients of the related elements of this auxiliary element, and according to the differences and the proportional normalization function, determine the occupancy coefficients of each auxiliary element for each generation time.
[0010] In one implementation manner, the determining the attention degree ranges of each auxiliary element for each generation time according to the occupancy coefficients includes: Respectively determine the original maximum attention degree and the original minimum attention degree of each auxiliary element for each generation time; Determine the second difference between the original maximum attention value and the original minimum attention value, and determine the second product of the second difference and the occupancy coefficient of the auxiliary element; When the occupancy coefficient of the auxiliary element is greater than the first preset value, determine the sum of the second product and the original minimum attention value as the adjusted original minimum attention value; When the occupancy coefficient of the auxiliary element is less than the first preset value, determine the sum of the second product and the original maximum attention value as the adjusted original maximum attention value; In each iteration, according to the adjusted original minimum attention value and the adjusted original maximum attention value, determine the attention range of each auxiliary element for each generation number.
[0011] In one implementation manner, the generating the target image according to the attention range includes: According to the attention range, control the attention of each auxiliary element in the iteration process, and continuously determine the attention span between the maximum attention and the minimum attention of the auxiliary element in each iteration; wherein, the attention range is re-determined after each iteration; When the attention span of the auxiliary element is less than the preset stability threshold, use the current attention range as the target attention range of the auxiliary element until the target attention ranges of all auxiliary elements are determined; Generate the target image according to the target attention ranges of all auxiliary elements.
[0012] In a second aspect, an embodiment of the present application provides a text-to-image system based on an AI agent, including: An acquisition module, configured to acquire an input text composed of multiple prompt words, and generate a plurality of original images according to the input text and the AI agent; An analysis module, configured to perform quality analysis on the original images, determine a reference image according to the quality analysis result and the plurality of original images, and use the reference image as a reference image for the next image generation, perform several iterations, and determine a plurality of original images and reference images for different generation numbers; A determination module, configured to determine the attention of each element in each original image for each generation number according to the plurality of original images for different generation numbers, where the attention characterizes the importance of the element in the original image; A generation module, configured to, in each iteration, determine the attention range of each auxiliary element for each generation number according to the attention of each element in each original image for each generation number and the reference image for each generation number, and generate a target image according to the attention range.
[0013] In a third aspect, an embodiment of the present application provides a text-to-image generation device based on an AI agent, including: a processor and a memory. Instructions are stored in the memory and loaded and executed by the processor to implement the method in any one of the above aspects.
[0014] The present invention has the following beneficial effects: By obtaining an input text composed of multiple prompt words, generating a number of original images according to the input text and the AI agent, performing quality analysis on the original images, determining a reference image according to the quality analysis result and the number of original images, and using the reference image as the reference image for the next image generation, performing several iterations, determining a number of original images and reference images with different generation times, determining the attention of each element in each original image for each generation time according to the number of original images with different generation times, and at each iteration, determining the attention range of each auxiliary element in each original image for each generation time according to the attention of each element in each original image for each generation time and the reference image for each generation time, considering the attention indicating the importance of the representative element in the original image to adjust the attention range of the auxiliary element, and generating a target image according to the attention range, which is beneficial to ensuring the quality and beauty of the target image and improving the generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is a schematic flowchart of the steps of a text-to-image generation method based on an AI agent provided by an embodiment of the present invention; Figure 2 It is a schematic diagram showing the change of the attention of auxiliary elements during the iteration provided by an embodiment of the present invention; Figure 3 It is a structural block diagram of a text-to-image generation system based on an AI agent provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in conjunction with the accompanying drawings and preferred embodiments, elaborate in detail on a text-to-image generation method, system, and device based on an AI agent according to the present invention, including its specific implementation manner, structure, features, and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0019] It should be noted that the "exemplary" in the embodiments of the present application refers to examples listed for convenience of explanation, and other embodiments are not limited to the listed examples.
[0020] It should be noted that to ensure the meaningfulness of the calculation results, in the fractional operations in the embodiments of the present invention, when encountering the situation where the denominator is 0, a tuning parameter factor greater than 0 needs to be added to the denominator to prevent the denominator from being 0. The value of the tuning parameter factor is set by the implementer according to the actual situation, and no special restrictions are imposed in this application.
[0021] The following will specifically describe the specific solutions of a text-to-image generation method, system, and device based on an AI agent provided by the present invention in conjunction with the accompanying drawings.
[0022] Please refer to Figure 1 , which shows the flowchart of the text-to-image generation method based on an AI agent provided by an embodiment of the present invention. The text-to-image generation method based on an AI agent can at least include steps S100 - S400: S100. Obtain an input text composed of multiple prompt words, and generate a plurality of original images according to the input text and the AI agent.
[0023] S200. Perform quality analysis on the original images, determine a reference image according to the quality analysis results and the plurality of original images, and use the reference image as the benchmark image for the next image generation, and perform several iterations to determine a plurality of original images and reference images with different generation times.
[0024] S300. Determine the attention degree of each element in each original image for each generation time according to the plurality of original images with different generation times.
[0025] Among them, the attention degree represents the importance degree of the element in the original image. The higher the attention degree, the higher the importance degree.
[0026] S400. Based on the attention of each element in each original image for each generation time and the reference image for each generation time, at each iteration, determine the attention range of each auxiliary element in each original image for each generation time, and generate a target image according to the attention range.
[0027] In the technical solution of the embodiment of the present application, by obtaining an input text composed of multiple prompt words, generating a plurality of original images according to the input text and the AI agent, performing quality analysis on the original images, determining a reference image according to the quality analysis result and the plurality of original images, and using the reference image as the benchmark image for the next image generation, performing several iterations, determining a plurality of original images and reference images with different generation times, determining the attention of each element in each original image for each generation time according to the plurality of original images with different generation times, based on the attention of each element in each original image for each generation time and the reference image for each generation time, at each iteration, determine the attention range of each auxiliary element in each original image for each generation time, adjust the attention range of the auxiliary element by considering the attention indicating the importance of the representative element in the original image, and generate a target image according to the attention range, which is beneficial to ensuring the quality and beauty of the target image and improving the generation effect.
[0028] In one implementation manner, in step S100, the user can perform an input operation based on their own needs to generate an input text. Usually, the input text is composed of multiple prompt words. The input input text can be obtained by the AI agent, and then the AI agent analyzes the input text to generate a plurality of original images. Optionally, the analysis process of the AI agent may include: the text understanding module of the AI agent understands the input text, for example, uses the BERT model based on the Transformer architecture to parse the input text, converts the natural language input text into a text encoding vector that can be understood by the machine, and performs word segmentation annotation and text cleaning to extract each semantic vector for generating an image. Then, the image generation module of the AI agent uses the generator of the diffusion model Stable Diffusion or the generative adversarial network GAN to generate a plurality of images through these semantic vectors, denoted as original images.
[0029] In one implementation, in step S200, through the AI agent evaluation and optimization module, the CLIP model or the GAN discriminator is used to perform quality analysis on the original images, evaluate the quality scores of the generated images (i.e., the quality analysis results). For example, taking the structural consistency score (whether the original image conforms to the content of the input text) as the quality score, then according to the high and low of the structural consistency score, several original images with the highest quality scores are presented to the user. For example, taking the presentation of 10 original images as an example, it can be adjusted according to the specific situation. Finally, the user selects one of the 10 presented original images as the reference image, thereby determining the reference image. It should be noted that this reference image is used as the benchmark image for the next image generation, and several iterations are performed to determine several original images and reference images with different generation times; among them, after each selection of the reference image by the user, new supplementary text can be input to add, delete, or adjust the content of the picture, and then several iterations are performed to determine several original images and reference images with different generation times, which is not specifically limited.
[0030] It should be noted that during the iteration process of text-to-image generation, due to the main content of the text that the user mainly focuses on when making adjustments, the non-user input content may be ignored, resulting in insufficient other content (auxiliary elements) for perfecting the picture in the generated original images, failing to achieve the desired effect of the user and reducing the efficiency of text-to-image generation. Therefore, the present invention analyzes the change situation of the non-user input content in the images generated during the iteration optimization process, restricts the change degree of the non-user input content, and obtains a relatively stable attention change range to ensure the picture effect of the non-user input content and ensure the richness of the picture.
[0031] Among them, for the text initially input by the user, the same element may be involved between different prompt words. For example, if the user inputs the text "mountains and waters, distant mountains, clouds and mists in the mountains", these three prompt words are all related to the element of mountains, indicating that the user has a relatively high attention to this element. When generating an image, the mountains should be made the main body with a relatively large picture proportion. Therefore, it is necessary to analyze the semantic relevance between the prompt words.
[0032] In one implementation, step S300 includes steps S301 - S303: S301. Process each original image with different generation times through a preset model, determine several separated element regions in each original image with different generation times, and determine the semantic vectors of each prompt word.
[0033] Optionally, the preset model includes a pre-trained YOLO neural network model. Each original image with different generation times is input into the pre-trained YOLO neural network model, and the pre-trained YOLO neural network model separates the elements in each original image, so as to determine several separated element regions in each original image with different generation times. Moreover, all the prompt words input by the user are transformed into the vector space, so as to determine the semantic vector of each prompt word.
[0034] S302. Determine the first cosine similarity between the semantic vectors of different prompt words, and determine the central coefficient of each prompt word according to the first cosine similarity and the number of prompt words.
[0035] Optionally, determine the first cosine similarity between the semantic vectors of different prompt words , be the cosine similarity between the semantic vectors corresponding to the th prompt word and the th prompt word; and according to the first cosine similarity and the number of prompt words , determine the central coefficient of each prompt word. The specific formula is: where is the central coefficient of the th prompt word, reflecting the correlation degree of each prompt word with other prompt words. When a prompt word has a high correlation degree with other prompt words, the cosine similarity between the vector corresponding to this prompt word and the vectors corresponding to other prompt words is relatively large, and the higher the central coefficient; represents the normalization function, specifically the linear normalization function.
[0036] It should be noted that the central coefficient is obtained based on the semantic correlation between prompt words, but semantic correlation does not fully represent the proportion of the corresponding elements in the picture. It is necessary to further combine the central coefficient with the proportion of all elements in the picture in each generated original image to obtain the attention of each element in each image. By comparing the cosine similarity between the feature vectors of the elements of the image, the correlation degree between each element in the generated original image and the prompt words input by the user can be measured. For elements that are more in line with the prompt words, their semantic similarity is higher (closer to 1). Since there may be elements in the picture that do not correspond to the prompt words, for example, when the user inputs the text "mountains and waters, distant mountains, clouds in the mountains", there may be an element of "flying birds" in the picture. Denote the elements in the user input text that do not correspond to the prompt words as auxiliary elements. Then, when comparing the similarity between the elements and the prompt words, the semantic similarity of the auxiliary elements will be relatively lower.
[0037] S303. Determine the attention of each element in each original image for each generation number according to the central coefficient and several separated element regions in each original image with different generation numbers.
[0038] First, process each element region through the CLIP model (Contrastive Language-Image Pre-training, a multi-modal machine learning model) in the preset model, so as to convert the separated element regions into a vector space based on the CLIP model and determine the corresponding feature vector of each element. Determine the second cosine similarity between each feature vector and the semantic vectors of each prompt word respectively, and take the element with the largest second cosine similarity as the dominant element, to obtain the dominant element corresponding to each prompt word and the auxiliary elements other than the dominant element in each original image with different generation numbers.
[0039] Second, modify the central coefficient corresponding to the auxiliary elements in each original image with different generation numbers to a first preset value. For example, the first preset value is 0, while the central coefficient corresponding to the dominant element remains unchanged.
[0040] Then, determine the area ratio of each element region in each original image with different generation numbers to the corresponding original image , denotes the th generation number of the th original image, the th element region accounts for the area ratio of the corresponding original image, and determine the sum value of each central coefficient and a second preset value (such as 1) respectively , is the central coefficient of the prompt word corresponding to the element of the th generation number of the th original image, the th element region (the prompt word corresponding to the element means that the cosine similarity between the feature vector of the element and the semantic vector of the prompt word is the highest). Finally, determine the attention of each element in each original image for each generation number according to the first product of the sum value and the area ratio and the linear normalization function , and the specific formula is: where, is the attention of the element (equivalent to each element) of the th generation number of the th original image, the th element region.
[0041] It should be noted that in each iteration, in some cases, the user can continuously input supplementary prompt words to control the direction of the iteratively generated image. With the user's multiple adjustments and feedback, the user's needs can be analyzed through the process of adjusting the image based on the user's feedback. For a certain auxiliary element, although the text input by the user does not involve this element, the user's selection during the iterative selection shows a tendency towards this element. For example, if the user's attention to a certain auxiliary element continuously increases when selecting reference images in multiple iterations, it indicates that the user tends to increase the proportion of this element in the image.
[0042] Among them, several original images generated each time are based on the reference images selected in the previous round, which are the results obtained by adjusting the feedback of the user input. In the embodiments of the present application, the user's needs are analyzed through the changes of elements in the images during the iterative process. Therefore, it is necessary to first determine the correspondence between the elements in each original image generated each time and the elements in the original image generated in the previous time. The dominant elements can be matched according to the same corresponding prompt words, while for the auxiliary elements, the same elements in different images can be respectively corresponding by comparing the similarity of the elements and their positions in the image.
[0043] In one implementation manner, in step S400, in each iteration, according to the attention of each element in each original image of each generation number and the reference image of each generation number, determine the attention range of each auxiliary element in each original image of each generation number, including steps S401 - S403: S401. In each iteration, according to the attention of each element in each original image of each generation number and the reference image of each generation number, determine the trend coefficient of each element in the reference image of each generation number.
[0044] First, in each iteration, respectively determine the correspondence coefficient between each element in any original image of each generation number (for example, the th generation number) and each element in the reference image of the previous generation number of this generation number (for example, the -1th generation number). It should be noted that in some embodiments, the correspondence coefficient here only considers the auxiliary elements: In the formula, is the correspondence coefficient between the element in the th element area of the th original image of the th generation number and the th element in the reference image of the -1th generation number; is the natural exponential function, is the the center of the element area of the th element of the reference image with the generation number of -1, and the distance between the centers of the element areas of the th original image with the generation number of and the characteristic vector of the element of the th element area of the th original image, and the cosine similarity between the characteristic vectors of the
[0045] th element of the reference image with the generation number of
[0046] -1. Add 2 to ensure that the denominator value of the normalization object is non - negative and non - zero, so as to ensure that the normalization result is between 0 and 1.
[0047] Secondly, determine the variance of the attention of each element in all the original images of each generation number , and determine the first difference between the attention of each element in each original image of any generation number and the attention of the element corresponding to each element in the reference image of the previous generation number of this generation number .
[0048] Among them, is the total number of original images of each generation number. For example = 10, is the The variance of the attention of the -th element among all the original images of the -th generation count, is the average attention of the elements in the -th element region among all the original images of the -th generation count; is the attention of the element corresponding to the -th element in the original image of the -th generation count in the reference image of the
[0049] Furthermore, according to the first difference and the number of original images, calculate the arithmetic mean , and respectively according to the ratio of the arithmetic mean to the variance and the proportional normalization function , determine the trend coefficient of each element in the reference image of each generation count. The specific formula is: where, is the trend coefficient of the -th element in the reference image of the -th generation count (i.e., the element corresponding to the -th element in the original image), representing the change trend of the -th element from the reference image of the -th generation count to the reference image of the -th generation count. The larger the absolute value of the trend coefficient, the more obvious the change trend (increase or decrease) of the attention of the element.
[0050] It should be noted that during the iteration process, when the user selects a reference image each time, they tend to consider the visual effect of the dominant elements. As a result, in this process, due to the excessive proportion of the dominant elements, the performance of the auxiliary elements in the picture may be ignored, leading to too low attention of the auxiliary elements to ensure the richness of the picture, and the final visual effect is not good. Therefore, it is necessary to determine a suitable lower threshold for the auxiliary elements in the picture to ensure that the auxiliary elements can play a role in enriching the picture. Compared with the dominant elements, the attention proportion of each auxiliary element in the picture may fluctuate greatly. For example, in some images, the attention is relatively large, while in some images, there is no corresponding element; when the dominant element with a large correlation with a certain auxiliary element changes greatly, resulting in a significant reduction in the attention of the auxiliary element, it is necessary to control the change degree of the auxiliary element (i.e., the dominant element occupies too much of the picture, making the weight of the auxiliary element too low).
[0051] S402. Determine the occupancy coefficients of each auxiliary element for each generation count based on the trend coefficients of each element in the reference image for each generation count.
[0052] First, determine the maximum second cosine similarity between the feature vectors corresponding to each auxiliary element and the semantic vectors of each prompt word in the reference image for each generation count, and determine the dominant element corresponding to the prompt word with the maximum second cosine similarity as the related element of the auxiliary element.
[0053] Second, determine the trend coefficients of each auxiliary element in the reference image for each generation count and the difference from the trend coefficient of the related element of the auxiliary element , and based on the difference and the direct proportional normalization function , determine the occupancy coefficients of each auxiliary element in each original image for each different generation count. The formula is: In the formula, represents the occupancy coefficient of the th auxiliary element for the th generation count, represents the trend coefficient of the related element of the th auxiliary element in the reference image for the th generation count, represents the trend coefficient of the th auxiliary element in the reference image for the th generation count.
[0054] It should be noted that when the absolute value of the occupancy coefficient is larger, it indicates that the attention of the auxiliary element has changed more significantly under the influence of the related element. When the occupancy coefficient is greater than 0, the attention of the auxiliary element is reduced under the influence of its related element. At this time, the lower limit of the attention of the auxiliary element should be increased to ensure the visual effect of the auxiliary element; when the occupancy coefficient is less than 0, the attention of the auxiliary element is increased under the influence of its related element. At this time, the upper limit of the attention of the auxiliary element should be reduced to prevent the auxiliary element from overwhelming the main element in the image.
[0055] S403. Determine the attention range of each auxiliary element for each generation count based on the occupancy coefficient.
[0056] First, based on the above calculation , since covers all elements, that is, it also includes auxiliary elements. Therefore, based on the attention of each auxiliary element in several original images for the th generation count, the maximum attention value of each auxiliary element for each generation count can be determined, denoted as the original maximum attention value (i.e., the maximum value of the original attention of the th generation times of the th auxiliary element), and determine the minimum value of the attention of each auxiliary element for each generation time, denoted as the original minimum attention value (i.e., the original minimum attention value of the th generation times of the th auxiliary element).
[0057] Secondly, determine the second difference between the original maximum attention value and the original minimum attention value , and determine the second product of the second difference and the occupancy coefficient of this auxiliary element . .
[0058] Furthermore, taking the first preset value as 0 as an example, when the occupancy coefficient of the auxiliary element is greater than the first preset value 0, determine the sum of the second product and the original minimum attention value , as the adjusted original minimum attention value, which is equivalent to the adjustment lower limit.
[0059] Then, when the occupancy coefficient of the auxiliary element is less than the first preset value 0, determine the sum of the second product and the original maximum attention value , as the adjusted original maximum attention value, which is equivalent to the adjustment upper limit.
[0060] Finally, in each iteration, according to the adjusted original minimum attention value and the adjusted original maximum attention value, determine the attention range of each auxiliary element for each generation time.
[0061] As Figure 2 shown, it is the data of the attention change of a certain auxiliary element in 4 iterations in each original image, and the corresponding schematic diagram generated based on the data.
[0062] In one implementation manner, in step S400, generating a target image according to the attention range includes steps S404 - S406: S404. According to the attention range, control the attention of each auxiliary element in the iteration process, and continuously determine the attention span between the maximum attention and the minimum attention of the auxiliary element in each iteration.
[0063] In the embodiments of the present application, according to the attention range, the attention of each auxiliary element in the iterative process is controlled. For example, after adjusting the upper and lower limits, it means that when the attention of the auxiliary element generated next time is greater than the upper limit, the upper limit value is taken to regenerate the image. Similarly, when the attention of the auxiliary element is less than the lower limit, the lower limit value is taken to regenerate the image. And after each iteration, the upper and lower limits of adjustment are re-determined based on the foregoing calculation process to determine the attention range. At the same time, the maximum attention of the auxiliary element is continuously determined during each iteration. and the minimum attention of the attention span .
[0064] S405. When the attention span of the auxiliary element is less than the preset stability threshold, the current attention range is used as the target attention range of the auxiliary element until the target attention ranges of all auxiliary elements are determined.
[0065] Optionally, taking the preset stability threshold as 0.02 as an example, there is no specific limitation; when the attention span of the auxiliary element is less than the preset stability threshold of 0.02, the current attention range is used as the target attention range of the auxiliary element, and the target attention range of the auxiliary element is no longer changed. Based on this principle, the target attention ranges of all auxiliary elements can be finally determined.
[0066] S406. Generate a target image according to the target attention ranges of all auxiliary elements.
[0067] Finally, when the AI agent generates the original image, it generates the image based on the target attention ranges of all auxiliary elements, and finally determines at least one target image. The content of the auxiliary elements in the target image can be controlled within a suitable proportion, making the picture quality higher, more beautiful, and conducive to ensuring the richness of the picture.
[0068] Referring to Figure 3 , a structural block diagram of a text-to-image system based on an AI agent according to an embodiment of the present application is shown. The system may include: An acquisition module, configured to acquire an input text composed of multiple prompt words, and generate a plurality of original images according to the input text and the AI agent; An analysis module, configured to perform quality analysis on the original images, determine a reference image according to the quality analysis result and the plurality of original images, and use the reference image as a reference image for the next image generation, and perform several iterations to determine a plurality of original images and reference images with different generation times; A determination module, configured to determine the attention of each element in each original image for each generation time according to a plurality of original images with different generation times, where the attention represents the importance of the element in the original image; A generation module, configured to, at each iteration, determine the attention range of each auxiliary element for each generation time according to the attention of each element in each original image for each generation time and the reference image for each generation time, and generate a target image according to the attention range.
[0069] In the embodiments of the present application, the functions of the modules in the system can refer to the corresponding descriptions in the above method, which will not be elaborated here.
[0070] In one implementation manner, the embodiments of the present application further provide a text-to-image generation device based on an AI agent, including: a processor and a memory, where instructions are stored in the memory, and the instructions are loaded and executed by the processor to implement the above-mentioned text-to-image generation method based on an AI agent.
[0071] It should be noted that the above sequence of the embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired result. In some implementation manners, multi-task processing and parallel processing are also possible or may be advantageous.
[0072] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments.
Claims
1. A method for creating a text map based on an AI agent, characterized in that: The method comprises: Obtain input text consisting of multiple prompt words, and generate several original images based on the input text and the AI agent; Performing quality analysis on the original image, determining a reference image based on the quality analysis result and a plurality of the original images, and using the reference image as a benchmark image for next image generation, performing a plurality of iterations, and determining a plurality of original images and reference images generated at different times; Determine, according to a plurality of original images generated at different times, the attention degree of each element in each original image generated at each time, wherein the attention degree represents the importance of the element in the original image; In each iteration, the attention range of each auxiliary element at each generation number is determined according to the attention of each element in each original image at each generation number and the reference image at each generation number, and the target image is generated according to the attention range.
2. According to the AI agent-based Wensheng diagram method of claim 1, it is characterized by: Determining the attention level of each element in each original image of each generation number according to the plurality of original images of different generation numbers includes: Processing each original image generated at different times by a preset model, determining a plurality of separated element regions in each original image generated at different times, and determining a semantic vector of each of the prompt words; Determining first cosine similarities between semantic vectors of different prompt words, and determining a central coefficient of each prompt word according to the first cosine similarities and the number of the prompt words; The attention degree of each element in each original image of each generation number is determined according to the central coefficient and a plurality of separated element regions in each original image of different generation numbers.
3. The AI agent-based Wensheng diagram method according to claim 2, characterized in that: Determining the attention level of each element in each original image of each generation number according to the center coefficient and a plurality of separated element regions in each original image of different generation numbers includes: According to each of the element regions, a corresponding feature vector of each element is determined, and the second cosine similarity between each of the feature vectors and the semantic vectors of each of the prompt words is determined respectively, and the element with the largest second cosine similarity is taken as the dominant element, and the dominant element corresponding to each prompt word in each of the original images with different generation times and the auxiliary elements other than the dominant element are obtained; Modifying the central coefficient corresponding to the auxiliary element in each original image generated at different times to a first preset value; Determine the area ratio of each element region in the corresponding original image in each original image generated at different times, and determine the sum of each center coefficient and the second preset value respectively, and determine the attention degree of each element in each original image of each generation number according to the first product of the sum and the area ratio and the normalization function.
4. The AI agent-based Wensheng diagram method according to claim 3 is characterized by: In each iteration, according to the attention degree of each element in each original image of each generation number and the reference image of each generation number, determining the attention degree range of each auxiliary element of each generation number includes: At each iteration, according to the attention degree of each element in each original image of each generation number and the reference image of each generation number, the trend coefficient of each element in the reference image of each generation number is determined; Determine the occupation coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number; According to the occupancy coefficient, the attention range of each auxiliary element for each generation number is determined.
5. The AI agent-based Wensheng diagram method according to claim 4 is characterized by: Determining the trend coefficient of each element in the reference image of each generation number according to the attention degree of each element in each original image of each generation number and the reference image of each generation number includes: Determine the corresponding coefficients between each element in each original image of any number of generation times and each element in the reference image of the number of generation times before the number of generation times, and determine the elements in the reference image to which each element in each original image corresponds according to the corresponding coefficients; Determine the variance of the attention degree of each element in all original images of each generation number, and determine the first difference between the attention degree of each element of each original image of any generation number and the attention degree of the element corresponding to each element in the reference image of the previous generation number of the generation number; An arithmetic mean is calculated according to the first difference and the number of the original images, and a trend coefficient of each element in the reference image of each generation number is determined according to a ratio of the arithmetic mean to the variance and a direct proportional normalization function.
6. The AI agent-based method of claim 4, characterized in that: Determining the occupation coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number includes: Determine the maximum second cosine similarity between the feature vector corresponding to each auxiliary element and the semantic vector of each prompt word in the reference image of each generation number, and determine the dominant element corresponding to the prompt word corresponding to the maximum second cosine similarity as the related element of the auxiliary element; The difference between the trend coefficient of each auxiliary element and the trend coefficient of the related element of the auxiliary element in the reference image of each generation number is determined respectively, and the occupancy coefficient of each auxiliary element of each generation number is determined according to the difference and the direct proportional normalization function.
7. The AI agent-based text graph method according to claim 4, characterized in that: Determining the attention range of each auxiliary element of each generation number according to the occupancy coefficient includes: Determine the maximum original attention value and the minimum original attention value of each auxiliary element for each number of generation; Determine a second difference between the original maximum value of the attention degree and the original minimum value of the attention degree, and determine a second product of the second difference and the occupation coefficient of the auxiliary element; When the occupancy coefficient of the auxiliary element is greater than a first preset value, determining a sum of the second product and the original minimum attention value as the adjusted original minimum attention value; When the occupancy coefficient of the auxiliary element is less than a first preset value, determining the sum of the second product and the original maximum attention value as the adjusted original maximum attention value; In each iteration, the attention range of each auxiliary element for each generation number is determined according to the adjusted original attention minimum value and the adjusted original attention maximum value.
8. The AI agent-based Wensheng diagram method according to claim 1, characterized in that: Generating a target image according to the attention range includes: According to the attention range, the attention of each auxiliary element in the iteration process is controlled, and the attention span between the maximum attention and the minimum attention of the auxiliary element is continuously determined in each iteration; wherein the attention range is re-determined after each iteration; When the attention span of an auxiliary element is less than a preset stability threshold, the current attention span is used as the target attention span of the auxiliary element until the target attention spans of all auxiliary elements are determined; Generate a target image based on the target attention range of all auxiliary elements.
9. A text map system based on AI agent, characterized in that: include: An acquisition module is used to acquire an input text consisting of multiple prompt words and generate a number of original images according to the input text and the AI agent; An analysis module is used to perform quality analysis on the original image, determine a reference image based on the quality analysis result and a plurality of the original images, and use the reference image as a benchmark image for next image generation, perform a plurality of iterations, and determine a plurality of original images and reference images generated at different times; A determination module, used to determine the attention degree of each element in each original image of each generation number according to a plurality of original images of different generation numbers, wherein the attention degree represents the importance of the element in the original image; The generation module is used to determine the attention range of each auxiliary element at each generation time according to the attention of each element in each original image at each generation time and the reference image at each generation time in each iteration, and generate the target image according to the attention range.
10. A text-based image device based on AI agent, characterized in that: include: A processor and a memory, wherein the memory stores instructions, and the instructions are loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image generation method and device based on text graph
CN117994376A
Method and system for generating controllable high-quality AI drawing picture description
CN119151793A
Image generation method and device
CN119540381A
Vegetation graph model training method, image generation method, device and medium
CN119849658A
Generation of image corresponding to input text using multi-text guided image cropping
US20240153153A1