A method, system and device for cultural graphics based on AI intelligent agent
By analyzing and adjusting the range of elements' attention in the process of literary and artistic pictures, the problem of insufficient proportion of non-user input content is solved, the high quality and aesthetics of generated images are achieved, and the generation efficiency is improved.
Patent Information
- Application Number
- CN202510625525.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-05-15
AI Technical Summary
In the process of literary production, the proportion of non-user input content is too low, resulting in insufficient other content used to improve the picture in the generated image, affecting the image quality and aesthetics, and the fixed threshold setting may affect the demand of different pictures, reducing the generation efficiency.
By obtaining the input text, several original images are generated, quality analysis is performed, reference images are determined, and the attention of auxiliary elements is adjusted according to the attention range of the element during the iteration process, and the target image is generated to ensure the appropriate proportion of auxiliary elements in the picture.
It improves the quality and beauty of the generated images, ensures the richness of the picture, improves the efficiency of the generation of literary pictures, and adapts to the needs of different pictures.
Smart Images

Figure CN120147483B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a text-based graph method, system and device based on an AI agent. Background Art
[0002] AI agent refers to an artificial intelligence system that can perceive the environment, make decisions and perform tasks. The AI agent-based text-based graph method is a technology that uses artificial intelligence systems to convert natural language text descriptions into images. The main process of AI agent-based text-based graph is: the user inputs the text description, the AI agent model analyzes the text content entered by the user, extracts the image features to be generated from the text entered by the user, and the generator uses the image features to create the image. Multiple rounds of iterative optimization are required to ensure that the generated image is as consistent as possible with the text description.
[0003] Currently, when users input text to generate text-based images, the output often contains non-user-input content (called auxiliary elements). These elements are typically used to refine and enrich the image's details. However, during the iterative process, users may prioritize the primary content corresponding to the text and overlook this non-user-input content. This can lead to a low proportion of non-user-input content after multiple iterations, resulting in insufficient additional content in the generated image to improve the image and failing to achieve the desired effect. This necessitates further user feedback and adjustment, reducing the efficiency of text-based images. Some solutions limit the minimum proportion of non-user-input content by setting a fixed weight threshold. However, since the required proportion of non-user-input content varies across images, setting a fixed threshold can compromise image quality and aesthetics, resulting in poor results. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a method, system and device for cultural graphs based on AI agents. The technical solutions adopted are as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for a cultural graph based on an AI agent, comprising:
[0006] Obtain input text consisting of multiple prompt words, and generate several original images based on the input text and the AI agent;
[0007] Performing a quality analysis on the original image, determining a reference image based on the quality analysis result and a plurality of the original images, and using the reference image as a benchmark image for next image generation, performing a plurality of iterations to determine a plurality of original images and reference images generated at different times;
[0008] Determining, based on a plurality of original images generated at different times, the attention level of each element in each original image generated at each time, wherein the attention level represents the importance of the element in the original image;
[0009] In each iteration, the attention range of each auxiliary element at each generation time is determined according to the attention of each element in each original image at each generation time and the reference image at each generation time, and the target image is generated according to the attention range.
[0010] In one embodiment, determining the attention level of each element in each original image of each generation number based on a plurality of original images of different generation numbers includes:
[0011] Processing each original image generated at different times by a preset model, determining a plurality of separated element regions in each original image generated at different times, and determining a semantic vector of each prompt word;
[0012] Determining a first cosine similarity between semantic vectors of different prompt words, and determining a central coefficient of each prompt word according to the first cosine similarity and the number of the prompt words;
[0013] The attention degree of each element in each original image of each generation number is determined according to the central coefficient and a plurality of separated element regions in each original image of different generation numbers.
[0014] In one embodiment, determining the attention level of each element in each original image of each generation number based on the center coefficient and a plurality of separated element regions in each original image of different generation numbers includes:
[0015] Determining, based on each element region, a corresponding feature vector for each element, and determining a second cosine similarity between each feature vector and the semantic vector of each prompt word, respectively; taking the element with the largest second cosine similarity as a dominant element, and obtaining the dominant element corresponding to each prompt word in each original image generated at different times, as well as auxiliary elements other than the dominant element;
[0016] Modifying the center coefficient corresponding to the auxiliary element in each original image generated at different times to a first preset value;
[0017] Determine the area ratio of each element region in the corresponding original image in each original image generated at different times, and determine the sum of each center coefficient and the second preset value respectively, and determine the attention level of each element in each original image generated at each time based on the first product of the sum and the area ratio and the normalization function.
[0018] In one embodiment, determining the attention range of each auxiliary element at each generation number according to the attention of each element in each original image at each generation number and the reference image at each generation number during each iteration includes:
[0019] At each iteration, according to the attention degree of each element in each original image of each generation number and the reference image of each generation number, the trend coefficient of each element in the reference image of each generation number is determined;
[0020] Determining the occupancy coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number;
[0021] The attention range of each auxiliary element for each generation number is determined according to the occupancy coefficient.
[0022] In one embodiment, determining the trend coefficient of each element in the reference image at each generation number based on the attention level of each element in each original image at each generation number and the reference image at each generation number includes:
[0023] Determining the corresponding coefficients between each element in each original image of any number of generation times and each element in the reference image of the number of generation times before the generation times, and determining the elements in the reference image corresponding to each element in each original image based on the corresponding coefficients;
[0024] Determining the variance of the attention degree of each element in all original images of each generation number, and determining the first difference between the attention degree of each element in each original image of any generation number and the attention degree of the element corresponding to each element in the reference image of the generation number before the generation number;
[0025] An arithmetic mean is calculated according to the first difference and the number of the original images, and a trend coefficient of each element in the reference image of each generation number is determined according to a ratio of the arithmetic mean to the variance and a direct proportional normalization function.
[0026] In one embodiment, determining the occupancy coefficient of each auxiliary element at each generation number based on the trend coefficient of each element in the reference image at each generation number includes:
[0027] Determining the maximum second cosine similarity between the feature vector corresponding to each auxiliary element and the semantic vector of each prompt word in the reference image generated for each number of times, and determining the dominant element corresponding to the prompt word corresponding to the maximum second cosine similarity as the related element of the auxiliary element;
[0028] The difference between the trend coefficient of each auxiliary element and the trend coefficient of the related element of the auxiliary element in the reference image of each generation number is determined respectively, and the occupancy coefficient of each auxiliary element of each generation number is determined based on the difference and the direct proportional normalization function.
[0029] In one embodiment, determining the attention range of each auxiliary element for each generation number according to the occupancy coefficient includes:
[0030] Determine the maximum original attention value and the minimum original attention value of each auxiliary element for each generation number respectively;
[0031] Determine a second difference between the original maximum attention value and the original minimum attention value, and determine a second product of the second difference and the occupancy coefficient of the auxiliary element;
[0032] When the occupancy coefficient of the auxiliary element is greater than a first preset value, determining the sum of the second product and the original minimum attention value as the adjusted original minimum attention value;
[0033] When the occupancy coefficient of the auxiliary element is less than a first preset value, determining the sum of the second product and the original maximum attention value as the adjusted original maximum attention value;
[0034] In each iteration, the attention range of each auxiliary element for each generation number is determined according to the adjusted original attention minimum value and the adjusted original attention maximum value.
[0035] In one embodiment, generating a target image according to the attention range includes:
[0036] According to the attention range, the attention of each auxiliary element in the iteration process is controlled, and the attention span between the maximum attention and the minimum attention of the auxiliary element is continuously determined in each iteration; wherein the attention range is re-determined after each iteration;
[0037] When the attention span of an auxiliary element is less than a preset stability threshold, the current attention span is used as the target attention span of the auxiliary element until the target attention spans of all auxiliary elements are determined;
[0038] Generate a target image based on the target attention range of all auxiliary elements.
[0039] In a second aspect, an embodiment of the present application provides a cultural graph system based on an AI agent, including:
[0040] The acquisition module is used to obtain input text consisting of multiple prompt words and generate several original images based on the input text and the AI agent;
[0041] an analysis module, configured to perform a quality analysis on the original image, determine a reference image based on the quality analysis result and a plurality of the original images, and use the reference image as a benchmark image for the next image generation, perform a plurality of iterations, and determine a plurality of original images and reference images generated at different times;
[0042] a determination module, configured to determine, based on a plurality of original images generated at different times, the attention level of each element in each original image generated at each time, wherein the attention level represents the importance of the element in the original image;
[0043] The generation module is used to determine the attention range of each auxiliary element at each generation time according to the attention of each element in each original image at each generation time and the reference image at each generation time in each iteration, and generate the target image according to the attention range.
[0044] In a third aspect, an embodiment of the present application provides a literary image device based on an AI agent, comprising: a processor and a memory, wherein the memory stores instructions, and the instructions are loaded and executed by the processor to implement the method in any one of the above-mentioned embodiments.
[0045] The present invention has the following beneficial effects:
[0046] By obtaining an input text consisting of multiple prompt words, generating several original images according to the input text and an AI agent, performing quality analysis on the original images, determining a reference image according to the quality analysis results and the several original images, and using the reference image as the benchmark image for the next image generation, performing several iterations, determining several original images and reference images of different generation times, determining the attention of each element in each original image of each generation time according to the several original images of different generation times, and in each iteration, determining the attention range of each auxiliary element in each original image of each generation time according to the attention of each element in each original image of each generation time and the reference image of each generation time, adjusting the attention range of the auxiliary elements considering the attention of the characterizing elements' importance in the original image, and generating the target image according to the attention range, which is beneficial to ensuring the quality and beauty of the target image and improving the generation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0048] Figure 1 A schematic flow chart of the steps of a text-based graph method based on an AI agent provided by one embodiment of the present invention;
[0049] Figure 2 A schematic diagram of changes in the attention level of auxiliary elements during an iteration process provided by one embodiment of the present invention;
[0050] Figure 3 A structural block diagram of a text graph system based on an AI agent provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0051] To further illustrate the technical means and effects employed by the present invention to achieve the intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effects of a text-based graph method, system, and device based on an AI agent proposed by the present invention. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.
[0052] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0053] It should be noted that the term “exemplary” in the embodiments of the present application refers to examples listed for the convenience of explanation, and other embodiments are not limited to the examples listed.
[0054] It should be noted that in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiments of the present invention, when encountering a situation where the denominator is 0, it is necessary to add a parameter adjustment factor greater than 0 to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to actual conditions, and this application does not impose any special restrictions.
[0055] The following describes in detail a specific solution of a text-generated diagram method, system and device based on an AI agent provided by the present invention in conjunction with the accompanying drawings.
[0056] See also Figure 1 , which shows a flow chart of a culture graph method based on an AI agent provided by an embodiment of the present invention. The culture graph method based on an AI agent may include at least steps S100-S400:
[0057] S100: Obtain an input text consisting of multiple prompt words, and generate several original images based on the input text and an AI agent.
[0058] S200, perform quality analysis on the original image, determine a reference image based on the quality analysis result and several original images, and use the reference image as the benchmark image for the next image generation, perform several iterations, and determine several original images and reference images with different generation times.
[0059] S300 : Determine the attention level of each element in each original image of each generation number according to a plurality of original images of different generation numbers.
[0060] The attention degree represents the importance of an element in the original image. The higher the attention degree, the higher the importance.
[0061] S400. According to the attention level of each element in each original image of each generation number and the reference image of each generation number, at each iteration, the attention level range of each auxiliary element in each original image of each generation number is determined, and the target image is generated according to the attention level range.
[0062] The technical solution of the embodiment of the present application obtains an input text consisting of multiple prompt words, generates several original images according to the input text and an AI intelligent agent, performs quality analysis on the original images, determines a reference image according to the quality analysis results and the several original images, and uses the reference image as the benchmark image for the next image generation, performs several iterations, determines several original images and reference images of different generation times, determines the attention of each element in each original image of each generation time according to the several original images of different generation times, determines the attention range of each auxiliary element in each original image of each generation time according to the attention of each element in each original image of each generation time and the reference image of each generation time in each iteration, adjusts the attention range of the auxiliary elements considering the attention of the characterizing elements' importance in the original image, and generates the target image according to the attention range, which is beneficial to ensuring the quality and beauty of the target image and improving the generation effect.
[0063] In one embodiment, in step S100, the user can perform input operations based on his or her own needs to generate input text. Usually, the input text consists of multiple prompt words. The input text can be obtained by the AI agent, and then the AI agent analyzes the input text to generate several original images. Optionally, the analysis process of the AI agent may include: the text understanding module of the AI agent understands the input text, for example, using the BERT model based on the Transformer architecture to parse the input text, converting the natural language input text into a text encoding vector that the machine can understand, and performing word segmentation and annotation and text cleaning to extract each semantic vector used to generate the image. Then, the image generation module of the AI agent uses the diffusion model Stable Diffusion or the generator of the generative adversarial network GAN to generate several images through these semantic vectors, which are recorded as original images.
[0064] In one embodiment, in step S200, the AI agent evaluation and optimization module uses a CLIP model or a GAN discriminator to perform a quality analysis on the original image and evaluate the quality score of the generated image (i.e., the quality analysis result). For example, the quality score is a structural consistency score (whether the original image conforms to the content of the input text). Then, according to the structural consistency score, several original images with the highest quality scores are displayed to the user. For example, 10 original images are displayed as an example, and the adjustment can be made according to the specific situation. Finally, the user selects one of the 10 original images as a reference image, thereby determining the reference image. It should be noted that this reference image is used as the baseline image for the next image generation, and several iterations are performed to determine several original images and reference images with different generation times. After each reference image selection, the user can enter new supplementary text, add, delete, or adjust the image content, and then perform several iterations to determine several original images and reference images with different generation times, without specific limitation.
[0065] It should be noted that during the iteration process of the text-based image, the user's adjustments may prioritize the primary content corresponding to the text, while ignoring non-user input content. This can result in insufficient additional content (auxiliary elements) in the generated original image to improve the image, failing to achieve the desired effect and reducing the efficiency of the text-based image. Therefore, the present invention analyzes the changes in non-user input content in the image generated during the iterative optimization process, limits the degree of change in non-user input content, and obtains a relatively stable range of attention variation to ensure the visual effect and richness of non-user input content.
[0066] Among them, for the text initially input by the user, different prompt words may involve the same element. For example, if the user input text is "mountains and waters, distant mountains, and clouds and fog between mountains", these three prompt words are all associated with the element of mountain, which means that the user pays more attention to this element. When generating an image, the mountain as the main body should occupy a larger proportion of the screen. Therefore, it is necessary to analyze the semantic correlation between the prompt words.
[0067] In one embodiment, step S300 includes steps S301-S303:
[0068] S301 , processing each original image generated at different times by using a preset model, determining a number of separated element regions in each original image generated at different times, and determining a semantic vector of each prompt word.
[0069] Optionally, the preset model includes a pre-trained YOLO neural network model. Each original image generated at different times is input into the pre-trained YOLO neural network model. The pre-trained YOLO neural network model separates the elements in each original image, thereby determining a number of separated element regions in each original image generated at different times. Furthermore, all prompt words input by the user are converted into a vector space to determine a semantic vector for each prompt word.
[0070] S302: Determine the first cosine similarity between the semantic vectors of different prompt words, and determine the center coefficient of each prompt word according to the first cosine similarity and the number of prompt words.
[0071] Optionally, determine the first cosine similarity between the semantic vectors of different prompt words , For the The prompt word and The cosine similarity between the semantic vectors corresponding to the prompt words; and based on the first cosine similarity and the number of prompt words , determine the central coefficient of each prompt word, the specific formula is:
[0072]
[0073] in, For the The central coefficient of each prompt word reflects the correlation between each prompt word and other prompt words. When a prompt word is highly correlated with other prompt words, the cosine similarity of the vectors corresponding to the prompt word and other prompt words is larger, and the central coefficient is higher. Represents a normalization function, specifically a linear normalization function.
[0074] It should be noted that the central coefficient is obtained based on the semantic relevance between the prompt words. However, semantic relevance does not fully represent the proportion of the corresponding elements in the image. It is necessary to further combine the central coefficient with the proportion of all elements in the image in each generated original image to obtain the attention of each element in each image. By comparing the cosine similarity between the feature vectors of the elements in the image, the correlation degree between each element in the generated original image and the prompt words input by the user can be measured. For elements that more conform to the prompt words, their semantic similarity is higher (closer to 1). Since there may be elements in the image that do not correspond to the prompt words, for example, when the user inputs the text "mountains and waters, distant mountains, mountain mists", there may be an element of "flying birds" in the image. These elements that do not have corresponding prompt words in the user input text are recorded as auxiliary elements. Then, when comparing the similarity between the elements and the prompt words, the semantic similarity of the auxiliary elements will be relatively lower.
[0075] S303. Determine the attention of each element in each original image for each generation count according to the central coefficient and several separated element regions in each original image of different generation counts.
[0076] First, according to each element region, process it through the CLIP model (Contrastive Language-Image Pre-training, a multimodal machine learning model) in the preset model, so as to convert the separated element regions into a vector space based on the CLIP model and determine the corresponding feature vector of each element. Respectively determine the second cosine similarity between each feature vector and the semantic vectors of each prompt word, and take the element with the largest second cosine similarity as the dominant element, and obtain the dominant elements corresponding to each prompt word and the auxiliary elements other than the dominant elements in each original image of different generation counts.
[0077] Second, modify the central coefficient corresponding to the auxiliary element in each original image of different generation counts to the first preset value. For example, the first preset value is 0, while the central coefficient corresponding to the dominant element remains unchanged.
[0078] Then, determine the area proportion of each element region in the corresponding original image in each original image of different generation counts , denotes the th generation count, the th element region in the and respectively determine the sum value of each central coefficient and the second preset value (for example, 1) is the th generation count, the In the original image, The central coefficient of the prompt word corresponding to the element in the element area (the prompt word corresponding to the element refers to the element whose feature vector has the highest cosine similarity with the semantic vector of the prompt word). Finally, according to the first product of the sum value and the area ratio and the linear normalization function , determine the attention of each element in each original image for each generation number, the specific formula is:
[0079] in, For the The number of times generated In the original image, The attention level of the elements in the element area (equivalent to each element).
[0080] It should be noted that in some cases, during each iteration, the user can continuously input supplementary prompt words to control the direction of the iterative image generation. With multiple adjustments and feedback from the user, the user's needs can be analyzed through the process of adjusting the screen through user feedback. For a certain auxiliary element, although the text entered by the user does not refer to the element, the user's choice during the iterative selection process shows a preference for the element. For example, if the user's attention to a certain auxiliary element continues to increase when selecting reference images in multiple iterations, it means that the user tends to increase the screen proportion of the element.
[0081] Among them, the several original images generated each time are based on the reference images selected in the previous round, which is equivalent to the result obtained by adjusting the feedback input by the user. The embodiment of the present application analyzes user needs through the changes in the elements in the image during the iteration process. Therefore, it is first necessary to determine the correspondence between each element in the original image generated each time and the elements of the original image generated last time. The dominant elements can be matched according to the corresponding prompt words, and for the auxiliary elements, the same elements in different images can be matched separately by comparing the similarity of the elements and the position in the image.
[0082] In one embodiment, in each iteration of step S400, based on the attention level of each element in each original image at each generation number and the reference image at each generation number, the attention level range of each auxiliary element in each original image at each generation number is determined, including steps S401-S403:
[0083] S401 , in each iteration, determining the trend coefficient of each element in the reference image of each generation number according to the attention degree of each element in each original image of each generation number and the reference image of each generation number.
[0084] First, at each iteration, determine the number of times each original image is generated (for example, The number of times each element in the number of generations) is related to the number of generations before that number (for example, -1 generation times) of the reference image. It should be noted that in some embodiments, the corresponding coefficients here only consider auxiliary elements:
[0085]
[0086] Where, For the The number of times generated The original image The elements of the element region are -1 times the number of times the reference image is generated The corresponding coefficients between the elements; is the natural exponential function, For the The number of times generated The original image The center of the element area and the -1 times the number of times the reference image is generated The distance between the centers of the element regions of elements, For the The number of times generated The original image The characteristic vector of the element in the element region is the same as the -1 times the number of times the reference image is generated The cosine similarity between the eigenvectors of the elements is calculated by adding 2 to ensure that the denominator of the normalized object is non-negative and non-zero, so that the normalized result is between 0 and 1.
[0087] Furthermore, based on the corresponding coefficients, the corresponding elements in the reference image for each element in each original image are determined. Specifically, a preset corresponding threshold of 0.7 is used to determine the maximum corresponding coefficient corresponding to each element in each original image. If this maximum corresponding coefficient is greater than the preset corresponding threshold of 0.7, it is used as the target corresponding coefficient. The elements corresponding to the target corresponding coefficients are then determined in the reference image, thereby determining the corresponding elements in the reference image for each element in each original image. It should be noted that if there is no corresponding element in the reference image, it is recorded as a newly added element.
[0088] It should be noted that during the multiple iterations of generating the user-selected reference image, the attention level of each element in the image may exhibit a certain trend of change, such as an increase or decrease in attention level. When the attention level of an element changes consistently across the generated images, for example, if it shows an increase in attention level across all generated images, this indicates that the attention level of the element has exhibited a clear trend. Conversely, if the attention level of the element changes irregularly across the generated images and does not show a consistent change direction, this indicates that the change in attention level of the element has not exhibited a clear trend.
[0089] Secondly, determine the variance of the attention of each element in all original images for each generation number , and respectively determine the first difference between the attention degree of each element of each original image of any generation number and the attention degree of the element corresponding to each element in the reference image of the previous generation number of the generation number .
[0090]
[0091] in, The total number of original images for each generation, e.g. =10, For the Among all the original images generated times The variance of the attention of each element, For the Among all the original images generated times, The average attention level of elements in the element area; For the -1 times generated reference image, and The number of times the original image is generated The attention level of the element corresponding to the element.
[0092] Furthermore, according to the first difference and the number of original images, calculate the arithmetic mean , respectively according to the arithmetic mean and variance The ratio and the proportional normalization function , determine the trend coefficient of each element in the reference image for each generation number. The specific formula is:
[0093]
[0094] in, For the The number of times the reference image is generated element (i.e., the first element in the original image) The trend coefficient of the element corresponding to the element represents the -1 times the reference image is generated to The number of reference images generated The larger the absolute value of the trend coefficient is, the more the attention change of the element shows a certain trend (increase or decrease).
[0095] It should be noted that in the iterative process, each time the user selects a reference image, they tend to consider the picture effect of the dominant element. In this process, the performance of the auxiliary elements in the picture may be ignored due to the excessive proportion of the dominant element, resulting in too low attention to the auxiliary elements and the inability to ensure the richness of the picture, resulting in a poor final picture effect. Therefore, it is necessary to determine a suitable lower threshold for the auxiliary elements in the picture to ensure that the auxiliary elements can play a role in enriching the picture. Compared with the dominant elements, the proportion of attention of each auxiliary element in the picture may fluctuate greatly. For example, in some images, the attention is greater, while in some images, there is no corresponding element; when the dominant element with a high degree of correlation with a certain auxiliary element undergoes a large change, resulting in a significant decrease in the attention of the auxiliary element, it is necessary to control the degree of change of the auxiliary element (that is, the dominant element occupies too much of the picture, making the weight of the auxiliary element too low).
[0096] S402 : Determine the occupancy coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number.
[0097] First, the maximum second cosine similarity between the feature vector corresponding to each auxiliary element and the semantic vector of each prompt word in the reference image of each generation number is determined, and the dominant element corresponding to the prompt word corresponding to the maximum second cosine similarity is determined as the related element of the auxiliary element.
[0098] Secondly, determine the trend coefficient of each auxiliary element in the reference image for each generation number. Trend coefficient of the element related to the auxiliary element The difference , and normalize the function based on the difference and the direct proportion , determine the occupancy coefficient of each auxiliary element in each original image of each different generation number, the formula is:
[0099]
[0100] Where, Indicates the The number of times generated The occupancy factor of the auxiliary elements, Indicates the Among the reference images generated times, the The trend coefficient of the related elements of the auxiliary elements, Indicates the Among the reference images generated times, the The trend coefficient of the auxiliary element.
[0101] It should be noted that the larger the absolute value of the occupancy coefficient is, the greater the change in the degree of attention of the auxiliary element under the influence of related elements. When the occupancy coefficient is greater than 0, the degree of attention of the auxiliary element is reduced by the influence of its related elements. At this time, the lower limit of the attention of the auxiliary element should be increased to ensure the picture effect of the auxiliary element; when the occupancy coefficient is less than 0, the auxiliary element is increasingly affected by its related elements. At this time, the upper limit of the attention of the auxiliary element should be lowered to avoid the auxiliary element from taking the lead in the picture.
[0102] S403: Determine the attention range of each auxiliary element for each generation number according to the occupancy coefficient.
[0103] First, based on the above calculation ,because covers all elements, i.e. also auxiliary elements, so based on The attention degree of each auxiliary element in several original images of the generation times can be used to determine the maximum attention degree of each auxiliary element in each generation time, which is recorded as the original attention degree maximum value. (i.e. The number of times generated The original maximum value of the auxiliary element's attention) and the minimum value of the attention of each auxiliary element for each generation number are determined, which is recorded as the original minimum value of attention (i.e. The number of times generated the minimum original attention of auxiliary elements).
[0104] Secondly, determine the second difference between the original maximum value and the original minimum value of attention , determine the second difference and the occupancy coefficient of the auxiliary element The second product of .
[0105] Furthermore, taking the first preset value as 0 as an example, when the occupancy coefficient of the auxiliary element is If the value is greater than the first preset value 0, the sum of the second product and the original minimum value of attention is determined. , as the minimum value of the original attention after adjustment, is equivalent to the adjustment lower limit.
[0106] Then, when the occupancy factor of the auxiliary element is less than The first preset value is 0, and the sum of the second product and the original maximum value of attention is determined. , as the maximum value of the original attention after adjustment, is equivalent to the adjustment upper limit.
[0107] Finally, in each iteration, the attention range of each auxiliary element for each generation number is determined according to the adjusted original attention minimum value and the adjusted original attention maximum value.
[0108] like Figure 2 As shown, the data of the attention change of a certain auxiliary element in each original image for 4 iterations, and the corresponding schematic diagram generated based on the data.
[0109] In one embodiment, generating a target image according to the attention range in step S400 includes steps S404-S406:
[0110] S404: Control the attention of each auxiliary element in the iteration process according to the attention range, and continuously determine the attention span between the maximum attention and the minimum attention of the auxiliary element in each iteration.
[0111] In the embodiment of the present application, the attention of each auxiliary element in the iteration process is controlled according to the attention range. For example, after adjusting the upper and lower limits, it means that when the attention of the auxiliary element is greater than the upper limit next time, the upper limit value is used to regenerate the image. The same is true for the lower limit. When the attention of the auxiliary element is less than the lower limit, the lower limit value is used to regenerate the image. After each iteration, the upper and lower limits are re-adjusted based on the above calculation process to determine the attention range. At the same time, the maximum attention of the auxiliary element is continuously determined in each iteration. With minimal attention The span of attention .
[0112] S405: When the attention span of the auxiliary element is smaller than the preset stability threshold, the current attention range is used as the target attention range of the auxiliary element until the target attention ranges of all auxiliary elements are determined.
[0113] Optionally, the preset stability threshold is 0.02, which is not limited to a specific value. When it is less than the preset stability threshold of 0.02, the current attention range is used as the target attention range of the auxiliary element, and the target attention range of the auxiliary element is no longer changed. Based on this principle, the target attention range of all auxiliary elements can be finally determined.
[0114] S406: Generate a target image according to the target attention ranges of all auxiliary elements.
[0115] Finally, when generating the original image, the AI agent generates the image based on the target attention range of all auxiliary elements, and finally determines at least one target image. The content of the auxiliary elements in the target image can be controlled at an appropriate proportion, making the picture quality higher and more beautiful, which is conducive to ensuring the richness of the picture.
[0116] Reference Figure 3 , shows a structural block diagram of a cultural graph system based on an AI agent according to an embodiment of the present application, which may include:
[0117] The acquisition module is used to obtain input text consisting of multiple prompt words and generate several original images based on the input text and the AI agent;
[0118] An analysis module is used to perform quality analysis on the original image, determine a reference image based on the quality analysis result and several original images, and use the reference image as the benchmark image for the next image generation, perform several iterations, and determine several original images and reference images generated at different times;
[0119] a determination module, configured to determine, based on a plurality of original images generated at different times, the attention degree of each element in each original image generated at each time, wherein the attention degree represents the importance of the element in the original image;
[0120] The generation module is used to determine the attention range of each auxiliary element at each generation time according to the attention of each element in each original image at each generation time and the reference image at each generation time in each iteration, and generate the target image according to the attention range.
[0121] In the embodiment of the present application, the functions of each module in the system can be found in the corresponding description in the above method and will not be repeated here.
[0122] In one embodiment, the embodiment of the present application also provides an AI agent-based text map device, including: a processor and a memory, the memory stores instructions, and the instructions are loaded and executed by the processor to implement the above-mentioned AI agent-based text map method.
[0123] It should be noted that the order in which the embodiments of the present invention are described above is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0124] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
Claims
1. A method for creating a cultural graph based on an AI agent, characterized in that: The method comprises: Obtain input text consisting of multiple prompt words, and generate several original images based on the input text and the AI agent; Performing a quality analysis on the original image, determining a reference image based on the quality analysis result and a plurality of the original images, and using the reference image as a benchmark image for next image generation, performing a plurality of iterations to determine a plurality of original images and reference images generated at different times; Determining, based on a plurality of original images generated at different times, the attention level of each element in each original image generated at each time, wherein the attention level represents the importance of the element in the original image; In each iteration, the attention range of each auxiliary element at each generation time is determined according to the attention of each element in each original image at each generation time and the reference image at each generation time, and the target image is generated according to the attention range.
2. The AI agent-based cultural graph method according to claim 1, characterized in that: The determining of the attention level of each element in each original image of each generation number based on the plurality of original images of different generation numbers includes: Processing each original image generated at different times by a preset model, determining a plurality of separated element regions in each original image generated at different times, and determining a semantic vector of each prompt word; Determining a first cosine similarity between semantic vectors of different prompt words, and determining a central coefficient of each prompt word according to the first cosine similarity and the number of the prompt words; The attention degree of each element in each original image of each generation number is determined according to the central coefficient and a plurality of separated element regions in each original image of different generation numbers.
3. The AI agent-based cultural graph method according to claim 2, characterized in that: Determining the attention level of each element in each original image of each generation number based on the center coefficient and a plurality of separated element regions in each original image of different generation numbers includes: Determining, based on each element region, a corresponding feature vector for each element, and determining a second cosine similarity between each feature vector and the semantic vector of each prompt word, respectively; taking the element with the largest second cosine similarity as a dominant element, and obtaining the dominant element corresponding to each prompt word in each original image generated at different times, as well as auxiliary elements other than the dominant element; Modifying the center coefficient corresponding to the auxiliary element in each original image generated at different times to a first preset value; Determine the area ratio of each element region in the corresponding original image in each original image generated at different times, and determine the sum of each center coefficient and the second preset value respectively, and determine the attention level of each element in each original image generated at each time based on the first product of the sum and the area ratio and the normalization function.
4. The AI agent-based cultural graph method according to claim 3, characterized in that: In each iteration, determining the attention range of each auxiliary element at each generation number according to the attention of each element in each original image at each generation number and the reference image at each generation number includes: At each iteration, according to the attention degree of each element in each original image of each generation number and the reference image of each generation number, the trend coefficient of each element in the reference image of each generation number is determined; Determining the occupancy coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number; The attention range of each auxiliary element for each generation number is determined according to the occupancy coefficient.
5. The AI agent-based cultural graph method according to claim 4, characterized in that: The step of determining the trend coefficient of each element in the reference image at each generation number according to the attention degree of each element in each original image at each generation number and the reference image at each generation number includes: Determining the corresponding coefficients between each element in each original image of any number of generation times and each element in the reference image of the number of generation times before the generation times, and determining the elements in the reference image corresponding to each element in each original image based on the corresponding coefficients; Determining the variance of the attention degree of each element in all original images of each generation number, and determining the first difference between the attention degree of each element in each original image of any generation number and the attention degree of the element corresponding to each element in the reference image of the generation number before the generation number; An arithmetic mean is calculated according to the first difference and the number of the original images, and a trend coefficient of each element in the reference image of each generation number is determined according to a ratio of the arithmetic mean to the variance and a direct proportional normalization function.
6. The AI agent-based cultural graph method according to claim 4, characterized in that: Determining the occupancy coefficient of each auxiliary element at each generation number according to the trend coefficient of each element in the reference image at each generation number includes: Determining the maximum second cosine similarity between the feature vector corresponding to each auxiliary element and the semantic vector of each prompt word in the reference image generated for each number of times, and determining the dominant element corresponding to the prompt word corresponding to the maximum second cosine similarity as the related element of the auxiliary element; The difference between the trend coefficient of each auxiliary element and the trend coefficient of the related element of the auxiliary element in the reference image of each generation number is determined respectively, and the occupancy coefficient of each auxiliary element of each generation number is determined based on the difference and the direct proportional normalization function.
7. The AI agent-based cultural graph method according to claim 4, characterized in that: Determining the attention range of each auxiliary element for each generation number according to the occupancy coefficient includes: Determine the maximum original attention value and the minimum original attention value of each auxiliary element for each generation number respectively; Determine a second difference between the original maximum attention value and the original minimum attention value, and determine a second product of the second difference and the occupancy coefficient of the auxiliary element; When the occupancy coefficient of the auxiliary element is greater than a first preset value, determining the sum of the second product and the original minimum attention value as the adjusted original minimum attention value; When the occupancy coefficient of the auxiliary element is less than a first preset value, determining the sum of the second product and the original maximum attention value as the adjusted original maximum attention value; In each iteration, the attention range of each auxiliary element for each generation number is determined according to the adjusted original attention minimum value and the adjusted original attention maximum value.
8. The AI agent-based cultural graph method according to claim 1, characterized in that: Generating a target image according to the attention range includes: According to the attention range, the attention of each auxiliary element in the iteration process is controlled, and the attention span between the maximum attention and the minimum attention of the auxiliary element is continuously determined in each iteration; wherein the attention range is re-determined after each iteration; When the attention span of an auxiliary element is less than a preset stability threshold, the current attention span is used as the target attention span of the auxiliary element until the target attention spans of all auxiliary elements are determined; Generate a target image based on the target attention range of all auxiliary elements.
9. A cultural graph system based on AI agent, characterized by: include: The acquisition module is used to obtain input text consisting of multiple prompt words and generate several original images based on the input text and the AI agent; an analysis module, configured to perform a quality analysis on the original image, determine a reference image based on the quality analysis result and a plurality of the original images, and use the reference image as a benchmark image for the next image generation, perform a plurality of iterations, and determine a plurality of original images and reference images generated at different times; a determination module, configured to determine, based on a plurality of original images generated at different times, the attention level of each element in each original image generated at each time, wherein the attention level represents the importance of the element in the original image; The generation module is used to determine the attention range of each auxiliary element at each generation time according to the attention of each element in each original image at each generation time and the reference image at each generation time in each iteration, and generate the target image according to the attention range.
10. A cultural image device based on AI agent, characterized in that: include: A processor and a memory, wherein the memory stores instructions, and the instructions are loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Image generation method and device based on text graph
CN117994376A
Method and system for generating controllable high-quality AI drawing picture description
CN119151793A