Picture generation and optimization method based on AI
By extracting key phrases through text recognition and semantic analysis, generating relevant conditions using fully connected layers and LSTM networks, and combining the CLIP model for image enhancement, the problem of text conversion errors in AI image generation is solved, achieving efficient and accurate image generation.
Patent Information
- Application Number
- CN202510978520.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-07-16
AI Technical Summary
Existing AI image generation technology has difficulty in accurately converting text information into images, resulting in significant errors between the generated images and user needs, low efficiency and waste of computing resources.
Key phrases are extracted through text recognition and semantic analysis, relevant conditions are generated using the fully connected layer network and LSTM network, and image enhancement processing is performed in combination with the CLIP model to ensure that the image style and content are consistent with the requirements.
The accuracy of image generation and system efficiency are improved, errors are reduced, and computing resources are saved.
Smart Images

Figure CN120599074A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent image generation, and specifically relates to an AI-based image generation and optimization method. Background Art
[0002] Traditionally, holiday image production has relied on professional designers dedicating considerable time and effort to manual creation. From conceiving holiday themes and selecting appropriate elements to sketching out the composition and coordinating color palettes, every step requires the designer's extensive experience and refined aesthetic sense. However, this purely manual production model is not only inefficient and unable to cope with the large volume of image requests within a short period of time, but more importantly, due to differences in creative styles and understanding among designers, it is difficult to precisely meet the individual needs of each user.
[0003] With the rapid development of artificial intelligence (AI), AI has demonstrated remarkable potential in the field of image generation and optimization. AI tools that use text descriptions can rapidly generate visually appealing images by learning from massive amounts of image data. Although numerous methods for generating images from text have been proposed and applied, in practice, these methods often struggle to fully and accurately integrate textual information into the image generation process. The rich semantics contained in text, including specific object forms, scene atmosphere, and emotional tendencies, are difficult for AI systems to fully capture and translate into visual elements in an image. This results in significant discrepancies between the generated image and the desired image. This not only significantly reduces the quality of the generated image but also results in a significant waste of computer resources, leading to inefficient operation of the entire system. Summary of the Invention
[0004] The purpose of this invention is to solve the problem that there is an obvious error between the image generated according to the text and the image actually desired by the user, which not only greatly reduces the quality of the image generation, but also causes a serious waste of computer resources, making the operation efficiency of the entire system low, and proposes an AI-based image generation and optimization method.
[0005] In the embodiment of the present invention, an AI-based image generation and optimization method is proposed, the method comprising:
[0006] Obtaining a requirements document, performing text recognition and semantic analysis on the requirements document to obtain key phrases, and substituting the key phrases into a fully connected layer network to obtain a phrase vector representation; the key phrases include a subject phrase, an object phrase, and a relationship phrase; and the phrase vector representation includes a subject representation, an object representation, and a relationship representation;
[0007] Substituting the phrase vector representation into a long short-term memory network to obtain a target relationship prompt;
[0008] Substituting the target relationship prompt into a preset model to obtain an initial image;
[0009] Substituting the target relationship prompt into the horizontal LSTM network and the vertical LSTM network respectively to obtain horizontal correlation conditions and vertical correlation conditions, and substituting the requirement document into the CLIP model to obtain global correlation conditions;
[0010] The initial image is enhanced according to the horizontal correlation condition, the vertical correlation condition and the global correlation condition to obtain a target required image.
[0011] Optionally, performing text recognition and semantic analysis on the requirement document to obtain key phrases includes:
[0012] Preprocessing the demand document to obtain an initial manuscript, and performing word segmentation processing on the initial manuscript to obtain a short word set;
[0013] Removing stop words from the short word set to obtain a valid short word set;
[0014] The effective short word set is processed through syntactic analysis and semantic role labeling to obtain subject phrases, object phrases and relation phrases; the subject phrases, the object phrases and the relation phrases are used as key phrases.
[0015] Optionally, substituting the target relationship prompt into a preset model to obtain an initial image includes:
[0016] Initialize Gaussian noise, substitute the Gaussian noise into the fully connected layer for dimension adjustment to obtain target Gaussian noise;
[0017] Performing feature splicing on the target Gaussian noise and the target relationship hint to obtain an initial fusion feature;
[0018] Substituting the initial fusion features into an image generation module to obtain a first image;
[0019] Performing global maximum pooling and global average pooling on the first image to obtain a global maximum feature and a global average feature;
[0020] Concatenating the global maximum feature and the global average feature along the channel dimension to obtain a global fusion feature;
[0021] A 1×1 convolution operation is performed on the global fusion feature to obtain an initial image.
[0022] Optionally, substituting the initial fusion features into an image generation module to obtain a first image includes:
[0023] Performing a convolution operation on the initial fusion feature after upsampling to obtain a first feature, and substituting the initial fusion feature and the first feature into an improved text fusion layer to obtain a second feature;
[0024] A convolution operation is performed on the second feature to obtain a third feature, and the third feature and the initial fusion feature are substituted into an improved text fusion layer to obtain a first image.
[0025] Optionally, the improved text fusion layer includes two improved affine transformation layers and two activation layers alternately stacked, followed by a convolution layer; the working principle of the improved text fusion layer specifically includes:
[0026] Obtaining a first input feature and the initial fused feature, and substituting the initial fused feature into an improved affine transformation layer to obtain a first affine output;
[0027] Adjusting the first input feature according to the first affine output to obtain a first affine feature; the first input feature is the first feature or the third feature;
[0028] Substituting the first affine feature into the activation layer to obtain a first activation feature;
[0029] Substituting the first activation feature into an improved affine transformation layer to obtain a second affine output, and adjusting the first activation feature according to the second affine output to obtain a second affine feature;
[0030] Substituting the second affine feature into the activation layer, a 1×1 convolution operation is performed to obtain the first image.
[0031] Optionally, the improved affine transformation layer includes:
[0032] Obtaining input features, and learning scaling parameters and translation parameters from the input features through two multi-layer perceptrons;
[0033] The input feature is adjusted according to the scaling parameter and the translation parameter to obtain an output feature.
[0034] Optionally, performing enhancement processing on the initial image according to the horizontal correlation condition, the vertical correlation condition, and the global correlation condition to obtain the target required image includes:
[0035] Performing word-level radiation change on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map;
[0036] The global correlation condition performs a global radiation change on the first radiation change map to obtain a second radiation change map;
[0037] fusing the initial image and the second radiometric change map to obtain an initial fused change map;
[0038] The matching degree between the initial fused change map and the horizontal correlation condition, the vertical correlation condition and the global correlation condition is verified by a discriminator. If the matching degree is greater than a preset threshold, the initial fused change map is recorded as the target requirement image.
[0039] Optionally, performing word-level radiation change on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map includes:
[0040] Determine an object according to the horizontal correlation condition and the vertical correlation condition, and generate a scaling factor and an offset for each object by a first multilayer perceptron and a second multilayer perceptron;
[0041] Determine the dynamic weight corresponding to each object through the first multi-layer perceptron and Softmax;
[0042] The objects in the initial image are adjusted according to the scaling factor, offset and dynamic weight of each object to obtain a first radiometric change map.
[0043] Optionally, performing a global radiation change on the first radiation change graph using the global correlation condition to obtain a second radiation change graph includes:
[0044] Extracting sentence-level vectors in the global correlation conditions to obtain global affine parameters;
[0045] The first radiation change map is adjusted according to the global affine parameter to obtain a second radiation change map.
[0046] Optionally, after verifying the matching degree between the initial fused change map and the horizontal correlation condition, the vertical correlation condition, and the global correlation condition through a discriminator, the method further includes:
[0047] If the matching degree is less than or equal to the preset value, the scaling factor and offset of each object are adjusted through back propagation to re-optimize the first radiation change map.
[0048] Beneficial effects of the present invention:
[0049] The present invention proposes an AI-based image generation and optimization method, which obtains a requirement document, performs text recognition and semantic analysis on the requirement document to obtain key phrases, and substitutes the key phrases into a fully connected layer network to obtain a phrase vector representation; the key phrases include subject phrases, object phrases, and relationship phrases; the phrase vector representation includes subject representation, object representation, and relationship representation; the phrase vector representation is substituted into a long short-term memory network to obtain a target relationship prompt; the target relationship prompt is substituted into a preset model to obtain an initial image; the target relationship prompt is substituted into a horizontal LSTM network and a vertical LSTM network respectively to obtain horizontal correlation conditions and vertical correlation conditions, and the requirement document is substituted into a CLIP model to obtain global correlation conditions; the initial image is enhanced according to the horizontal correlation conditions, vertical correlation conditions, and global correlation conditions to obtain a target requirement image. By extracting key phrases through text recognition and semantic analysis, the core information of the requirement document can be accurately captured. Then, by converting the key phrases into vector representations and converting the implicit logical relationships in the text into computable structured information through an LSTM network, combined with global correlation conditions, the overall style and content of the image are ensured to be consistent with the requirement, thereby reducing the error between the image generated by the text and the image actually desired by the user, and improving the operating efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The present invention will be further described below with reference to the accompanying drawings.
[0051] Figure 1 A flowchart of an AI-based image generation and optimization method provided in an embodiment of the present invention;
[0052] Figure 2 A flowchart of a method for obtaining an initial image through a target relationship provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0054] Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of the present invention.
[0055] The embodiment of the present invention provides an AI-based image generation and optimization method. Figure 1 , Figure 1 A flowchart of an AI-based image generation and optimization method provided in an embodiment of the present invention. The method includes the following steps:
[0056] S101: Obtain the requirement document, perform text recognition and semantic analysis on the requirement document to obtain key phrases, and substitute the key phrases into the fully connected layer network to obtain the phrase vector representation;
[0057] S102, substitute the phrase vector representation into the long short-term memory network to obtain the target relationship prompt;
[0058] S103, substituting the target relationship prompt into the preset model to obtain an initial image;
[0059] S104: Substitute the target relationship hint into the horizontal LSTM network and the vertical LSTM network to obtain horizontal correlation conditions and vertical correlation conditions, and substitute the requirement document into the CLIP model to obtain global correlation conditions;
[0060] S105, performing enhancement processing on the initial image according to the horizontal correlation condition, the vertical correlation condition, and the global correlation condition to obtain the target required image;
[0061] Among them, key phrases include topic phrases, object phrases and relationship phrases; phrase vector representation includes topic representation, object representation and relationship representation.
[0062] An AI-based image generation and optimization method provided by an embodiment of the present invention can accurately capture the core information of a requirement document by extracting key phrases through text recognition and semantic analysis. By converting the key phrases into vector representations and converting the implicit logical relationships in the text into computable structured information through an LSTM network, the method combines global relevant conditions to ensure that the overall style and content of the image are consistent with the requirements, thereby reducing the error between the image generated by the text and the image actually desired by the user, and improving the operating efficiency of the system.
[0063] In one implementation, key phrases are extracted through text recognition and semantic analysis to accurately capture the core information of the requirements document. By adding a fully connected layer and LSTM, the implicit logical relationships in the text (spatial position, action association) can be converted into computable structured information, ensuring that the text semantics are fully encoded into a computable vector representation, avoiding subjective biases or omissions that may occur when manually understanding the requirements.
[0064] In one implementation, the initial image is generated through target relationship mapping to avoid the ambiguity of manually designed features in traditional methods; the horizontal and vertical structural constraints of the image are captured by horizontal and vertical LSTMs respectively, and the global semantic consistency extracted by CLIP is combined to perform local-global joint optimization on the initial image, thereby improving the image generation accuracy.
[0065] In one implementation, the global relevant conditions ensure that the overall style, content of the image and the global semantics of the requirements are consistent; the entire process from parsing the requirements document to image generation is automatically processed by the dependency model, reducing the links of manual annotation, design sketches and repeated modifications, realizing the efficient and accurate conversion from text requirements to high-quality images, and improving the system efficiency.
[0066] In one embodiment, the key phrases obtained by performing optical character recognition and semantic parsing on the requirements document include:
[0067] Preprocess the requirements document to obtain an initial manuscript, and perform word segmentation on the initial manuscript to obtain a short word set;
[0068] Remove the stop words in the short word set to obtain an effective short word set;
[0069] Process the effective short word set through syntactic analysis and semantic role labeling to obtain topic phrases, object phrases and relationship phrases; the topic phrases, object phrases and relationship phrases are used as key phrases.
[0070] In one implementation, preprocessing and word segmentation break down the text into basic language units. After removing stop words, the effective short word set filters out meaningless words such as "de", "le", "zai", etc., reducing the interference of redundant information and making subsequent analysis more focused on valuable content.
[0071] In one implementation, by extracting topic phrases, object phrases and relationship phrases through syntactic analysis and semantic role labeling, the logical structure of "who (topic)", "to whom / what (object)", "did what / has what relationship (relationship)" in the text can be clearly sorted out.
[0072] In one implementation, preprocessing is to remove redundant symbols (such as punctuation marks, special characters), correct spelling mistakes (perform context correction through language models such as BERT), unify case, format and other existing technical means.
[0073] In one implementation, the extraction of key phrases compresses long texts into structured core elements. Subsequently, tasks such as information retrieval, text classification, sentiment analysis or knowledge graph construction can be efficiently carried out based on these key phrases, avoiding repeated processing of the complete text and saving time and computing power costs.
[0074] In one implementation, word segmentation, syntactic analysis and semantic role labeling are implemented through natural language processing (NLP) technology.
[0075] In one implementation, compared with simple keyword extraction, the division of topics, objects, and relationship phrases is more in line with the semantics and syntactic rules of the language, can more comprehensively reflect the connotation of the text, reduce ambiguity caused by isolated keywords, and improve the accuracy of understanding the text theme, core content and logical relationships; the extracted key phrases can be directly used to construct text summaries, generate tags, match similar texts, etc.
[0076] In one embodiment, see Figure 2 , Figure 2 A flowchart of a method for obtaining an initial image through a target relationship is provided, further comprising:
[0077] S1031, initialize Gaussian noise, substitute the Gaussian noise into the fully connected layer to adjust the dimension to obtain the target Gaussian noise;
[0078] S1032, performing feature concatenation on the target Gaussian noise and the target relationship hint to obtain an initial fusion feature;
[0079] S1033, substituting the initial fusion features into an image generation module to obtain a first image;
[0080] S1034, performing global maximum pooling and global average pooling on the first image to obtain a global maximum feature and a global average feature;
[0081] S1035, concatenating the global maximum feature and the global average feature along the channel dimension to obtain a global fusion feature;
[0082] S1036: Perform a 1×1 convolution operation on the global fusion feature to obtain an initial image.
[0083] In one implementation, Gaussian noise is initialized as the input starting point, and its randomness is used to provide rich initial images for image generation, avoiding the uniformity of the generated results. The dimensionality adjustment of the fully connected layer can map the noise to the feature dimension that is adapted to the subsequent tasks, ensuring that the noise information can effectively participate in the image generation process, laying the foundation for diversified image output.
[0084] In one implementation, the dimension is adjusted to the scale of the final required image; by feature splicing the target Gaussian noise and the target relationship hint, abstract semantic relationships can be embedded in the initial stage of image generation. This fusion allows the generative model to clearly understand what relationship content to generate, avoids the disconnection between the generated results and the target semantics, and improves the match between the image and the expected relationship.
[0085] In one implementation, global maximum pooling and global average pooling are performed on the first image to capture the most significant features (areas with the highest brightness and strongest contrast) and overall trend features (average color and overall structure) in the image, respectively. The global fusion features of the two after splicing along the channel dimension can fully reflect the global information of the image, preventing the model from focusing too much on local details and ignoring the overall structure.
[0086] In one embodiment, substituting the initial fusion features into the image generation module to obtain the first image includes:
[0087] The initial fusion feature is upsampled and then convolved to obtain the first feature, and the initial fusion feature and the first feature are substituted into the improved text fusion layer to obtain the second feature;
[0088] A convolution operation is performed on the second feature to obtain the third feature, and the third feature and the initial fusion feature are substituted into the improved text fusion layer to obtain the first image.
[0089] In one implementation, the improved text fusion layer intervenes twice at different feature stages (initial fusion feature and first feature, third feature and initial fusion feature), which allows text information (such as semantics in target relationship cues) to penetrate into visual features in stages and multiple levels. Unlike the shallow information association that may be caused by a single fusion, this progressive fusion ensures that text semantics continue to take effect during the evolution of features from low dimension to high dimension and from coarse to fine, avoiding visual features from deviating from text guidance and improving the consistency of image and text semantics; the convolution operation is performed through a 1×1 convolution kernel.
[0090] In one implementation, the initial fused feature is upsampled and convolved to obtain the first feature, which is essentially to increase the dimension and enhance the details of the feature; and the initial fused feature is fused with the first feature to obtain the second feature, which is equivalent to using the underlying semantics retained in the initial feature to constrain the evolution direction of the detail feature, avoiding semantic distortion caused by excessive expansion of the details after upsampling.
[0091] In one implementation, the second feature is subsequently convolved to obtain the third feature, which is then fused with the initial fused feature for a second time. The global structure of the high-order features is further calibrated through the core semantics of the initial feature, ensuring that during feature iteration, while the richness of details is improved, the overall semantic coherence is not destroyed. The final generated image contains both rich details and fits the core semantics.
[0092] In one implementation, the two applications of the text fusion layer are improved, and differentiated fusion is performed on features at different stages (low-order initial features, mid-order first features, and high-order third features). Early fusion focuses on injecting text semantics into basic features to guide the general direction of feature generation; late fusion focuses on using text semantics to constrain the detailed expression of high-order features, avoiding the dilution of text information after it only plays a role in the initial stage, and enhancing the overall guiding power of text on image generation.
[0093] In one embodiment, the improved text fusion layer includes two improved affine transformation layers and two activation layers stacked alternately, followed by a convolution layer. The working principle of the improved text fusion layer specifically includes:
[0094] Obtaining a first input feature and an initial fusion feature, and substituting the initial fusion feature into an improved affine transformation layer to obtain a first affine output;
[0095] Adjusting the first input feature according to the first affine output to obtain a first affine feature; the first input feature is the first feature or the third feature;
[0096] Substituting the first affine feature into the activation layer to obtain the first activation feature;
[0097] Substituting the first activation feature into the improved affine transformation layer to obtain a second affine output, and adjusting the first activation feature according to the second affine output to obtain a second affine feature;
[0098] Substitute the second affine feature into the activation layer and perform a 1×1 convolution operation to obtain the first image.
[0099] In one implementation method, the improved affine transformation layer generates affine transformation parameters (first and second affine outputs) based on the initial fusion features (including text semantics and noise features), and dynamically adjusts the first input feature (first feature or third feature) and the first activation feature respectively. The essence of this adjustment is to calibrate the visual features (input features, activation features) under the guidance of text semantics. For example, according to the "color" and "shape" information described in the text, the scale, offset or distribution of the features is dynamically adjusted to make the visual features more consistent with the text semantics, reduce the deviation between vision and text, and improve the consistency between the two.
[0100] In one implementation, the alternating use of affine transformation layers and activation layers (affine adjustment, activation, re-affine adjustment, and reactivation) forms a progressive feature optimization chain of linear transformation and nonlinear mapping. The affine transformation adjusts the basic distribution of features through linear scaling and offset, and the activation layer (such as ReLU, Swish, etc.) introduces nonlinearity to enhance the feature's ability to express complex patterns; the quadratic affine adjustment further performs fine-tuning calibration based on the nonlinear features, allowing the features to break through linear limitations while maintaining direction under semantic constraints, thereby preventing the features from deviating from the core semantics during nonlinear transformation.
[0101] In one implementation, the initial fused features are reused twice as the semantic benchmark for affine transformation, ensuring that the text semantics continue to play a role in different stages of feature processing (input feature adjustment, post-activation feature adjustment), avoiding the dilution of semantic information after a single adjustment; at the same time, the first input feature (the first feature or the third feature) undergoes two affine calibrations and activation optimizations, and the visual details it contains (such as texture and edges) are gradually refined and fused with semantics, realizing the efficient flow and integration of visual information and semantic information.
[0102] In one implementation, the parameters of the improved affine transformation layer are dynamically generated by the initial fused features rather than fixed parameters. This means that the transformation rules will be adaptively adjusted as the input text semantics or initial noise changes, making the feature processing scene-adaptive and improving the model's response accuracy to different text conditions.
[0103] In one implementation, after two affine adjustments, nonlinear enhancement of the activation layer and the final 1×1 convolution, the output first image is based on the features of "semantic calibration + nonlinear optimization + channel compression", which not only retains key visual details but also deeply integrates text semantics. In addition, the 1×1 convolution reduces redundant channels and improves the compactness of features. It can be directly converted into a clearer and more expected image, reducing the burden of subsequent processing and improving generation efficiency and quality.
[0104] In one embodiment, improving the affine transformation layer includes:
[0105] Get input features and learn scaling and translation parameters from them through two multi-layer perceptrons;
[0106] The input features are adjusted according to the scaling parameters and translation parameters to obtain the output features.
[0107] In one implementation, the scaling and translation parameters are not fixed values, but are dynamically learned from the input features through a multi-layer perceptron (MLP). This means that the adjustment rules will adaptively change according to the specific content of the input features. For example, when the input features contain bright areas, the MLP may learn a larger scaling parameter to enhance the contrast of the area; when the input features have background interference, the feature distribution may be offset by the translation parameter to highlight the subject. The dynamic nature allows the feature adjustment to better fit the characteristics of the input data, avoiding the limitations brought by fixed rules.
[0108] In one embodiment, performing enhancement processing on the initial image according to the horizontal correlation condition, the vertical correlation condition, and the global correlation condition to obtain the target required image includes:
[0109] Perform word-level radiation changes on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map;
[0110] Performing a global radiation change on the first radiation change map using a global correlation condition to obtain a second radiation change map;
[0111] Fusing the initial image and the second radiometric change map to obtain an initial fused change map;
[0112] The discriminator is used to verify the matching degree between the initial fusion change map and the horizontal correlation conditions, vertical correlation conditions and global correlation conditions. If the matching degree is greater than the preset threshold, the initial fusion change map is recorded as the target demand image.
[0113] In one implementation, horizontally related conditions (such as the horizontal position of objects and left-right relationships) and vertically related conditions (such as the upper and lower positions of objects and hierarchical relationships) perform word-level affine transformations on the initial image from the local spatial dimension to ensure that the image conforms to specific constraints in the detailed spatial layout; global related conditions (such as the overall scene layout and viewing angle range) adjust the first radial change map from the global dimension to ensure that the overall structure of the image meets macro requirements; the discriminator is the MA-GP discriminator.
[0114] In one implementation, word-level affine transformation focuses on the precise adjustment of local elements (position and angle correction of individual objects), while global affine transformation coordinates the coordination of the overall layout (avoiding overlap and disproportion between objects after local adjustments). The two are combined to form an optimization chain of "local calibration and global coordination", which not only retains the precise response of local details to conditions, but also maintains the rationality and coherence of the overall image through global adjustments, so that the final image maintains high quality both in details and global levels.
[0115] In one implementation, the discriminator verifies the match between the initial fused change map and three conditions. Only when an image satisfies the horizontal, vertical, and global conditions is it considered the target image. This mechanism filters out unqualified images caused by transformation errors, condition conflicts, and other factors, preventing the output of results that do not meet the requirements and improving the reliability and effectiveness of the final image.
[0116] In one implementation, horizontal, vertical, and global conditions can be independently set or adjusted. This layered processing of word-level and global transformations allows the model to flexibly respond to requirements at varying granularities. For example, modifying only the horizontal condition allows for individual adjustments to the left and right positions of objects without affecting the overall layout. Modifying the global condition allows for a shift in overall scene perspective while preserving the relative relationships of local objects. This controllability enables image generation to adapt to diverse and personalized needs.
[0117] In one implementation, the initial image and the second radiometric change map are fused instead of directly using the transformed image. This not only retains the valid information in the initial image that meets the conditions, but also corrects the parts that do not meet the conditions through the second radiometric change map, achieving the dual effects of "valid information retention and deviation correction" and improving feature utilization efficiency and transformation accuracy.
[0118] In one embodiment, performing word-level radiation change on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map includes:
[0119] Determine the object according to the horizontal correlation condition and the vertical correlation condition, and generate a scaling factor and an offset of each object through a first multilayer perceptron and a second multilayer perceptron;
[0120] Determine the dynamic weight corresponding to each object through the first multi-layer perceptron and Softmax;
[0121] The objects in the initial image are adjusted according to the scaling factor, offset and dynamic weight of each object to obtain a first radiometric change map.
[0122] In one implementation, a first multilayer perceptron is used to generate a scaling factor; a second multilayer perceptron is used to generate an offset; the first multilayer perceptron is combined with Softmax to generate a dynamic weight for each object, and adjustment priorities are assigned to different objects based on the importance of spatial conditions or the relevance of objects; when there is a potential conflict in spatial conditions, objects with higher weights can be prioritized to ensure that their conditions are met, while objects with lower weights can be appropriately compromised to maintain overall rationality, thereby improving the flexibility and feasibility of adjustments.
[0123] In one implementation, adjusting the objects in the initial image according to the scaling factor, offset, and dynamic weight of each object to obtain the first radiation change map specifically includes the following steps: Get each adjusted object, and get the first radiation change map based on all the adjusted objects. For the pixel point set M in the area where each object is located, adjust each pixel point using the above formula, where AFF is the adjusted object, M is the pixel point set of the object, and m is the pixel point. is the dynamic weight of pixel m, is the scaling factor of pixel m, is the original value of pixel m, is the offset of pixel m.
[0124] In one embodiment, performing a global radiation change on the first radiation change graph under the global correlation condition to obtain a second radiation change graph includes:
[0125] Extract the sentence-level vectors in the global correlation conditions to obtain the global affine parameters;
[0126] The first radiation change map is adjusted according to the global affine parameter to obtain a second radiation change map.
[0127] In one implementation, the global relevant conditions usually include a description of the overall scene. Extracting the sentence-level vector (i.e., the overall encoding of the global semantics) can transform the abstract global requirements into quantifiable global affine parameters.
[0128] In one implementation, global semantic constraints are directly applied to the overall spatial adjustment of the image, ensuring that the first radial change map (the image after local object adjustment) meets global requirements in terms of overall layout, perspective, scale, and other macro levels, avoiding the problem of local objects meeting the conditions but the overall scene deviating.
[0129] In one implementation, the global affine parameters include a first coefficient A and a second coefficient B; all objects in the first radiation change map are modified using the formula ADF=A*T(AFF)+B to obtain the second radiation change map ADF.
[0130] In one implementation method, the first radial change map focuses on the adjustment of the horizontal / vertical relationship of the local object, which may result in local optimization but overall incoordination. The adjustment of the global affine parameters is coordinated from a macro perspective. The combination of the two forms a closed loop of "local precise adjustment and global coordinated optimization", allowing the image to maintain consistency and rationality in both local details and global structure.
[0131] In one implementation, sentence-level vectors contain rich information about global semantics, and the global affine parameters generated based on this information can adapt to the characteristics of different scenarios. Compared with fixed global adjustment rules, this semantically driven parameter generation method allows the model to automatically optimize and adjust strategies according to different global needs, improving its adaptability to diverse scenarios.
[0132] In one embodiment, after verifying the matching degree between the initial fused change map and the horizontal correlation condition, the vertical correlation condition, and the global correlation condition through the discriminator, the following steps are further included:
[0133] If the matching degree is less than or equal to the preset value, the scaling factor and offset of each object are adjusted through back propagation to re-optimize the first radiation change map.
[0134] In one implementation, if the degree of match is less than or equal to a preset value, the back-propagation mechanism can trace the "error signal" of the match (i.e., the degree of deviation between the image and the condition) along the parameter generation path (multi-layer perceptron), accurately locating which objects' scaling factors or offsets need to be corrected. By adjusting the parameters in a targeted manner and regenerating the first radial change map, a closed loop of "generation, verification, correction, and regeneration" is formed, gradually narrowing the deviation between the image and the condition until the degree of match is met, significantly improving the accuracy of the final result.
[0135] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.
Claims
1. An AI-based image generation and optimization method, characterized in that: The method comprises: Obtaining a requirements document, performing text recognition and semantic analysis on the requirements document to obtain key phrases, and substituting the key phrases into a fully connected layer network to obtain a phrase vector representation; the key phrases include a subject phrase, an object phrase, and a relationship phrase; and the phrase vector representation includes a subject representation, an object representation, and a relationship representation; Substituting the phrase vector representation into a long short-term memory network to obtain a target relationship prompt; Substituting the target relationship prompt into a preset model to obtain an initial image; Substituting the target relationship prompt into the horizontal LSTM network and the vertical LSTM network respectively to obtain horizontal correlation conditions and vertical correlation conditions, and substituting the requirement document into the CLIP model to obtain global correlation conditions; The initial image is enhanced according to the horizontal correlation condition, the vertical correlation condition and the global correlation condition to obtain a target required image.
2. The AI-based image generation and optimization method according to claim 1, characterized in that: Key phrases obtained by performing text recognition and semantic analysis on the requirement document include: Preprocessing the demand document to obtain an initial manuscript, and performing word segmentation processing on the initial manuscript to obtain a short word set; Removing stop words from the short word set to obtain a valid short word set; The effective short word set is processed through syntactic analysis and semantic role labeling to obtain subject phrases, object phrases and relation phrases; the subject phrases, the object phrases and the relation phrases are used as key phrases.
3. The AI-based image generation and optimization method according to claim 1, characterized in that: Substituting the target relationship prompt into a preset model to obtain an initial image includes: Initialize Gaussian noise, substitute the Gaussian noise into the fully connected layer for dimension adjustment to obtain target Gaussian noise; Performing feature splicing on the target Gaussian noise and the target relationship hint to obtain an initial fusion feature; Substituting the initial fusion features into an image generation module to obtain a first image; Performing global maximum pooling and global average pooling on the first image to obtain a global maximum feature and a global average feature; Concatenating the global maximum feature and the global average feature along the channel dimension to obtain a global fusion feature; A 1×1 convolution operation is performed on the global fusion feature to obtain an initial image.
4. The AI-based image generation and optimization method according to claim 3, characterized in that: Substituting the initial fusion features into an image generation module to obtain a first image includes: Performing a convolution operation on the initial fusion feature after upsampling to obtain a first feature, and substituting the initial fusion feature and the first feature into an improved text fusion layer to obtain a second feature; A convolution operation is performed on the second feature to obtain a third feature, and the third feature and the initial fusion feature are substituted into an improved text fusion layer to obtain a first image.
5. The AI-based image generation and optimization method according to claim 4, characterized in that: The improved text fusion layer includes two improved affine transformation layers and two activation layers alternately stacked and followed by a convolution layer; The working principle of the improved text fusion layer specifically includes: Obtaining a first input feature and the initial fused feature, and substituting the initial fused feature into an improved affine transformation layer to obtain a first affine output; Adjusting the first input feature according to the first affine output to obtain a first affine feature; the first input feature is the first feature or the third feature; Substituting the first affine feature into the activation layer to obtain a first activation feature; Substituting the first activation feature into an improved affine transformation layer to obtain a second affine output, and adjusting the first activation feature according to the second affine output to obtain a second affine feature; Substituting the second affine feature into the activation layer, a 1×1 convolution operation is performed to obtain the first image.
6. The AI-based image generation and optimization method according to claim 5, characterized in that: Improvements to the affine transformation layer include: Obtaining input features, and learning scaling parameters and translation parameters from the input features through two multi-layer perceptrons; The input feature is adjusted according to the scaling parameter and the translation parameter to obtain an output feature.
7. The AI-based image generation and optimization method according to claim 1, characterized in that: Performing enhancement processing on the initial image according to the horizontal correlation condition, the vertical correlation condition, and the global correlation condition to obtain a target required image includes: Performing word-level radiation change on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map; The global correlation condition performs a global radiation change on the first radiation change map to obtain a second radiation change map; fusing the initial image and the second radiometric change map to obtain an initial fused change map; The matching degree between the initial fused change map and the horizontal correlation condition, the vertical correlation condition and the global correlation condition is verified by a discriminator. If the matching degree is greater than a preset threshold, the initial fused change map is recorded as the target requirement image.
8. The AI-based image generation and optimization method according to claim 7, characterized in that: Performing word-level radiation change on the initial image according to the horizontal correlation condition and the vertical correlation condition to obtain a first radiation change map includes: Determine an object according to the horizontal correlation condition and the vertical correlation condition, and generate a scaling factor and an offset for each object by a first multilayer perceptron and a second multilayer perceptron; Determine the dynamic weight corresponding to each object through the first multi-layer perceptron and Softmax; The objects in the initial image are adjusted according to the scaling factor, offset and dynamic weight of each object to obtain a first radiometric change map.
9. The AI-based image generation and optimization method according to claim 7, characterized in that: The global correlation condition performs a global radiation change on the first radiation change graph to obtain a second radiation change graph, including: Extracting sentence-level vectors in the global correlation conditions to obtain global affine parameters; The first radiation change map is adjusted according to the global affine parameter to obtain a second radiation change map.
10. The AI-based image generation and optimization method according to claim 8, characterized in that: After verifying the matching degree between the initial fusion change map and the horizontal correlation condition, the vertical correlation condition and the global correlation condition through the discriminator, the method further includes: If the matching degree is less than or equal to the preset value, the scaling factor and offset of each object are adjusted through back propagation to re-optimize the first radiation change map.
Citation Information
Patent Citations
Phrase-level text image generation method and system based on self-attention mechanism
CN115587160A
Text image generation method and system based on local detail editing
CN116245967A
Digital human image AI generation method and device
CN117011435A
CLIP text-to-image synthesis method based on cyclic affine transformation
CN118967877A
Text image generation method and system based on cross-modal information guide fusion
CN120147479A