Semantic style separation image generation method, device, equipment and medium
Patent Information
- Application Number
- CN202611058626.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-09-01
AI Technical Summary
[0005]本发明的主要目的在于提供一种语义风格分离的图像生成方法、装置、设备及存储介质,旨在解决现有文本到图像生成技术将内容语义与视觉风格耦合在单一文本嵌入表示中,导致图像生成过程中难以在保持语义内容稳定的同时独立控制风格表达的技术问题
[0010]有益效果:本发明涉及语义解析技术领域,公开了一种语义风格分离的图像生成方法、装置、设备及介质,包括:获取文本提示信息并生成标准化文本嵌入序列;识别语义文本单元和风格文本单元,生成语义候选位置集合和风格候选位置集合;基于语义候选位置集合生成语义嵌入向量,并基于语义候选位置集合和风格候选位置集合生成门控矩阵,进行门控风格编码,生成风格嵌入向量;将语义嵌入向量和风格嵌入向量映射至兼容空间,生成兼容语义表示和兼容风格表示;依据风格强度参数缩放兼容风格表示,并与兼容语义表示组合生成组合文本表示;将组合文本表示输入目标图像生成模型,生成候选图像,并依据图像分析结果输出目标图像。本发明可应用于金融科技及医疗健康等业务场景中,通过分离文本提示信息中的语义文本单元和风格文本单元,使语义嵌入向量和风格嵌入向量分别形成;通过门控矩阵减少风格编码中的语义干扰,并通过风格强度参数调节兼容风格表示的贡献,因此能够降低内容语义与视觉风格的耦合程度,提高语义保持和风格控制的稳定性。
Smart Images

Figure CN122676012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic parsing technology, and in particular to an image generation method, apparatus, device, and medium for semantic style separation. Background Technology
[0002] A key shortcoming of text-to-image generation technology lies in the fact that existing text encoding processes typically compress content semantics and visual style into a single embedded representation, causing the subject object, attribute relationships, scene content, and stylistic information such as color, material, composition, and artistic expression to become entangled. Consequently, when generating images, the model struggles to independently adjust stylistic expression while maintaining the core semantics, easily leading to issues such as subject semantic shift, loss of key attributes, altered scene relationships, or insufficient stylistic expression. This problem is particularly prominent in business image generation scenarios that require simultaneous emphasis on content accuracy and style control.
[0003] In the fintech field, text-to-image generation is commonly used for financial product promotional images, investor education illustrations, insurance business explanations, credit risk control tips, asset allocation explanations, and compliance risk warnings. Existing models, when processing textual prompts related to wealth management products, investment risks, insurance claims, credit assessments, and asset allocation, tend to mix the financial business meaning with visual styles such as business style, hand-drawn style, and data visualization style. This leads to shifts in the financial entity, risk level, business relationship, or compliance information as the style changes. For financial image generation tasks, the most prominent problem with existing technologies is their inability to consistently maintain the financial semantic content while simultaneously controlling the image style.
[0004] In the healthcare field, text-to-image generation is commonly used for health education illustrations, disease prevention posters, medication tips, rehabilitation guidance diagrams, medical process diagrams, and patient education materials. Existing models, when processing text prompts related to chronic disease management, post-operative rehabilitation, pediatric medication tips, and mental health education, are prone to altering the relationships between medical objects, health behaviors, applicable populations, or scenarios due to semantic and stylistic coupling. For healthcare image generation tasks, the most prominent problem with existing technologies is the difficulty in independently controlling image styles such as approachable, gentle, cartoonish, and realistic while maintaining the accuracy of healthcare semantics. Summary of the Invention
[0005] The main objective of this invention is to provide a semantic style separation image generation method, apparatus, device, and storage medium, aiming to solve the technical problem that existing text-to-image generation technologies couple content semantics and visual style into a single text embedding representation, making it difficult to independently control style expression while maintaining semantic content stability during image generation.
[0006] To achieve the above objectives, the present invention provides an image generation method for semantic style separation, comprising: Obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; Identify semantic text units and style text units in the text unit sequence, and generate a set of semantic candidate positions and a set of style candidate positions according to the positions of the semantic text units and style text units in the standardized text embedding sequence; Semantic attention is aggregated based on the standardized text embedding sequence and the set of semantic candidate positions to generate a semantic embedding vector. A gating matrix is generated based on the standardized text embedding sequence, the set of semantic candidate positions, and the set of style candidate positions. The gating matrix is then used to perform gating style encoding to generate a style embedding vector. Determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; Obtain style intensity parameters, scale the compatible style representation based on the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The combined text representation is input into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet the preset output conditions, the target image is output.
[0007] Furthermore, to achieve the above objectives, the present invention provides an image generation apparatus for semantic style separation, comprising: The text embedding preprocessing module is used to obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; A semantic style tagging and recognition module is used to identify semantic text units and style text units in the text unit sequence, and generate a semantic candidate position set and a style candidate position set according to the positions of the semantic text units and style text units in the standardized text embedding sequence; The dual-branch gated coding module is used to aggregate semantic attention based on the standardized text embedding sequence and the semantic candidate position set to generate a semantic embedding vector, and to generate a gated matrix based on the standardized text embedding sequence, the semantic candidate position set and the style candidate position set, and to perform gated style coding using the gated matrix to generate a style embedding vector. The model compatibility mapping module is used to determine the target image generation model and compatibility space containing the cross attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; The style intensity combination module is used to obtain style intensity parameters, scale the compatible style representation according to the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The image generation, analysis, and output module is used to input the combined text representation into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, analyzes the semantic and style consistency between the candidate images and the text prompt information, generates image analysis results, and outputs the target image when the image analysis results meet the preset output conditions.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a semantic style separation image generation program stored in the memory and executable on the processor, wherein when the semantic style separation image generation program is executed by the processor, it implements the steps of the semantic style separation image generation method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a semantic style separation image generation program, wherein the semantic style separation image generation program, when executed by a processor, implements the steps of the semantic style separation image generation method as described above.
[0010] Beneficial Effects: This invention relates to the field of semantic parsing technology and discloses a method, apparatus, device, and medium for image generation with semantic style separation. The method includes: acquiring text prompt information and generating a standardized text embedding sequence; identifying semantic text units and style text units, generating a set of semantic candidate positions and a set of style candidate positions; generating a semantic embedding vector based on the set of semantic candidate positions, and generating a gating matrix based on the set of semantic candidate positions and the set of style candidate positions, performing gating style encoding, and generating a style embedding vector; mapping the semantic embedding vector and the style embedding vector to a compatible space, generating a compatible semantic representation and a compatible style representation; scaling the compatible style representation according to a style intensity parameter, and combining it with the compatible semantic representation to generate a combined text representation; inputting the combined text representation into a target image generation model to generate candidate images, and outputting the target image based on image analysis results. This invention can be applied to business scenarios such as fintech and healthcare. By separating the semantic text units and style text units in the text prompt information, the semantic embedding vector and style embedding vector are formed separately; the gating matrix reduces semantic interference in style encoding, and the style intensity parameter adjusts the contribution of the compatible style representation, thus reducing the coupling between content semantics and visual style, and improving the stability of semantic preservation and style control. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an image generation method with semantic style separation according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the semantic style separation image generation method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the semantic style separation image generation device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The semantic style separation image generation method provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain text prompt information from the client and generate a standardized text embedding sequence; identify semantic text units and style text units, and generate a set of semantic candidate positions and a set of style candidate positions; generate semantic embedding vectors based on the set of semantic candidate positions, and generate a gating matrix based on the set of semantic candidate positions and the set of style candidate positions, perform gating style encoding, and generate style embedding vectors; map the semantic embedding vectors and style embedding vectors to a compatible space to generate compatible semantic representations and compatible style representations; scale the compatible style representation according to the style intensity parameter, and combine it with the compatible semantic representation to generate a combined text representation; input the combined text representation into the target image generation model to generate candidate images, and output the target image based on the image analysis results. This invention can be applied to business scenarios such as fintech and healthcare. By separating the semantic text units and style text units in the text prompt information, the semantic embedding vectors and style embedding vectors are formed separately; the gating matrix reduces semantic interference in style encoding, and the style intensity parameter adjusts the contribution of the compatible style representation, thus reducing the coupling degree between content semantics and visual style, and improving the stability of semantic preservation and style control. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the semantic style separation image generation method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the semantic style separation image generation method proposed in this invention includes the following steps: S10, obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; In this embodiment, the text prompt information can be obtained from user input, template filling, business field concatenation, or extraction from interaction records. The content typically includes the main object, attribute modifiers, scene description, action relationships, and visual style description. When receiving the text prompt information, the boundary information between punctuation, conjunctions, parentheses, and modifier phrases can be preserved, while duplicate whitespace, undisplayed characters, and formatting characters unrelated to the image content can be removed, ensuring stable input for subsequent word segmentation.
[0016] Tokenization is used to divide continuous text into a sequence of text units. Text units can be words, subwords, phrase fragments, compound terms, or style descriptions. For compound content that is easily split incorrectly, the entire fragment can be preserved through word list matching, phrase boundary detection, and maximum length matching. For example, risk warning diagrams, asset allocation instructions, rehabilitation guidance illustrations, and watercolor-style content can be preserved as continuous text units. The text unit sequence is arranged according to the order of appearance in the text prompts. Each text unit can have a position number, character range, and fragment type, ensuring that the internal structural relationships of the text are preserved during subsequent mapping.
[0017] Mapping is used to convert a sequence of text units into a normalized text embedding sequence. Each text unit can be converted into a vector representation through an embedding table, a subword encoder, or a text encoding layer. Normalization may include standardizing sequence length, standardizing vector dimensions, configuring position vectors, configuring fragment vectors, and configuring padding masking markers. Longer text can be truncated while preserving the main content and style description, while shorter text can be padded to the required length. The normalized text embedding sequence is obtained by fusing text unit vectors, position vectors, and fragment vectors, providing a consistent input format for subsequent recognition of semantic and style text units.
[0018] In one implementation, the text prompt information, after character cleaning, enters a word segmenter. The word segmenter uses a basic vocabulary and a phrase vocabulary for matching. The basic vocabulary handles general descriptive words, while the phrase vocabulary handles compound phrases and business terms. The segmentation results are written into a text unit sequence in the order of appearance, with each text unit recording the character range and sequence position. During embedding mapping, the text unit sequence is converted into a word vector sequence, and then position vectors and fragment vectors are superimposed. Low-priority modifier fragments exceeding a preset length are truncated, and insufficient length is padded with padding vectors to generate a standardized text embedding sequence.
[0019] In another implementation, the text prompt information first undergoes phrase boundary detection, and consecutive modifying phrases, parallel phrases, and style phrases are marked as indivisible segments. The word segmenter performs sub-word segmentation within indivisible segments and word-level segmentation outside the segments. During embedding mapping, sub-words within the same indivisible segment share the segment vector, while the sub-word position vector is preserved. This implementation is suitable for inputs containing long, complex descriptions, such as promotional images for low-risk financial products in fintech or illustrations reminding children of medication in healthcare.
[0020] Alternatively, a dynamic loading approach using domain-specific vocabularies can be employed. For fintech scenarios, vocabularies for terms such as risk warnings, insurance claims, credit assessment, and asset allocation are loaded; for healthcare scenarios, vocabularies for terms such as chronic disease management, rehabilitation training, medication reminders, and nutritional advice are loaded. The word segmenter prioritizes matching the dynamically loaded vocabularies before matching the basic vocabulary. During embedding mapping, text units that hit the domain vocabulary are configured with domain fragment vectors, ensuring that the standardized text embedding sequence retains the distinction between business context and visual description.
[0021] In fintech business scenarios, the input text is an illustration for investor education on low-risk financial products, with a simple business style and a blue background. After word segmentation, it can be divided into text units such as "low-risk financial products," "investor education illustration," "simple business style," and "blue background." After embedding and mapping, a standardized text embedding sequence of uniform length is formed, allowing the financial business content and visual style description to have different positions in the sequence.
[0022] In healthcare scenarios, the input text is a health education illustration for chronic disease management, featuring a gentle cartoon style and a light-colored background. After word segmentation, it can be divided into text units such as chronic disease management, health education illustration, gentle cartoon style, and light-colored background. After embedding and mapping, a standardized text embedding sequence is formed, ensuring that the health-related content and image presentation requirements retain distinguishable embedding positions.
[0023] This embodiment segments text prompts into a sequence of text units, preserving the subject, attributes, scene, and style descriptions within continuous text as locationable text units. By mapping the text unit sequence to a standardized text embedding sequence, input text of different lengths and structures can be converted into vector input in a unified format. Consequently, subsequent processing can distinguish between content semantics and visual style based on stable position and fragment information, reducing semantic and style mixing during the input stage.
[0024] S20, identify semantic text units and style text units in the text unit sequence, and generate a semantic candidate position set and a style candidate position set according to the positions of the semantic text units and style text units in the standardized text embedding sequence; In this embodiment, after the text unit sequence is formed, each text unit needs to be assigned a discernible role attribute. Semantic text units are filtered out from text units representing subject objects, attribute modifiers, action relationships, and scene relationships, such as characters, equipment, backgrounds, actions, and object relationships. Stylistic text units are filtered out from text units representing the visual presentation style, such as color, material, lighting, composition, brushstrokes, and artistic style. During filtering, the sequential position of the text unit in the text unit sequence, adjacent text units, fragment boundaries, part-of-speech tags, and context windows can be read, allowing the same text unit to obtain different role judgments in different contexts. For example, cool colors are closer to style text units in visual presentation, and "cold" in cold compress products is closer to attribute modifier content.
[0025] The identification of semantic and stylistic text units can be achieved through role tagging. Role tags can be generated jointly by lexical matching, phrase boundary judgment, context classifier, and embedding similarity discrimination. Lexical matching is used to identify the included subject words, style words, and business words; phrase boundary judgment is used to avoid generating incorrect roles after compound phrases are broken down; the context classifier is used to determine the actual use of the text unit in the local window; and embedding similarity discrimination is used to handle new words or compound words not included in the database. Multiple judgment results can be synthesized into semantic role tags and stylistic role tags, and the case where the same text unit has both semantic and stylistic meanings can be handled by using confidence values.
[0026] The positions in the standardized text embedding sequence are used to convert the text-level recognition results into a set of positions usable for subsequent vector processing. A position mapping table can be established between the text unit sequence and the standardized text embedding sequence, recording the embedding index range corresponding to each text unit. Semantic text units are converted into a set of semantic candidate positions via the position mapping table, and stylistic text units are converted into a set of style candidate positions via the position mapping table. For text units that are split into multiple sub-word vectors, the position set can record continuous embedding intervals; for text units formed by merging multiple phrase fragments, the position set can record multiple discrete embedding intervals. This ensures consistency between the positions of the text recognition results and the standardized text embedding sequence.
[0027] In one implementation, the text unit sequence undergoes vocabulary matching and context window discrimination. The processing end reads the text units preceding and following each text unit and generates a local window based on segment boundaries. The local window is input to a text role classifier, which outputs semantic role probabilities and style role probabilities. When the semantic role probability is higher than the style role probability, the text unit is written into a semantic text unit; when the style role probability is higher than the semantic role probability, the text unit is written into a style text unit; when the two probabilities are close, phrase boundaries and adjacent modification relationships are read for secondary discrimination. The text units and the standardized text embedding sequence are connected through an embedding index table, ultimately converting the semantic text units and style text units into a set of semantic candidate positions and a set of style candidate positions.
[0028] In another implementation, text unit sequences are identified using role prototype vectors. The processing end configures prototype vectors for the main object, attribute relationships, scene relationships, composition, color, lighting, material, and artistic expression. The embedding vector corresponding to each text unit is compared for similarity with each prototype vector, and smoothing is performed by combining the role results of adjacent text units. Text units with high semantic similarity form semantic text units, and text units with high style similarity form style text units. This implementation is suitable for scenarios where prompt text contains new compound words, such as visualizations of yield curves and risk warning illustrations in fintech images, and illustrations of exercise interventions and nutrition advice posters in medical and health images.
[0029] Another approach is to combine rule constraints with a classifier. Rule constraints handle explicit phrases; for example, phrases like "blue background," "low saturation," "hand-drawn style," and "business style" are categorized into style text units, while phrases like "risk warning," "asset allocation," "rehabilitation guidance," and "medication reminder" are categorized into semantic text units. The classifier handles text units that cannot be directly categorized by rules. When rule results and classifier results conflict, the integrity of phrase boundaries is prioritized, and the attribution is determined based on the position of the text unit within the segment. After attribution, a position mapping table converts the attribution results into a set of semantic candidate positions and a set of style candidate positions.
[0030] In fintech business scenarios, input text includes investor education illustrations for low-risk financial products, featuring a simple business style and a blue background. The processing unit can identify low-risk financial products and investor education illustrations as semantic text units, and the simple business style and blue background as style text units. A position mapping table writes the embedding intervals of semantic text units in the standardized text embedding sequence into a semantic candidate position set, and writes the embedding intervals of style text units in the standardized text embedding sequence into a style candidate position set, thus separating financial business content and visual presentation requirements at the position level.
[0031] In healthcare scenarios, input text includes illustrations promoting health knowledge about chronic disease management, featuring a gentle cartoon style and a light-colored background. The processing unit can identify chronic disease management and health education illustrations as semantic text units, and the gentle cartoon style and light-colored background as stylistic text units. The corresponding positions in the standardized text embedding sequence are written into the semantic candidate position set and the stylistic candidate position set, respectively, allowing health-related content and visual style requirements to have different candidate position ranges.
[0032] This embodiment identifies semantic and stylistic text units in a text unit sequence, enabling the separation of the subject, attributes, scene relationships, and visual presentation requirements within the text content into different categories. By mapping semantic and stylistic text units to positions in a standardized text embedding sequence, text-level role judgments can be transformed into a set of positions usable for vector encoding. Consequently, subsequent processing can read the semantic candidate position set and the stylistic candidate position set separately, reducing the mixing of semantic content and stylistic expression at embedding positions.
[0033] S30, based on the standardized text embedding sequence and the semantic candidate position set, aggregate semantic attention to generate a semantic embedding vector, and generate a gating matrix based on the standardized text embedding sequence, the semantic candidate position set and the style candidate position set, and use the gating matrix to perform gating style encoding to generate a style embedding vector; In this embodiment, the standardized text embedding sequence carries the vector, position, and fragment information corresponding to the text units. The semantic candidate position set is used to mark the embedding positions related to the subject object, attribute relationships, action relationships, and scene content. Semantic attention aggregation can be accomplished through query representation, key representation, and value representation. During processing, semantic query representation, semantic key representation, and semantic value representation are generated from the standardized text embedding sequence, and then position masking information is generated based on the semantic candidate position set to concentrate attention on the embedding region corresponding to the semantic candidate position set. The semantic value representation is weighted and aggregated according to the semantic attention distribution to form a semantic embedding vector that can centrally express the subject object, attribute relationships, and scene relationships.
[0034] When semantic candidate positions participate in semantic attention aggregation, position masking, candidate position weighting, and boundary constraints can be used together. Position masking reduces the degree to which non-semantic positions participate in aggregation, candidate position weighting distinguishes the contributions of main subjects, auxiliary attributes, and scene relationships, and boundary constraints preserve the integrity of continuous embeddings within phrases. In this way, content such as financial products, risk warnings, and asset allocation instructions in fintech texts, and chronic disease management, rehabilitation guidance, and medication reminders in medical and health texts, can form a relatively concentrated vector representation in the semantic embedding vector.
[0035] The gating matrix is used to adjust the attention distribution during style encoding. It can be generated jointly from the positional representations in the standardized text embedding sequence, the set of semantic candidate positions, and the set of style candidate positions. The set of semantic candidate positions identifies semantic positions where style attention occupancy needs to be reduced, while the set of style candidate positions identifies style positions where style attention needs to be enhanced. The standardized text embedding sequence provides the vector content and adjacency relationships for each position. The gating matrix can include semantic suppression, style enhancement, and overlap adjustment terms. The semantic suppression term reduces the excessive absorption of main content positions by style encoding; the style enhancement term increases the participation of color, texture, composition, lighting, and artistic expression positions; and the overlap adjustment term handles cases where the same text unit carries both content and style meanings.
[0036] Gated style coding can be adjusted after the style attention distribution is generated. The style coding branch generates style query representations, style key representations, and style value representations from the standardized text embedding sequence, and obtains the initial style attention distribution. The gating matrix is combined with the initial style attention distribution to form the gated style attention distribution. The style value representations are aggregated according to the gated style attention distribution to form a style embedding vector. The style embedding vector mainly carries the visual expression style, such as business style, hand-drawn style, low saturation, blue gradient background, gentle cartoon style, light background, and other visual expression content.
[0037] During encoding branch training or parameter fine-tuning, an adversarial contrastive loss function can be constructed based on semantic embedding vectors and style embedding vectors. Semantic embedding vectors represent the main object, attribute relationships, and scene content in the text prompts, while style embedding vectors represent color, composition, texture, lighting, and artistic expression. To reduce cross-penetration between the two types of embeddings, a similarity discrimination objective can be set, enabling the discriminator to learn the pairing relationship between semantic and style embedding vectors in the same text sample. Adversarial updates then reduce the recognizability of this pairing relationship between the semantic and style encoding branches. This processing can be added during the training phase without changing the order in which semantic and style embedding vectors are used during image generation.
[0038] The semantic style embedding adversarial contrast constraint formula can be expressed as:
[0039] in, denoted as the adversarial contrastive loss function, used to construct a pairwise discriminative objective between semantic embedding vectors and style embedding vectors. N represents the number of text samples in the same batch. i represents the index of the current text sample. j represents the index of the text sample participating in the contrast within the same batch. Let represent the semantic embedding vector corresponding to the i-th text sample. This represents the style embedding vector corresponding to the i-th text sample. This represents the style embedding vector corresponding to the j-th text sample. `s(,)` represents the similarity function, which can use cosine similarity to measure the closeness between two embedding vectors. `τ` represents the temperature parameter, used to adjust the sharpness of the distribution between comparison items; the smaller the temperature parameter, the more obvious the difference in scores between different comparison items; the larger the temperature parameter, the smoother the difference in scores between different comparison items. `exp()` represents the exponential mapping, used to convert the similarity results into non-negative weights. Numerator term This represents the pairwise similarity weight between semantic embedding vectors and style embedding vectors within the same text sample. (Denominator term) This represents the sum of the comparative similarity weights between the current semantic embedding vector and the style embedding vectors in the same batch.
[0040] During training, this loss function can be used in the similarity discriminator or adversarial branch, enabling the similarity discriminator to learn the pairing relationship between semantic embedding vectors and style embedding vectors. The semantic encoding branch and style encoding branch, through backward gradient updates, reduce the discriminator's ability to recognize the pairing relationship between semantic and style embedding vectors. This adversarial constraint reduces the components directly related to style expression in the semantic embedding vectors and the components directly related to the main content in the style embedding vectors, thus allowing subsequent gated style encoding to obtain style embedding vectors with clearer sources.
[0041] In one implementation, a standardized text embedding sequence is processed through a semantic encoding branch to generate a semantic query representation, a semantic key representation, and a semantic value representation. The set of semantic candidate positions is converted into a position masking vector with the same embedding length, which is then expanded into a semantic position masking matrix. The semantic position masking matrix is combined with the attention result formed by the semantic query representation and the semantic key representation to obtain a semantic attention distribution. This semantic attention distribution is then aggregated with the semantic value representation to generate a semantic embedding vector. This implementation is suitable for scenarios where the main object in the text prompt is relatively clear, such as insurance claim flowcharts and credit risk control prompts in fintech businesses, and also suitable for medication reminders and rehabilitation guidance diagrams in healthcare businesses.
[0042] Alternatively, a hierarchical weighted approach to candidate positions can be used. The semantic candidate position set is divided into subject positions, attribute positions, and relation positions, with different candidate weights assigned to different positions. Subject positions correspond to the main object, attribute positions correspond to the modifying content, and relation positions correspond to the action or scene relationship between objects. During semantic attention aggregation, subject positions receive a higher participation weight, while attribute and relation positions participate in supplementation according to the text order. This implementation is suitable for describing complex text prompts, such as educational illustrations for asset allocation in fintech businesses, where the text simultaneously includes product objects, risk relationships, and explanatory scenarios; it is also suitable for popular science illustrations for chronic disease management in healthcare businesses, where the text simultaneously includes population, behavior, health goals, and scene relationships.
[0043] Alternatively, a gating matrix partitioning method can be used. Each embedding position in the normalized text embedding sequence is assigned a position vector and a fragment vector. The semantic candidate position set generates a semantic suppression vector, and the style candidate position set generates a style enhancement vector. The two vectors, along with the position overlap information, are synthesized into a gating matrix. After the style encoding branch generates the initial style attention distribution, the gating matrix adjusts the initial style attention distribution, prioritizing the aggregation of style fragments in the style value representation. This implementation is suitable for texts with extensive style descriptions, such as concise business style, data visualization style, gentle cartoon style, and low-saturation illustrations.
[0044] In fintech business scenarios, text prompts include investor education illustrations of low-risk financial products, featuring a simple business style and a blue gradient background. A semantic candidate position set corresponds to the content locations of low-risk financial products and investor education illustrations, and semantic attention aggregation generates semantic embedding vectors. A style candidate position set corresponds to style locations such as the simple business style and the blue gradient background. A gating matrix reduces the occupancy of semantic locations related to low-risk financial products in style encoding and enhances locations related to the business style and blue background, generating style embedding vectors.
[0045] In healthcare scenarios, text prompts include illustrations promoting health education on chronic disease management, featuring a gentle cartoon style and a light-colored background. A semantic candidate location set corresponds to the content locations such as chronic disease management and health education illustrations, and semantic attention aggregation generates semantic embedding vectors. A style candidate location set corresponds to visual expression locations such as the gentle cartoon style and light-colored background. A gating matrix adjusts the attention distribution in style encoding, ensuring that style embedding vectors originate primarily from style candidate locations, reducing interference between health-related content and style expression.
[0046] This embodiment constrains semantic attention aggregation by using a set of semantic candidate positions, allowing semantic embedding vectors to come more from embedding positions corresponding to the main object, attribute relationships, and scene content. By generating a gating matrix through a standardized text embedding sequence, a set of semantic candidate positions, and a set of style candidate positions, the style encoding process can reduce the interference of the main content position on style expression and increase the participation of style candidate positions. By adjusting the style attention distribution through the gating matrix, the style embedding vectors can more centrally carry information on color, composition, material, lighting, and artistic expression, thereby reducing the mixing of semantic content and visual style during the encoding stage.
[0047] S40, determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; In this embodiment, the target image generation model needs to have a cross-attention layer. The cross-attention layer receives text side vectors and participates in conditional injection during the image generation process. When determining the target image generation model, the model structure configuration, attention layer identifier, text conditional input port, and intermediate latent space specifications can be read. The compatibility space is determined by the input formats that the cross-attention layer can receive, typically involving vector length, number of channels, arrangement direction, batch dimension, and conditional input position. The purpose of the compatibility space is to ensure that semantic embedding vectors and style embedding vectors can enter the same generated side vector format, avoiding inconsistencies between the text encoding side vector dimension and the image generation side conditional input dimension.
[0048] The mapping of semantic embedding vectors to the compatibility space can be accomplished through a projection layer, a normalization layer, and residual connections. The projection layer converts the dimension of the semantic embedding vectors to the number of channels required by the compatibility space, the normalization layer adjusts the vector distribution, and the residual connections preserve the main semantic structures in the subject objects, attribute relationships, and scene relationships. The compatibility semantic representation is formed by the mapped semantic vectors and has a format that matches the text conditional input of the cross-attention layer. For text content such as risk warnings, insurance claims, and asset allocation instructions in fintech scenarios, the compatibility semantic representation carries financial business objects and business relationships. For text content such as rehabilitation guidance, medication reminders, and nutritional advice in healthcare scenarios, the compatibility semantic representation carries health objects, behavioral relationships, and scene content.
[0049] The mapping of style embedding vectors to a compatible space needs to be adjusted by compatible semantic representation. Compatible semantic representation can generate semantic conditional gating vectors, semantic bias vectors, or conditional scaling vectors, used to change the channel weights and positional weights after the style embedding vectors are projected. The style embedding vectors, after style projection, form an initial style projection representation, which is then adjusted by the conditional control variables formed by the compatible semantic representation to obtain the compatible style representation. The compatible style representation and the compatible semantic representation reside in the same compatible space, preserving stylistic content such as color, composition, texture, lighting, and artistic expression, while reducing deviations in style mapping from semantic content.
[0050] In one implementation, the processing end reads the cross-attention layer configuration of the target image generation model to obtain the vector length and number of channels of the text conditional input. The semantic embedding vector undergoes linear projection and normalization to form a semantic projection result consistent with the text conditional input. The semantic projection result and the semantic embedding vector are then residually fused to obtain a compatible semantic representation. The style embedding vector is passed through another projection layer to form a style projection result. The compatible semantic representation is passed through a gating layer to generate a semantic conditional gating vector. The semantic conditional gating vector is multiplied or weighted by channels with the style projection result to obtain a compatible style representation.
[0051] Layered projection can also be used. When the target image generation model contains multiple cross-attention layers, different layers can have different channel configurations. Semantic embedding vectors are mapped to compatible semantic representations at multiple levels, and style embedding vectors are mapped to style projection results at multiple levels. Layers closer to the image structure use stronger semantic conditional control, while layers closer to the texture and color use stronger style conditional control. This implementation is suitable for generation tasks that require maintaining both object structure and visual style, such as controlling icon structure and business style in fintech promotional images, and controlling character actions and mild style in medical and health science popularization images.
[0052] Alternatively, a shared compatibility space can be used. The processing unit projects the semantic embedding vector and the style embedding vector onto the same vector dimension, employing the same length calibration strategy. The compatible semantic representation is used to generate a semantic bias vector, which is then added to the style projection result, ensuring that the style projection result is constrained by semantic content within the same compatibility space. This implementation is suitable for deployment environments with limited inference resources and can reduce the computational overhead of multi-layer projection.
[0053] In fintech business scenarios, text content includes illustrations of risk warnings for wealth management products and a clean, business-like style. After the cross-attention layer of the target image generation model reads the text's conditional input specifications, the semantic embedding vector is mapped to a compatible semantic representation to preserve business content such as wealth management products, risk warnings, and investor education. The style embedding vector, adjusted by the compatible semantic representation, generates a compatible style representation, ensuring that visual requirements such as a clean, business-like style, blue background, and low saturation are consistent with the financial business content.
[0054] In healthcare scenarios, text content includes illustrations for rehabilitation training guidance and a gentle cartoon style. Semantic embedding vectors are mapped to a compatible semantic representation to preserve rehabilitation actions, training objects, and health guidance scenarios. Style embedding vectors, adjusted by the compatible semantic representation, generate a compatible style representation, ensuring that the gentle cartoon style, light background, and soft lighting surround the health content, reducing the shift in action relationships caused by style transitions.
[0055] This embodiment determines the compatibility space based on the cross-attention layer, enabling semantic embedding vectors and style embedding vectors to be converted into an input format acceptable to the target image generation model. By mapping semantic embedding vectors to compatible semantic representations, the subject object, attribute relationships, and scene content remain stable before entering the image generation condition input. By adjusting the mapping of style embedding vectors to the compatibility space with compatible semantic representations, style content is constrained by semantic content during the conversion process, thereby reducing the influence of style expression on semantic content deviation.
[0056] S50, obtain style intensity parameters, scale the compatible style representation according to the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; In this embodiment, the style intensity parameter controls the degree to which the compatible style representation participates in the combined text representation. The style intensity parameter can come from the style intensity option in the user interface, a preset configuration file, or be derived from degree words in text prompts. After processing, the style intensity parameter can be converted into a scaling factor consistent with the dimensions of the compatible style representation, a channel-distributed scaling vector, or a position-distributed scaling matrix. The scaling factor is suitable for overall enhancement or reduction of style expression, the scaling vector is suitable for situations where different channels have different style contributions, and the scaling matrix is suitable for situations where different text positions correspond to different style intensities.
[0057] The compatible style representation is already within the vector space that the target image generation model can accept. When scaling the compatible style representation, the vector length and channel arrangement can remain unchanged, only the vector magnitude, channel weights, or positional weights can be changed. The scaled compatible style representation still retains the original style meaning, but the degree of style contribution is controlled by the style intensity parameter. For cases with low style intensity parameters, the impact of the scaled compatible style representation on the combined text representation is weakened; for cases with high style intensity parameters, the participation of the scaled compatible style representation in the combined text representation is enhanced.
[0058] When combining compatible semantic representations and scaled compatible style representations, consistency in vector length, channel arrangement, and input format must be maintained. Combination methods can include concatenation, interleaving, segmentation, or labeled parallel arrangement. Concatenation is suitable for preserving the complete boundaries of semantic and style representations; interleaving is suitable for situations where semantic and style fragments need to correspond positionally; and segmentation is suitable for target image generation models that read input conditions by fragment. During the combination process, it's possible to record whether each combined fragment originates from the compatible semantic representation or the scaled compatible style representation, enabling subsequent read structures to distinguish the content source.
[0059] Component identifiers are used to mark the different sources of components in the combined text representation. Component identifiers can be position labels, channel labels, fragment labels, or masking labels. Position labels mark the sequence range of semantic and style fragments in the combined text representation; channel labels mark semantic and style channels; fragment labels mark the fragment categories in the segmented arrangement structure; and masking labels control the visible area during subsequent reading. After the component identifiers are configured, the combined text representation simultaneously contains a compatible semantic representation, a scaled compatible style representation, and source-distinguishing labeling information, ensuring that the combined input still retains the semantic and stylistic separation boundaries.
[0060] In one implementation, the style intensity parameter is converted into a scalar scaling factor. The processor multiplies each style vector in the compatible style representation by the scalar scaling factor to obtain the scaled compatible style representation. The compatible semantic representation and the scaled compatible style representation are concatenated in a segmented manner, with the first segment recorded as the semantic component and the second segment as the style component. Component identifiers are configured using fragment tags, with a semantic tag or style tag written at each position, ultimately forming a combined text representation. This implementation is suitable for generation services with high response speed requirements.
[0061] Channel-based scaling can also be used. The processor generates channel scaling vectors based on the style intensity parameters, and each channel scaling vector corresponds one-to-one with a channel in the compatible style representation. Different scaling values can be obtained for the style texture channel, color channel, composition channel, and lighting channel. The scaled compatible style representation and compatible semantic representation are combined in a channel-side manner, and the component identifiers are configured using channel labels. This implementation is suitable for image generation scenarios that require separate adjustment of the contributions of color, lighting, and composition.
[0062] Positional scaling can also be used. The processor generates a positional scaling matrix based on the positions of semantic fragments in the compatible semantic representation and style fragments in the compatible style representation. The positional scaling matrix controls the influence range of different style fragments on different semantic fragments. The scaled compatible style representation and compatible semantic representation are arranged alternately according to corresponding fragments, and the component identifiers are configured using both position labels and fragment labels. This implementation is suitable for situations where the text prompt information contains multiple main objects and multiple style requirements, such as image generation tasks where different objects require different visual representations.
[0063] In fintech business scenarios, text prompts include asset allocation diagrams and a concise business style. When the style intensity parameter is low, the compatible style representation is weakened, and the compatible semantic representation corresponding to the financial business content in the combined text representation makes the main contribution, keeping the relationship between asset allocation, risk warnings, and explanations clear. When the style intensity parameter is high, the scaled compatible style representation enhances the business feel, layout, and low-saturation background, while the component identifiers retain the boundaries between semantic and style components, ensuring that the financial semantics can still be distinguished and read even when the business style is enhanced.
[0064] In healthcare scenarios, text prompts include illustrations of rehabilitation training guidance and a gentle cartoon style. The style intensity parameter controls the degree to which cartoon lines, light backgrounds, and soft lighting are incorporated into the compatible style representation. The compatible semantic representation retains the rehabilitation actions, training objects, and guidance scenarios, while the scaled compatible style representation maintains the gentle presentation requirements. The combination of these two elements is configured with component identifiers, ensuring that the health content and style representation each have distinct positions, reducing the risk of style enhancement shifting the meaning of actions.
[0065] This embodiment controls the amplitude, channel weight, or position weight of the compatible style representation through style intensity parameters, allowing the style contribution level to be adjusted before entering the combined text representation. By combining the compatible semantic representation and the scaled compatible style representation with component identifiers, the combined text representation retains the source boundaries of both semantic and style information while merging them. Therefore, subsequent reading structures can distinguish between semantic and style components, reducing the interference of style intensity changes on semantic content and improving the stability of style expression adjustment.
[0066] S60, the combined text representation is input into the target image generation model, and the cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet the preset output conditions, the target image is output.
[0067] In this embodiment, after the combined text representation enters the target image generation model, component identifiers are used to distinguish the position range, segment category, or channel category of the compatible semantic representation and the scaled compatible style representation in the input sequence. When the cross-attention layer reads the combined text representation, it can establish semantic reading regions and style reading regions based on the component identifiers. The semantic reading region provides conditional information such as the main object, attribute relationships, and scene relationships, while the style reading region provides conditional information such as color, material, composition, lighting, brushstrokes, and artistic expression. The cross-attention layer can generate different attention weights for the two types of reading regions and then write the reading results into the image generation process, so that the candidate image is simultaneously constrained by semantic content and style expression.
[0068] Candidate images can be generated using a diffusion-based generation structure, a latent variable decoding structure, or a conditional image decoding structure. A combined text representation provides the textual conditional input, and the target image generation model forms an intermediate image representation based on the readings from the cross-attention layer, before generating candidate images. After candidate images are generated, semantic and stylistic representations need to be extracted from the image content. Semantic representations can come from a visual encoder, region feature extractor, or image-text matching network, reflecting the subject object, attribute relationships, and scene content. Stylistic representations can come from color statistics, texture encoding, compositional features, illumination distribution, or a style encoder, reflecting the image's expressive style.
[0069] Semantic consistency and style consistency analysis are used to determine the degree of matching between candidate images and text prompts. Text prompts can generate semantic reference representations and style reference representations. Semantic consistency is obtained by comparing the image semantic representation with the text semantic reference representation, and style consistency is obtained by comparing the image style representation with the text style reference representation. Image analysis results can be formed by both semantic and style consistency. Preset output conditions can include semantic consistency reaching a set range, style consistency reaching a set range, or the combined analysis result meeting the output requirements. When the image analysis result meets the preset output conditions, the candidate image is determined as the target image and output.
[0070] In one implementation, component identification uses fragment labels. The combined text representation is divided into semantic fragments and style fragments, and a cross-attention layer reads the semantic fragments and style fragments respectively. Semantic fragments generate semantic conditional features, and style fragments generate style conditional features. The target image generation model injects these two types of conditional features into the intermediate image representation to generate candidate images. The candidate images enter the visual encoder to obtain the image semantic representation, and enter the style encoder to obtain the image style representation. The text prompt information is processed by the text encoder to form a text semantic reference representation and a text style reference representation, which are compared with the image side representation to generate the image analysis result.
[0071] A layered reading approach can also be used. In the target image generation model, the shallow cross-attention layer reads more compatible semantic representations, while the deep cross-attention layer reads more scaled compatible style representations. The shallow reading results constrain the subject outline, object relationships, and scene layout, while the deep reading results constrain texture, color, and lighting. This implementation is suitable for image generation tasks involving complex subjects and clear style requirements, such as asset allocation illustrations in fintech, which need to maintain clarity of chart elements, risk warning text, and business relationships while presenting a concise business style. It is also suitable for rehabilitation action illustrations in healthcare, which need to maintain clear action postures and health guidance scenes while presenting a soft cartoon style.
[0072] A dual-path analysis approach can also be used. After candidate images are generated, the semantic analysis branch detects the main object, scene relationships, and key attributes, while the style analysis branch extracts color distribution, texture patterns, and compositional features. In fintech, the semantic analysis branch can detect elements such as financial product descriptions, risk warning scenarios, and insurance claim processes, while the style analysis branch can detect business-style elements, low saturation, and data visualization. In healthcare, the semantic analysis branch can detect rehabilitation actions, medication prompts, and health management themes, while the style analysis branch can detect gentle color tones, cartoon lines, and light-colored backgrounds.
[0073] In fintech business scenarios, the text prompt includes an illustration of low-risk financial products for investor education, featuring a simple business style and a blue background. The cross-attention layer reads semantic content such as "low-risk financial products" and "investor education" from the compatible semantic representation based on component identifiers, and also reads the simple business style and blue background from the scaled compatible style representation. After candidate images are generated, semantic consistency analysis checks whether the financial business object and the educational scenario match the text prompt, and style consistency analysis checks whether the business style and blue background match the style description. The target image is output when the image analysis results meet the preset output conditions.
[0074] In healthcare scenarios, text prompts include illustrations for chronic disease management and health education, featuring a gentle cartoon style and a light-colored background. A cross-attention layer reads the compatible semantic representation of the health topic and the scaled compatible style representation of the visual presentation. After candidate images are generated, semantic consistency analysis checks the content related to chronic disease management, health education, and the scene; style consistency analysis checks the gentle cartoon style and light-colored background; and the target image is output when the image analysis results meet preset output conditions.
[0075] In this embodiment, component identification enables the cross-attention layer to read both the compatible semantic representation and the scaled compatible style representation, allowing the image generation process to distinguish between semantic and style conditions on the input side. Through semantic and style consistency analysis between candidate images and text prompts, candidate images undergo content matching and style matching before output. By comparing the image analysis results with preset output conditions, the target image output is constrained by both semantic preservation and style matching, thereby reducing semantic content shift and style expression imbalance.
[0076] In one embodiment, step S10 includes: S101, extract the original character sequence from the text prompt information, and generate an original character position table based on the arrangement order of each character in the original character sequence; S102, Generate a standard character sequence based on the original character sequence, and retain separator characters, modifier characters and phrase connector characters used to identify semantic boundaries and style boundaries in the standard character sequence; S103, adjust the standardized character sequence according to the preset length boundary to generate a fixed-length standardized character sequence; S104, Generate a text unit boundary table based on the fixed-length standard character sequence, and mark the boundary positions of continuous phrases, modifying phrases and parallel phrases in the text unit boundary table; S105, based on the text unit boundary table, the fixed-length standard character sequence is segmented to generate a text unit sequence, and a source position marker and a boundary marker are configured for the text units in the text unit sequence. The source position marker is used to record the position range of the text unit originating from the text prompt information in the original character position table, and the boundary marker is determined by the text unit boundary table. S106, the text unit sequence is mapped to a text unit vector sequence according to a preset text unit vocabulary, a position representation is generated according to the source position marker, a fragment representation is generated according to the text unit boundary table, and a boundary representation is generated according to the boundary marker; S107, the text unit vector sequence, the position representation, the segment representation and the boundary representation are fused to generate a standardized text embedding sequence.
[0077] In this embodiment, after the text prompt information is read, the character content retains its original arrangement order, forming an original character sequence. The original character sequence can contain Chinese characters, English words, numbers, punctuation, hyphens, parentheses, spaces, and decorative symbols. An original character position table is generated based on the position of each character in the original character sequence, used to record the positional correspondence between characters and the input text. When generating the original character position table, a character index, character type, and adjacent character relationship can be configured for each character. The character index is used to subsequently determine the source range of the text unit, the character type is used to distinguish between text, punctuation, spaces, and hyphens, and the adjacent character relationship is used to determine whether characters belong to the same phrase segment.
[0078] The canonical character sequence is obtained by tidying up the original character sequence. During the tidying process, repeating whitespace characters, invisible control characters, and format specifiers unrelated to the image content can be deleted or replaced, while delimiters, modifier characters, and phrase connectors that can identify semantic and style boundaries are retained. Delimiters can be used to distinguish different descriptive fragments, modifier characters can be used to preserve attribute modification relationships, and phrase connectors can be used to maintain the integrity of compound phrases. After retaining these characters in the canonical character sequence, consecutive phrases, modifier phrases, and parallel phrases are less likely to be incorrectly broken up during subsequent segmentation.
[0079] Preset length boundaries are used to limit the length range of the canonical character sequence. When the canonical character sequence exceeds the preset length boundary, it can be truncated according to the importance of text units, phrase integrity, and character positional relationships, retaining the main description, key modifiers, and style descriptions. When the canonical character sequence is shorter than the preset length boundary, placeholder characters or padding markers can be added to ensure consistent length conditions for subsequent word segmentation and embedding mapping. Fixed-length canonical character sequences are formed by truncating or padding the canonical character sequence, maintaining the original positional relationships of valid characters and providing consistent input for subsequent vectorization.
[0080] The text unit boundary table is generated based on a fixed-length, standardized character sequence. The boundary positions of consecutive phrases mark the complete meaning fragment formed by multiple adjacent characters; the boundary positions of modifying phrases mark the combination range between the modifier and the modified object; and the boundary positions of parallel phrases mark the separation range between multiple parallel descriptions. The text unit boundary table can record the boundary start position, boundary end position, boundary type, and boundary priority. Boundary priority is used to select the granularity to retain when multiple boundaries overlap; for example, the overall boundary of a compound phrase takes precedence over single-character segmentation boundaries.
[0081] The text unit sequence is obtained by segmenting a fixed-length, standardized character sequence according to a text unit boundary table. Text units can be words, subwords, phrase fragments, compound descriptive fragments, or placeholder fragments. Each text unit in the sequence is configured with a source location marker and a boundary marker. The source location marker records the position range of the text unit originating from the text prompt information in the original character location table, used to maintain a traceable relationship between the text unit and the input text. The boundary marker, determined by the text unit boundary table, identifies whether the text unit belongs to a continuous phrase, modifying phrase, parallel phrase, or placeholder fragment. For placeholder fragments formed by padding, the source location marker can be recorded as an empty position or a placeholder position, and the boundary marker can be recorded as a filled boundary.
[0082] A pre-defined text unit vocabulary is used to convert a sequence of text units into a sequence of text unit vectors. The pre-defined text unit vocabulary can include word indexes, subword indexes, phrase indexes, and special tag indexes. When a text unit has a corresponding entry in the vocabulary, the corresponding vector can be read directly; when a text unit does not match the vocabulary, it can be split into subword fragments and the subword vectors can be combined; when a text unit is a placeholder fragment, the filler vector can be read. The text unit vector sequence is generated according to the order of arrangement in the text unit sequence, and the vector dimensions remain consistent.
[0083] Position representations are generated based on source location markers and are used to preserve the relative position and source range of text units in the input text. Fragment representations are generated based on text unit boundary tables and are used to distinguish fragment types such as continuous phrases, modifying phrases, and parallel phrases. Boundary representations are generated based on boundary markers and are used to enhance the boundary attributes of text units, enabling the text encoder to identify whether a text unit is located inside a phrase, at the edge of a phrase, or in a parallel separating position. Text unit vector sequences, position representations, fragment representations, and boundary representations can be combined using vector summation, vector concatenation followed by projection, gated fusion, or normalized fusion to generate a standardized text embedding sequence. The standardized text embedding sequence has a uniform length and uniform vector dimension and preserves the source location, fragment structure, and boundary attributes of the text units.
[0084] This embodiment records the source positions of characters through an original character position table, ensuring a clear positional correspondence between text units and the input text. It preserves separator characters, modifier characters, and phrase connector characters through a standardized character sequence, ensuring that semantic and stylistic boundaries remain identifiable after character organization. A fixed-length standardized character sequence unifies the input length, enabling subsequent word segmentation and embedding mapping to be completed under consistent length conditions. A text unit boundary table segments the fixed-length standardized character sequence, allowing continuous phrases, modifier phrases, and parallel phrases to enter the vectorization process as relatively complete text units. Through the combined participation of source position markers, boundary markers, position representations, fragment representations, and boundary representations, the standardized text embedding sequence simultaneously preserves text content, positional relationships, and boundary structures, thereby reducing confusion regarding semantic content and stylistic descriptions during the input stage.
[0085] In one embodiment, step S20 above includes: S201, Read the sequential position of each text unit and adjacent text units in the text unit sequence, and determine the boundary marker according to the boundary position of each text unit in the text unit sequence, and generate a text unit context window set; S202, Generate semantic role tags, style role tags, and role confidence tags for each text unit based on the text unit context window set; S203, based on the semantic role marker, filter object text units, attribute text units, action text units, scene text units and relation text units from the text unit sequence, and merge the object text units, attribute text units, action text units, scene text units and relation text units into a semantic text unit candidate set; S204, based on the style role marker, filter style expression text units, composition text units, material text units, lighting text units, color text units and artistic expression text units from the text unit sequence, and merge the style expression text units, the composition text units, the material text units, the lighting text units, the color text units and the artistic expression text units into a style text unit candidate set; S205, for a text unit that simultaneously has the semantic role tag and the style role tag, a master attribution tag is generated based on the role confidence tag and the boundary tag, and the text unit that simultaneously has the semantic role tag and the style role tag is retained in the semantic text unit candidate set or the style text unit candidate set based on the master attribution tag, and removed from the other candidate set that has never been retained; S206, the set of candidate semantic text units is determined as semantic text units, and the set of candidate style text units is determined as style text units; S207, Establish the embedding position mapping relationship between the text unit sequence and the standardized text embedding sequence; S208, determine the position of the semantic text unit in the standardized text embedding sequence according to the embedding position mapping relationship, generate a semantic candidate position set, and determine the position of the style text unit in the standardized text embedding sequence according to the embedding position mapping relationship, generate a style candidate position set.
[0086] In this embodiment, each text unit in the text unit sequence has an ordered position and adjacent text units. The ordered position indicates the order in which the text units are arranged in the text input, while the adjacent text units provide local context. The boundary position indicates whether the text unit is located inside a phrase, at the edge of a phrase, or at a phrase separator. After generating boundary markers based on the ordered position, adjacent text units, and boundary positions, a text unit can be judged within its surrounding context. For example, a cool color tone appearing alone is more likely to be judged as visual content; it retains its style attribute when appearing against a cool color background; and when it appears in the context of cold compresses, it should be judged as object or attribute content in conjunction with adjacent text units. The text unit context window set can use a fixed window length or a phrase boundary adaptive window. A fixed window length is suitable for inputting shorter prompt text, while a phrase boundary adaptive window is suitable for input text containing compound phrases and parallel descriptions.
[0087] Semantic role tags, style role tags, and role confidence tags can be generated jointly from word meaning vectors, boundary markers, relationships between adjacent text units, and fragment types within the context window. Semantic role tags record whether a text unit carries the meaning of an object, attribute, action, scene, or relationship. Style role tags record whether a text unit carries the meaning of style expression, composition, material, lighting, color, or artistic expression. Role confidence tags record the credibility of role judgments and provide numerical basis for subsequent processing of dual-role text units. Role generation can utilize classifier output, role prototype vector comparison, or a fusion of lexical matching results and contextual discrimination results. Classifier output is suitable for text units where word meaning changes with context, role prototype vector comparison is suitable for unlisted words and compound words, and lexical matching results are suitable for well-structured phrase fragments.
[0088] Object text units, attribute text units, action text units, scene text units, and relation text units are selected from the text unit sequence based on semantic role tags. Object text units typically represent the subject or object in the generated image; attribute text units describe the object's color, shape, quantity, or state; action text units describe behavior or posture; scene text units describe the environment; and relation text units describe the positional, subordinate, inclusive, or interactive relationships between objects. After the selected text units are merged into a semantic text unit candidate set, the semantic content is centralized in the same candidate set, facilitating the subsequent unified determination of its position range in the standardized text embedding sequence.
[0089] Style expression text units, composition text units, material text units, lighting text units, color text units, and artistic expression text units are selected from the text unit sequence based on style role tags. Style expression text units record the overall style of the image; composition text units record the layout and perspective; material text units record surface texture; lighting text units record brightness, direction, and atmosphere; color text units record the primary color, color scheme, and saturation; and artistic expression text units record the drawing method, brushstrokes, or artistic style. After merging the selected text units into a style text unit candidate set, the image expression requirements are concentrated in the same candidate set, reducing the mixing of style content and semantic content at the candidate level.
[0090] When a text unit has both semantic role and style role tags, a primary attribution tag needs to be generated based on role confidence tags and boundary tags. Role confidence tags reflect the relative strength of the semantic and style roles, while boundary tags indicate whether the text unit is within a compound phrase. The primary attribution tag determines whether the text unit is retained in the semantic or style text unit candidate set, and is not removed from the other candidate set that was never retained. For text units within the boundaries of style phrases, the primary attribution tag may favor the style text unit candidate set; for text units within subject noun phrases or action phrases, the primary attribution tag may favor the semantic text unit candidate set. This avoids the same text unit appearing repeatedly in both candidate sets, preventing subsequent position overlap.
[0091] The candidate set of semantic text units is determined as semantic text units after dual-role processing, and the candidate set of style text units is determined as style text units after dual-role processing. The embedding position mapping relationship between the text unit sequence and the standardized text embedding sequence is used to record the transformation result from text unit to embedding position. A text unit can correspond to one embedding position or multiple consecutive embedding positions; a text unit composed of multiple sub-words can correspond to one embedding interval; a text unit formed by parallel phrases can correspond to multiple discrete embedding intervals. Based on the embedding position mapping relationship, the embedding positions corresponding to semantic text units are written into the semantic candidate position set, and the embedding positions corresponding to style text units are written into the style candidate position set. Thus, the role recognition result at the text level is converted into the position result in the standardized text embedding sequence.
[0092] This embodiment generates a text unit context window set by sequential position, adjacent text units, and boundary markers, enabling individual text units to complete role determination based on local context. Text units are classified using semantic role markers, style role markers, and role confidence markers, allowing for the separation of subject objects, attribute relationships, action scenes, and visual content such as color, material, lighting, and composition. Text units with both semantic and style role markers are processed using master attribution markers, reducing duplicate attribution among candidate sets. Semantic candidate position sets and style candidate position sets are generated by embedding position mapping relationships, converting semantic and style content at the text level into different position ranges in a vector sequence, thereby reducing semantic and style mixing in subsequent encoding stages.
[0093] In one embodiment, step S30 above includes: S301, the standardized text embedding sequence is input into the semantic encoding branch to generate a semantic query representation, a semantic key representation, and a semantic value representation; S302, generate a semantic location masking matrix and a candidate semantic weight vector based on the semantic candidate location set, and generate a semantic attention distribution based on the semantic query representation, the semantic key representation, the semantic value representation, the semantic location masking matrix and the candidate semantic weight vector; S303, based on the semantic attention distribution, aggregate the semantic value representations corresponding to the semantic candidate position set to generate a semantic candidate aggregate representation, and convert the semantic candidate aggregate representation into a semantic embedding vector; S304, the standardized text is embedded into the sequence input style encoding branch to generate a style query representation, a style key representation, and a style value representation, and an initial style attention distribution is generated based on the style query representation and the style key representation; S305, generate a semantic suppression vector based on the standardized text embedding sequence and the semantic candidate position set, generate a style enhancement vector based on the standardized text embedding sequence and the style candidate position set, and generate a candidate position overlap marker based on the position overlap relationship between the semantic candidate position set and the style candidate position set; S306, Generate a gating matrix based on the semantic suppression vector, the style enhancement vector, and the candidate position overlap marker; S307, the initial style attention distribution is adjusted using the gating matrix to generate a gated style attention distribution, and the style value representation is aggregated based on the gated style attention distribution to generate a style embedding vector.
[0094] In this embodiment, after the standardized text embedding sequence is input into the semantic encoding branch, it can be processed through a linear projection layer to form a semantic query representation, a semantic key representation, and a semantic value representation, respectively. The semantic query representation is used to initiate association retrieval between semantic locations, the semantic key representation is used to provide a matching semantic index, and the semantic value representation is used to carry aggregateable vector content. The three types of representations can be obtained from the same standardized text embedding sequence through different projection parameters, or they can be obtained by combining a shared projection layer and a branch projection layer. The semantic encoding branch can adopt a multi-head attention structure, so that different attention heads focus on different semantic regions such as the subject object, attribute relationships, action relationships, and scene relationships.
[0095] A set of semantic candidate locations is used to form a semantic location masking matrix and candidate semantic weight vectors. The semantic location masking matrix can be constructed according to the length of the standardized text embedding sequence, retaining the markers involved in the attention operation in the regions corresponding to semantic candidate locations, and configuring suppression markers in the regions corresponding to non-semantic candidate locations. The candidate semantic weight vectors are generated based on the role category, boundary integrity, and candidate confidence of different locations in the semantic candidate location set, assigning different weights to main object locations, key attribute locations, and relationship locations. Semantic query representation, semantic key representation, semantic value representation, along with the semantic location masking matrix and candidate semantic weight vectors, jointly participate in the generation of semantic attention distribution, concentrating attention on the embedding regions corresponding to the semantic candidate location set and reducing the influence of irrelevant locations on semantic aggregation.
[0096] Semantic attention distributions are used to aggregate semantic value representations corresponding to a set of semantic candidate positions. During aggregation, the semantic value representations can be weighted and summed according to the semantic attention distribution, or local semantic aggregation results can be formed separately within multiple attention heads before concatenation and projection. The aggregated semantic candidate representations can preserve the relative contribution relationships between different semantic positions and are converted into semantic embedding vectors through projection layers, normalization layers, or residual fusion layers. The semantic embedding vectors are formed by aggregating the semantic value representations corresponding to the set of semantic candidate positions, which can reduce the interference of style words, color words, material words, etc., on semantic expression.
[0097] After standardizing the text embedding sequence input style encoding branch, style query representation, style key representation, and style value representation can be generated. The style query representation is used to retrieve style-related locations, the style key representation provides style location indexes, and the style value representation carries aggregateable style content. The style query representation and style key representation generate an initial style attention distribution. Before being adjusted by a gating matrix, this distribution may still be influenced by object words, action words, and scene words; therefore, further gating control is needed by combining the semantic candidate location set and the style candidate location set.
[0098] Semantic suppression vectors can be generated based on a standardized text embedding sequence and a set of semantic candidate positions. During generation, the embedding content, position range, and adjacent position relationships corresponding to the semantic candidate position set can be read. Higher suppression values are configured at semantic positions, and lower suppression values are configured at non-semantic positions. Style enhancement vectors can be generated based on a standardized text embedding sequence and a set of style candidate positions. Enhancement values are configured at style candidate positions, and attenuation values are configured at non-style positions. Candidate position overlap markers are generated based on the positional overlap relationships between the semantic and style candidate position sets, used to handle cases where the same embedding position simultaneously possesses both semantic and stylistic meanings. Independent adjustment markers can be configured at overlapping positions, so that the gating matrix does not simply suppress or enhance at that position, but instead performs a transitional adjustment based on the degree of overlap.
[0099] The gating matrix can be obtained by combining semantic suppression vectors, style enhancement vectors, and candidate position overlap markers. The combination method can employ position-wise multiplication, weighted superposition, gated projection, or normalized mapping. The row and column dimensions of the gating matrix can be the same as the initial style attention distribution, or the positional gating vectors can be generated first and then expanded into a matrix of the same size as the initial style attention distribution. Regions in the gating matrix corresponding to semantic candidate positions are used to reduce the attention given to semantic positions in the initial style attention distribution, regions corresponding to style candidate positions are used to increase the attention given to style positions in the initial style attention distribution, and regions corresponding to overlapping positions are used for balancing adjustments.
[0100] The gating matrix adjusts the initial style attention distribution to generate a gated style attention distribution. Adjustment can be achieved through element-wise multiplication, attention bias stacking, or pre-gating before normalization. The gated style attention distribution is then used to aggregate style value representations, forming a style embedding vector. This vector is primarily generated by the style value representations corresponding to the style candidate positions, preserving necessary local context. This ensures that style expression focuses on content related to color, lighting, material, composition, and artistic expression, while reducing the infiltration of the subject object and scene relationships into the style encoding results.
[0101] For example, the formula for generating the candidate position gating matrix is:
[0102] in, This represents the gating matrix. This represents the activation function, which constrains the values in the gating matrix to a range that can be used for attention tuning. This represents the gating projection weights, used to map the input candidate position control information into a gating matrix. This indicates the gating projection bias. This represents the semantic suppression vector. Represents a style enhancement vector. Indicates overlapping candidate positions. This indicates that the semantic suppression vector, style enhancement vector, and candidate position overlap markers are concatenated to form a gated input.
[0103] The gating style attention aggregation formula is:
[0104] in, This represents the result of gated style attention calculation, corresponding to the processing result represented by the aggregated style value based on the gated style attention distribution. This indicates the style query representation. This indicates the style key. This indicates the style value. This represents the gating matrix. This represents the initial style attention distribution, corresponding to the initial style attention distribution generated based on the style query representation and the style key representation. This indicates element-wise multiplication, used to adjust the initial style attention distribution using a gating matrix. This indicates a normalization process used to restore the adjusted attention results to a distribution form that can be used for weighted aggregation. This indicates the channel dimension represented by the style key.
[0105] This embodiment constrains the semantic attention distribution through a semantic location masking matrix and candidate semantic weight vectors, ensuring that the semantic embedding vectors are concentrated from the semantic value representations corresponding to the set of semantic candidate locations. A gating matrix is generated using semantic suppression vectors, style enhancement vectors, and candidate location overlap markers, allowing the initial style attention distribution to be differentiated based on semantic location, style location, and overlap location. By aggregating style value representations through the gating style attention distribution, the style embedding vectors can centrally carry the visual expression information from the style candidate locations. Thus, the semantic embedding vectors and style embedding vectors achieve a clearer source distinction during the encoding stage, reducing mutual interference between semantic content and style expression during vector aggregation.
[0106] In one embodiment, step S40 above includes: S401, determine the cross-attention layer in the target image generation model, and read the text input structure of the cross-attention layer; S402, determine the compatibility space based on the text input structure, the compatibility space including the representation length and channel arrangement; S403, Input the semantic embedding vector into the semantic residual projection path to generate a semantic projection representation; S404, calibrate the representation length and channel arrangement of the semantic projection representation according to the compatibility space, and generate a compatible semantic representation; S405, Generate a semantic conditional gating vector based on the compatible semantic representation; S406, The style is embedded into the vector input style projection path to generate a style projection representation; S407, calibrate the representation length and channel arrangement of the style projection representation according to the compatibility space, and adjust the calibrated style projection representation using the semantic conditional gating vector to generate a compatible style representation.
[0107] In this embodiment, the cross-attention layer in the target image generation model can be located through model structure configuration, layer name, conditional input port, and attention parameter table. The text input structure of the cross-attention layer typically includes input vector length, channel arrangement, batch dimension, conditional position arrangement, and masking format. When reading the text input structure, the size requirements that the text-side conditional vectors need to meet can be determined from the query input, key input, and value input configuration of the cross-attention layer, or the interface specifications between the conditional encoder output and the cross-attention layer input can be read through the inference interface. The compatibility space of the target image generation model is obtained by transforming the text input structure. The representation length in the compatibility space is used to constrain the length of the vector sequence, and the channel arrangement is used to constrain the vector dimension and channel order at each position, so that the semantic embedding vector and style embedding vector can be converted into conditional inputs that the cross-attention layer can read.
[0108] After the semantic embedding vector is input into the semantic residual projection path, it can be processed through linear projection, nonlinear mapping, normalization, and residual fusion to generate a semantic projection representation. Linear projection is used to change the number of channels in the semantic embedding vector; nonlinear mapping is used to enhance the expression differences of subject objects, attribute relationships, and scene relationships; normalization is used to adjust the vector distribution range; and residual fusion is used to preserve the original semantic structure in the semantic embedding vector. After the semantic projection representation is generated, the representation length and channel arrangement are calibrated according to the compatibility space. Representation length calibration can be achieved through truncation, padding, resampling, or pooling sampling; channel arrangement calibration can be achieved through channel rearrangement, channel projection, or channel normalization. The calibrated result forms a compatible semantic representation, which is consistent with the text input structure of the cross-attention layer.
[0109] The compatible semantic representation can be further used to generate semantic conditional gating vectors. These vectors can be generated through a gated projection layer, a channel compression layer, or a position-aware projection layer. The gated projection layer converts the compatible semantic representation into channel-level control variables, the channel compression layer aggregates multiple semantic locations into a global control variable, and the position-aware projection layer preserves the local influence of semantic locations on style locations. The semantic conditional gating vector can contain channel weights, position weights, or a joint weight of both, used to adjust the style projection representation so that the style content is constrained by semantic content when transforming to the compatible space.
[0110] After the style embedding vector is input into the style projection path, it can be processed through a style projection layer, a style normalization layer, and a channel transformation layer to generate a style projection representation. The style projection layer transforms the style embedding vector into a vector form close to the compatibility space; the style normalization layer adjusts the numerical distribution of the style vector; and the channel transformation layer matches the style channels with the channel arrangement in the compatibility space. The style projection representation is then calibrated in terms of representation length and channel arrangement according to the compatibility space. Length calibration ensures that the style projection representation and the compatible semantic representation have a composable sequence size, and channel calibration makes the style projection representation consistent with the text input channels of the cross-attention layer.
[0111] The calibrated style projection representation is adjusted by a semantic conditional gating vector. Adjustment can be achieved through channel-wise multiplication, position-wise weighting, gated residual superposition, or pre-normalization bias. Channel-wise multiplication is suitable for controlling the contribution of style channels, position-wise weighting is suitable for controlling the participation degree of different style segments, gated residual superposition is suitable for preserving the style projection representation while adding semantic conditional constraints, and pre-normalization bias is suitable for changing the direction of style expression before adjusting the numerical distribution. After adjustment by the semantic conditional gating vector, the style projection representation generates a compatible style representation. The compatible style representation and the compatible semantic representation reside in the same compatibility space, which can preserve style expression and reduce the deviation of style mapping from semantic content while maintaining the consistency of the input format of the target image generation model.
[0112] For example, the formula for compatible semantic conditional style mapping is:
[0113] in, This indicates a compatible style. Represents the style embedding vector. This indicates a compatible semantic representation. This represents the style projection path, used to convert style embedding vectors into style projection representations. This represents the gated projection path, used to generate semantically conditional gated vectors based on compatible semantic representations. This represents a semantic conditional gating vector. This indicates calibration processing performed based on the compatibility space, used to calibrate the representation length and channel arrangement. Indicates compatibility space. This indicates element-wise multiplication, used to adjust the calibrated style projection representation using semantically conditional gating vectors.
[0114] This embodiment reads the text input structure of the cross-attention layer and determines the compatibility space. Semantic embedding vectors and style embedding vectors can be converted into a conditional input format readable by the target image generation model. Through semantic residual projection paths and compatibility space calibration, the compatible semantic representation preserves the subject objects, attribute relationships, and scene relationships in the semantic embedding vectors. Semantic conditional gating vectors are generated from the compatible semantic representation, and these vectors are used to adjust the calibrated style projection representation. The compatible style representation is thus constrained by semantic content within the same compatibility space. Therefore, the semantic and style representations maintain consistency in input format and coordination in content relationships, reducing semantic shifts caused by excessive style mapping during subsequent image generation.
[0115] In one embodiment, step S50 above includes: S501, Obtain style intensity parameters, and generate style intensity control information based on the style intensity parameters; S502, determine the style components based on the style intensity control information and the compatible style representation, and generate the style scaling control vector corresponding to the style components; S503, adjust the style components according to the style scaling control vector to generate a scaled compatible style representation; S504, extract the semantic component positions in the compatible semantic representation and the style component positions in the scaled compatible style representation, and establish a combination correspondence based on the semantic component positions and the style component positions; S505, the compatible semantic representation and the scaled compatible style representation are combined according to the combination correspondence to form each combination component and generate candidate combination representations; S506, Based on the semantic component position and the style component position, configure component identifiers for each combination component in the candidate combination representation, and generate a combination text representation.
[0116] In this embodiment, the style intensity parameter can be provided by the interactive interface, configuration items, or text parsing results, and is used to control the degree of participation of the compatible style representation in the combined text representation. After processing, the style intensity parameter can be converted into style intensity control information. The style intensity control information can adopt a single scaling value, a channel scaling vector, a position scaling vector, or a control matrix that combines channels and positions. A single scaling value is suitable for increasing or decreasing the overall style expression, a channel scaling vector is suitable for situations where different style channels have different contributions, and a position scaling vector is suitable for situations where different style segments require different adjustment intensities. The style intensity control information needs to be consistent with the length and channel arrangement of the compatible style representation to avoid changing the input format of the compatible style representation during scaling.
[0117] The compatible style representation can be divided into style components based on channel response, position response, and fragment boundaries. Style components can represent color components, material components, lighting components, composition components, and artistic expression components, or style fragments corresponding to different position ranges within a style embedding sequence. When determining style components, the channel distribution, position distribution, and fragment markers in the compatible style representation can be read, giving each style component a clear scope of application. The style scaling control vector is generated based on style intensity control information and style components; each control value in the vector is associated with a corresponding style component. The style scaling control vector can be obtained through linear mapping, normalized mapping, or gating mapping, and is used to control the amplification or reduction of different style components.
[0118] When adjusting style components using the style scaling control vector, component-wise multiplication, weighted superposition, gated scaling, or pre-normalization adjustment can be employed. Component-wise multiplication directly changes the amplitude of style components; weighted superposition preserves the basic contribution of style components while adding an intensity bias; gated scaling limits the impact of excessively strong style components on the combined text representation; and pre-normalization adjustment alters the intensity of style expression before the numerical distribution becomes uniform. The adjusted style components are then recombined to form a scaled compatible style representation. This scaled compatible style representation retains its original length and channel arrangement, only changing the degree of style contribution, facilitating subsequent combination with compatible semantic representations.
[0119] The semantic component positions are extracted from the compatible semantic representation and can represent the location range of the main object, attribute relationships, action relationships, and scene relationships within the compatible semantic representation. The style component positions are extracted from the scaled compatible style representation and can represent the location range of color, material, lighting, composition, and artistic expression. The semantic and style component positions can be recorded using position indices, fragment boundaries, channel labels, or embedding sequence ranges. The combination correspondence is established based on the semantic and style component positions and is used to determine the arrangement of the compatible semantic representation and the scaled compatible style representation during combination. The combination correspondence can employ segmented correspondence, segment-by-segment correspondence, or staggered correspondence. Segmented correspondence preserves the overall boundaries of the semantic and style components; segment-by-segment correspondence arranges each semantic component close to its related style component; and staggered correspondence is suitable for situations where multiple semantic fragments and multiple style fragments need to be locally associated.
[0120] The compatible semantic representation and the scaled compatible style representation are combined according to a combination correspondence to form various combined components and generate candidate combined representations. Each combined component can include semantic combined components derived from the compatible semantic representation and style combined components derived from the scaled compatible style representation. The candidate combined representation retains the sequence arrangement, channel arrangement, and fragment boundaries of the combined components. Component identifiers are configured for each combined component in the candidate combined representation based on the positions of the semantic and style components. Component identifiers can be fragment labels, position labels, channel labels, or masking labels to distinguish between semantic and style combined components. After configuring component identifiers, the combined text representation simultaneously contains the combined vector content and the source information of each combined component, enabling subsequent reading to distinguish between the compatible semantic representation and the scaled compatible style representation.
[0121] For example, the formula for semantic style combination with component identifiers is:
[0122] in, This indicates a combined text representation. This indicates a compatible semantic representation. This indicates a compatible style. This represents the style intensity parameter. This indicates a scaled-down, compatible style representation. The corresponding component identifier indicates compatibility semantics. The component identifier represents the corresponding component in the scaled-down compatibility style. This indicates a combination process used to combine compatible semantic representations and scaled compatible style representations into a combined text representation with component identifiers.
[0123] This embodiment converts style intensity parameters into control variables that can be applied to compatible style representations through style intensity control information. Different style components can have their intensity adjusted according to their corresponding style scaling control vectors. By establishing a combination correspondence between semantic component positions and style component positions, the compatible semantic representation and the scaled compatible style representation can maintain their respective positional boundaries and fragment relationships when combined. By configuring component identifiers for each combination component in the candidate combined representation, the combined text representation retains source differentiation after merging semantic and style content. Thus, the style expression intensity can be adjusted, and semantic and style components remain distinguishable in the combined text representation, reducing interference with semantic content during style enhancement or reduction.
[0124] In one embodiment, step S60 above includes: S601, input the combined text representation into the target image generation model, and determine the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation from the combined text representation based on the component identifier; S602, control the cross-attention layer to read the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation according to the component identifier, and generate semantic condition features and style condition features; S603, based on the semantic condition features and the style condition features, drive the target image generation model to generate candidate images, and extract image semantic representation and image style representation from the candidate images; S604, Generate a text semantic reference representation and a text style reference representation based on the text prompt information, compare the image semantic representation with the text semantic reference representation to obtain semantic consistency, and compare the image style representation with the text style reference representation to obtain style consistency; S605, Generate image analysis results based on the semantic consistency and style consistency, and compare the image analysis results with preset output conditions; S606, when the image analysis result meets the preset output condition, the candidate image is determined as the target image and output.
[0125] In this embodiment, after the combined text representation is input into the target image generation model, it needs to be split into components based on component identifiers. This allows the target image generation model to distinguish between information segments carrying semantic constraints and those carrying style constraints. Component identifiers can be attached to embedding locations, channel locations, or segment boundaries. Therefore, before reading, it is necessary to scan each combined component in the combined text representation, identify the category to which the component identifier belongs, and assign combined components with semantic category labels to the corresponding part of the compatible semantic representation, while assigning combined components with style category labels to the corresponding part of the scaled compatible style representation. To avoid positional confusion during the reading stage, semantic index tables and style index tables can be formed after classification, allowing subsequent cross-attention layers to extract input content according to the index range. The corresponding part of the compatible semantic representation retains constraints related to objects, attributes, actions, scenes, and relationships, while the corresponding part of the scaled compatible style representation retains constraints related to color, material, lighting, composition, and artistic expression. This ensures that the text conditions input to the target image generation model remain separated in terms of content and expression dimensions.
[0126] When reading the corresponding parts of the compatible semantic representation and the scaled corresponding parts of the compatible style representation, the cross-attention layer can set independent attention projection channels or independent reading masking rules. The corresponding parts of the compatible semantic representation can participate in the generation of semantic conditional features, used to constrain the subject composition, spatial relationships, and semantic integrity in the candidate image. The corresponding parts of the scaled compatible style representation can participate in the generation of style conditional features, used to constrain the texture tendency, tone distribution, lighting and shadow performance, and artistic expression in the candidate image. During the reading process, the corresponding parts of the compatible semantic representation can be mapped to a set of semantic key values, and the corresponding parts of the scaled compatible style representation can be mapped to a set of style key values. Then, the cross-attention layer calculates the attention response for each set separately. After the semantic conditional features and style conditional features are generated, they can be injected into different levels of the target image generation model. Semantic conditional features are suitable for injection into locations that affect image structure and object layout, while style conditional features are suitable for injection into locations that affect surface appearance and visual quality. This dual-conditional constraint can reduce the dilution of semantic content during style enhancement and also reduce the compression of style content during semantic enhancement.
[0127] After receiving semantic and style conditional features, the target image generation model uses these two types of conditions to drive the image generation process and output candidate images. Candidate images can be the result of a single generation, or the result of multiple denoising iterations. For subsequent analysis of the candidate images, it is necessary to extract image semantic representations and image style representations. Image semantic representations can be obtained through visual coding networks, region semantic extraction structures, or object relationship extraction structures, and are used to represent the main objects, attribute states, action states, scene layout, and element relationships in the candidate images. Image style representations can be obtained through style analysis networks, texture decomposition structures, tone distribution extraction structures, or composition analysis structures, and are used to represent the color distribution, brushstroke tendencies, material texture, lighting style, and visual composition in the candidate images. Image semantic representations and image style representations respectively serve as the content and representational basis for subsequent comparisons; therefore, it is necessary to maintain the stability of the representation dimensions during the extraction process so that they can be uniformly compared with the reference representations generated from the text side.
[0128] The text prompts also need to be converted into semantic reference representations and style reference representations. Semantic reference representations can be obtained by aggregating semantic parsing results, and their content can cover the subject object, object attributes, action direction, scene information, and relationship descriptions. Style reference representations can be obtained by aggregating style parsing results, and their content can cover style categories, color tendencies, composition descriptions, material modifications, and artistic expression requirements. When comparing image semantic representations with text semantic reference representations, semantic matching scores, object coverage scores, relationship consistency scores, and scene consistency scores can be calculated, and these scores are summarized as semantic consistency. When comparing image style representations with text style reference representations, style matching scores, tone matching scores, composition matching scores, and expression matching scores can be calculated, and these scores are summarized as style consistency. Semantic consistency reflects the degree to which candidate images meet text content constraints, while style consistency reflects the degree to which candidate images meet text expression constraints; both together constitute the basic input for image analysis results.
[0129] Image analysis results can be represented using scalar scores, dual-score combinations, grade labels, or multi-field evaluation records. After semantic consistency and style consistency are input into the analysis unit, image analysis results containing content matching and style matching information can be generated. Preset output conditions can be set as a combination of semantic consistency threshold, style consistency threshold, and comprehensive judgment threshold, or as a judgment rule prioritizing semantic consistency or satisfying both semantic and style consistency. When comparing the image analysis results with the preset output conditions, if both semantic and style consistency requirements are met, the candidate image is determined as the target image and output; if the requirements are not met, the candidate image can be retained for subsequent filtering or regeneration. When outputting the target image, the image analysis results can be output together, facilitating the recording of the semantic and style satisfaction of the candidate image. Through the above processing, the compatible semantic representation and scaled compatible style representation in the combined text representation can be distinguished and utilized within the target image generation model. Candidate images can also complete result filtering through semantic and style consistency, maintaining a more stable constraint relationship between the output results and the text prompt information.
[0130] This embodiment decomposes the combined text representation into a compatible semantic representation and a scaled compatible style representation using component identifiers. The cross-attention layer can read the two types of input content separately. Semantic and style conditional features maintain independent constraints within the target image generation model. By extracting image semantic and style representations and comparing them with text semantic and style reference representations respectively, semantic and style consistency can be achieved. Based on this, image analysis results are generated and compared with preset output conditions. Thus, the candidate image generation and result selection processes can be simultaneously constrained by content and style, making it easier to maintain the integrity of text content and consistency of style expression when outputting the target image.
[0131] In one embodiment, a semantic style separation image generation apparatus is provided, which corresponds one-to-one with the semantic style separation image generation method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the semantic style separation image generation device of the present invention. The modules include a text embedding preprocessing module 10, a semantic style marker recognition module 20, a dual-branch gating coding module 30, a generation model compatible mapping module 40, a style intensity combination module 50, and an image generation analysis output module 60. Detailed descriptions of each functional module are as follows: The text embedding preprocessing module 10 is used to acquire text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; The semantic style tagging and recognition module 20 is used to identify semantic text units and style text units in the text unit sequence, and generate a semantic candidate position set and a style candidate position set according to the positions of the semantic text units and style text units in the standardized text embedding sequence. The dual-branch gating coding module 30 is used to aggregate semantic attention based on the standardized text embedding sequence and the semantic candidate position set, generate a semantic embedding vector, generate a gating matrix based on the standardized text embedding sequence, the semantic candidate position set and the style candidate position set, and use the gating matrix to perform gating style coding to generate a style embedding vector. The model compatibility mapping module 40 is used to determine the target image generation model and compatibility space containing the cross attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; The style intensity combination module 50 is used to obtain style intensity parameters, scale the compatible style representation according to the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation. The image generation, analysis, and output module 60 is used to input the combined text representation into the target image generation model, and the cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, analyzes the semantic and style consistency between the candidate images and the text prompt information, generates image analysis results, and outputs the target image when the image analysis results meet the preset output conditions.
[0132] In one embodiment, the text embedding preprocessing module 10 is specifically used for: Extract the original character sequence from the text prompt information, and generate an original character position table based on the arrangement order of each character in the original character sequence; A canonical character sequence is generated based on the original character sequence, and the separator characters, modifier characters, and phrase connector characters used to identify semantic and style boundaries are retained in the canonical character sequence; The standardized character sequence is adjusted according to a preset length boundary to generate a fixed-length standardized character sequence; A text unit boundary table is generated based on the fixed-length standard character sequence, and the boundary positions of continuous phrases, modifying phrases and parallel phrases are marked in the text unit boundary table; Based on the text unit boundary table, the fixed-length standard character sequence is segmented to generate a text unit sequence. A source position marker and a boundary marker are configured for the text units in the text unit sequence. The source position marker is used to record the position range of the text unit originating from the text prompt information in the original character position table. The boundary marker is determined by the text unit boundary table. The text unit sequence is mapped to a text unit vector sequence according to a preset text unit vocabulary, a position representation is generated according to the source position marker, a fragment representation is generated according to the text unit boundary table, and a boundary representation is generated according to the boundary marker. By fusing the text unit vector sequence, the position representation, the segment representation, and the boundary representation, a standardized text embedding sequence is generated.
[0133] In one embodiment, the semantic style tagging recognition module 20 is specifically used for: Read the sequential position of each text unit and adjacent text units in the text unit sequence, and determine the boundary marker based on the boundary position of each text unit in the text unit sequence to generate a text unit context window set; Based on the set of text unit context windows, semantic role tags, style role tags, and role confidence tags are generated for each text unit; Based on the semantic role markers, object text units, attribute text units, action text units, scene text units, and relation text units are filtered from the text unit sequence, and the object text units, attribute text units, action text units, scene text units, and relation text units are merged into a semantic text unit candidate set; Based on the style role markers, style expression text units, composition text units, material text units, lighting text units, color text units, and artistic expression text units are selected from the text unit sequence, and the style expression text units, composition text units, material text units, lighting text units, color text units, and artistic expression text units are merged into a style text unit candidate set; For a text unit that simultaneously possesses the semantic role tag and the style role tag, a master attribution tag is generated based on the role confidence tag and the boundary tag. Based on the master attribution tag, the text unit that simultaneously possesses the semantic role tag and the style role tag is retained in the semantic text unit candidate set or the style text unit candidate set, and is removed from the other candidate set that has never been retained. The candidate set of semantic text units is determined as semantic text units, and the candidate set of style text units is determined as style text units; Establish the embedding position mapping relationship between the text unit sequence and the standardized text embedding sequence; The position of the semantic text unit in the standardized text embedding sequence is determined based on the embedding position mapping relationship, generating a set of semantic candidate positions. The position of the style text unit in the standardized text embedding sequence is determined based on the embedding position mapping relationship, generating a set of style candidate positions.
[0134] In one embodiment, the dual-branch gating coding module 30 is specifically used for: The standardized text embedding sequence is input into the semantic encoding branch to generate semantic query representation, semantic key representation, and semantic value representation; A semantic location masking matrix and a candidate semantic weight vector are generated based on the set of semantic candidate locations, and a semantic attention distribution is generated based on the semantic query representation, the semantic key representation, the semantic value representation, the semantic location masking matrix, and the candidate semantic weight vector; Based on the semantic attention distribution, the semantic value representations corresponding to the set of semantic candidate positions are aggregated to generate a semantic candidate aggregate representation, and the semantic candidate aggregate representation is converted into a semantic embedding vector; The standardized text embedding sequence is input into the style encoding branch to generate a style query representation, a style key representation, and a style value representation, and an initial style attention distribution is generated based on the style query representation and the style key representation; A semantic suppression vector is generated based on the standardized text embedding sequence and the semantic candidate position set; a style enhancement vector is generated based on the standardized text embedding sequence and the style candidate position set; and a candidate position overlap marker is generated based on the position overlap relationship between the semantic candidate position set and the style candidate position set. A gating matrix is generated based on the semantic suppression vector, the style enhancement vector, and the candidate position overlap markers; The initial style attention distribution is adjusted using the gating matrix to generate a gated style attention distribution, and the style value representation is aggregated based on the gated style attention distribution to generate a style embedding vector.
[0135] In one embodiment, the model-compatible mapping module 40 is specifically used for: Identify the cross-attention layer in the target image generation model and read the text input structure of the cross-attention layer; The compatibility space is determined based on the text input structure, and the compatibility space includes the representation length and channel arrangement. The semantic embedding vector is input into the semantic residual projection path to generate a semantic projection representation; Based on the compatibility space, the representation length and channel arrangement of the semantic projection representation are calibrated to generate a compatible semantic representation; Generate semantic conditional gating vectors based on the aforementioned compatible semantic representation; The style embedding vector is input into the style projection path to generate a style projection representation; The style projection representation is calibrated according to the compatibility space, and the calibrated style projection representation is adjusted using the semantic conditional gating vector to generate a compatible style representation.
[0136] In one embodiment, the style intensity combination module 50 is specifically used for: Obtain style intensity parameters and generate style intensity control information based on the style intensity parameters; Based on the style intensity control information and the compatible style representation, style components are determined, and style scaling control vectors corresponding to the style components are generated; The style components are adjusted according to the style scaling control vector to generate a scaled compatible style representation; Extract the semantic component positions in the compatible semantic representation and the style component positions in the scaled compatible style representation, and establish a combined correspondence based on the semantic component positions and the style component positions; Based on the aforementioned combination correspondence, the compatible semantic representation and the scaled compatible style representation are combined to form various combination components and generate candidate combination representations; Based on the semantic component position and the style component position, component identifiers are configured for each combination component in the candidate combination representation to generate a combined text representation.
[0137] In one embodiment, the image generation and analysis output module 60 is specifically used for: The combined text representation is input into the target image generation model, and the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation are determined from the combined text representation based on the component identifier; The cross-attention layer is controlled to read the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation according to the component identifier, and generate semantic condition features and style condition features; The target image generation model is driven to generate candidate images based on the semantic condition features and the style condition features, and image semantic representation and image style representation are extracted from the candidate images; Based on the text prompt information, a text semantic reference representation and a text style reference representation are generated. The image semantic representation is compared with the text semantic reference representation to obtain semantic consistency. The image style representation is compared with the text style reference representation to obtain style consistency. Image analysis results are generated based on the semantic consistency and style consistency, and the image analysis results are compared with preset output conditions; When the image analysis result meets the preset output conditions, the candidate image is determined as the target image and output.
[0138] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the server-side functions or steps of a semantic style separation image generation method.
[0139] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a semantic style separation image generation method.
[0140] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; Identify semantic text units and style text units in the text unit sequence, and generate a set of semantic candidate positions and a set of style candidate positions according to the positions of the semantic text units and style text units in the standardized text embedding sequence; Semantic attention is aggregated based on the standardized text embedding sequence and the set of semantic candidate positions to generate a semantic embedding vector. A gating matrix is generated based on the standardized text embedding sequence, the set of semantic candidate positions, and the set of style candidate positions. The gating matrix is then used to perform gating style encoding to generate a style embedding vector. Determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; Obtain style intensity parameters, scale the compatible style representation based on the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The combined text representation is input into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet the preset output conditions, the target image is output.
[0141] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; Identify semantic text units and style text units in the text unit sequence, and generate a set of semantic candidate positions and a set of style candidate positions according to the positions of the semantic text units and style text units in the standardized text embedding sequence; Semantic attention is aggregated based on the standardized text embedding sequence and the set of semantic candidate positions to generate a semantic embedding vector. A gating matrix is generated based on the standardized text embedding sequence, the set of semantic candidate positions, and the set of style candidate positions. The gating matrix is then used to perform gating style encoding to generate a style embedding vector. Determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; Obtain style intensity parameters, scale the compatible style representation based on the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The combined text representation is input into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet the preset output conditions, the target image is output.
[0142] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0143] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0144] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0145] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0146] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An image generation method for semantic style separation, characterized in that, Includes the following steps: Obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; Identify semantic text units and style text units in the text unit sequence, and generate a set of semantic candidate positions and a set of style candidate positions according to the positions of the semantic text units and style text units in the standardized text embedding sequence; Semantic attention is aggregated based on the standardized text embedding sequence and the set of semantic candidate positions to generate a semantic embedding vector. A gating matrix is generated based on the standardized text embedding sequence, the set of semantic candidate positions, and the set of style candidate positions. The gating matrix is then used to perform gating style encoding to generate a style embedding vector. Determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; Obtain style intensity parameters, scale the compatible style representation based on the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The combined text representation is input into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet the preset output conditions, the target image is output.
2. The image generation method for semantic style separation as described in claim 1, characterized in that, Obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a normalized text embedding sequence, including: Extract the original character sequence from the text prompt information, and generate an original character position table based on the arrangement order of each character in the original character sequence; A canonical character sequence is generated based on the original character sequence, and the separator characters, modifier characters, and phrase connector characters used to identify semantic and style boundaries are retained in the canonical character sequence; The standardized character sequence is adjusted according to a preset length boundary to generate a fixed-length standardized character sequence; A text unit boundary table is generated based on the fixed-length standard character sequence, and the boundary positions of continuous phrases, modifying phrases and parallel phrases are marked in the text unit boundary table; Based on the text unit boundary table, the fixed-length standard character sequence is segmented to generate a text unit sequence. A source position marker and a boundary marker are configured for the text units in the text unit sequence. The source position marker is used to record the position range of the text unit originating from the text prompt information in the original character position table. The boundary marker is determined by the text unit boundary table. The text unit sequence is mapped to a text unit vector sequence according to a preset text unit vocabulary, a position representation is generated according to the source position marker, a fragment representation is generated according to the text unit boundary table, and a boundary representation is generated according to the boundary marker. By fusing the text unit vector sequence, the position representation, the segment representation, and the boundary representation, a standardized text embedding sequence is generated.
3. The image generation method for semantic style separation as described in claim 1, characterized in that, Identify semantic text units and style text units in the text unit sequence, and generate a semantic candidate position set and a style candidate position set according to the positions of the semantic text units and style text units in the normalized text embedding sequence, including: Read the sequential position of each text unit and adjacent text units in the text unit sequence, and determine the boundary marker based on the boundary position of each text unit in the text unit sequence to generate a text unit context window set; Based on the set of text unit context windows, semantic role tags, style role tags, and role confidence tags are generated for each text unit; Based on the semantic role markers, object text units, attribute text units, action text units, scene text units, and relation text units are filtered from the text unit sequence, and the object text units, attribute text units, action text units, scene text units, and relation text units are merged into a semantic text unit candidate set; Based on the style role markers, style expression text units, composition text units, material text units, lighting text units, color text units, and artistic expression text units are selected from the text unit sequence, and the style expression text units, composition text units, material text units, lighting text units, color text units, and artistic expression text units are merged into a style text unit candidate set; For a text unit that simultaneously possesses the semantic role tag and the style role tag, a master attribution tag is generated based on the role confidence tag and the boundary tag. Based on the master attribution tag, the text unit that simultaneously possesses the semantic role tag and the style role tag is retained in the semantic text unit candidate set or the style text unit candidate set, and is removed from the other candidate set that has never been retained. The candidate set of semantic text units is determined as semantic text units, and the candidate set of style text units is determined as style text units; Establish the embedding position mapping relationship between the text unit sequence and the standardized text embedding sequence; The position of the semantic text unit in the standardized text embedding sequence is determined based on the embedding position mapping relationship, generating a set of semantic candidate positions. The position of the style text unit in the standardized text embedding sequence is determined based on the embedding position mapping relationship, generating a set of style candidate positions.
4. The image generation method for semantic style separation as described in claim 1, characterized in that, Semantic attention is aggregated based on the standardized text embedding sequence and the semantic candidate position set to generate a semantic embedding vector. A gating matrix is then generated based on the standardized text embedding sequence, the semantic candidate position set, and the style candidate position set. Gated style encoding is performed using the gating matrix to generate a style embedding vector, including: The standardized text embedding sequence is input into the semantic encoding branch to generate semantic query representation, semantic key representation, and semantic value representation; A semantic location masking matrix and a candidate semantic weight vector are generated based on the set of semantic candidate locations, and a semantic attention distribution is generated based on the semantic query representation, the semantic key representation, the semantic value representation, the semantic location masking matrix, and the candidate semantic weight vector; Based on the semantic attention distribution, the semantic value representations corresponding to the set of semantic candidate positions are aggregated to generate a semantic candidate aggregate representation, and the semantic candidate aggregate representation is converted into a semantic embedding vector; The standardized text embedding sequence is input into the style encoding branch to generate a style query representation, a style key representation, and a style value representation, and an initial style attention distribution is generated based on the style query representation and the style key representation; A semantic suppression vector is generated based on the standardized text embedding sequence and the semantic candidate position set; a style enhancement vector is generated based on the standardized text embedding sequence and the style candidate position set; and a candidate position overlap marker is generated based on the position overlap relationship between the semantic candidate position set and the style candidate position set. A gating matrix is generated based on the semantic suppression vector, the style enhancement vector, and the candidate position overlap markers; The initial style attention distribution is adjusted using the gating matrix to generate a gated style attention distribution, and the style value representation is aggregated based on the gated style attention distribution to generate a style embedding vector.
5. The image generation method for semantic style separation as described in claim 1, characterized in that, Determine the target image generation model and compatibility space containing the cross-attention layer, map the semantic embedding vectors to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vectors to the compatibility space using the compatible semantic representation to generate a compatible style representation, including: Identify the cross-attention layer in the target image generation model and read the text input structure of the cross-attention layer; The compatibility space is determined based on the text input structure, and the compatibility space includes the representation length and channel arrangement. The semantic embedding vector is input into the semantic residual projection path to generate a semantic projection representation; Based on the compatibility space, the representation length and channel arrangement of the semantic projection representation are calibrated to generate a compatible semantic representation; Generate semantic conditional gating vectors based on the aforementioned compatible semantic representation; The style embedding vector is input into the style projection path to generate a style projection representation; The style projection representation is calibrated according to the compatibility space, and the calibrated style projection representation is adjusted using the semantic conditional gating vector to generate a compatible style representation.
6. The image generation method for semantic style separation as described in claim 1, characterized in that, Obtain style intensity parameters, scale the compatible style representation based on the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation, configure component identifiers, and generate a combined text representation, including: Obtain style intensity parameters and generate style intensity control information based on the style intensity parameters; Based on the style intensity control information and the compatible style representation, style components are determined, and style scaling control vectors corresponding to the style components are generated; The style components are adjusted according to the style scaling control vector to generate a scaled compatible style representation; Extract the semantic component positions in the compatible semantic representation and the style component positions in the scaled compatible style representation, and establish a combined correspondence based on the semantic component positions and the style component positions; Based on the aforementioned combination correspondence, the compatible semantic representation and the scaled compatible style representation are combined to form various combination components and generate candidate combination representations; Based on the semantic component position and the style component position, component identifiers are configured for each combination component in the candidate combination representation to generate a combined text representation.
7. The image generation method for semantic style separation as described in claim 1, characterized in that, The combined text representation is input into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation from the combined text representation according to the component identifiers, generates candidate images, and analyzes the semantic and style consistency between the candidate images and the text prompt information to generate image analysis results. When the image analysis results meet preset output conditions, the target image is output, including: The combined text representation is input into the target image generation model, and the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation are determined from the combined text representation based on the component identifier; The cross-attention layer is controlled to read the corresponding part of the compatible semantic representation and the corresponding part of the scaled compatible style representation according to the component identifier, and generate semantic condition features and style condition features; The target image generation model is driven to generate candidate images based on the semantic condition features and the style condition features, and image semantic representation and image style representation are extracted from the candidate images; Based on the text prompt information, a text semantic reference representation and a text style reference representation are generated. The image semantic representation is compared with the text semantic reference representation to obtain semantic consistency. The image style representation is compared with the text style reference representation to obtain style consistency. Image analysis results are generated based on the semantic consistency and style consistency, and the image analysis results are compared with preset output conditions; When the image analysis result meets the preset output conditions, the candidate image is determined as the target image and output.
8. An image generation apparatus for semantic style separation, characterized in that, The semantic style separation image generation device includes: The text embedding preprocessing module is used to obtain text prompt information, segment the text prompt information into words to generate a text unit sequence, and map the text unit sequence into a standardized text embedding sequence; A semantic style tagging and recognition module is used to identify semantic text units and style text units in the text unit sequence, and generate a semantic candidate position set and a style candidate position set according to the positions of the semantic text units and style text units in the standardized text embedding sequence; The dual-branch gated coding module is used to aggregate semantic attention based on the standardized text embedding sequence and the semantic candidate position set to generate a semantic embedding vector, and to generate a gated matrix based on the standardized text embedding sequence, the semantic candidate position set and the style candidate position set, and to perform gated style coding using the gated matrix to generate a style embedding vector. The model compatibility mapping module is used to determine the target image generation model and compatibility space containing the cross attention layer, map the semantic embedding vector to the compatibility space to generate a compatible semantic representation, and adjust the mapping of the style embedding vector to the compatibility space with the compatible semantic representation to generate a compatible style representation; The style intensity combination module is used to obtain style intensity parameters, scale the compatible style representation according to the style intensity parameters, combine the compatible semantic representation and the scaled compatible style representation and configure component identifiers to generate a combined text representation; The image generation, analysis, and output module is used to input the combined text representation into the target image generation model. The cross-attention layer reads the compatible semantic representation and the scaled compatible style representation in the combined text representation according to the component identifier, generates candidate images, analyzes the semantic and style consistency between the candidate images and the text prompt information, generates image analysis results, and outputs the target image when the image analysis results meet the preset output conditions.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a semantic style separation image generation program stored in the memory and executable on the processor. When executed by the processor, the semantic style separation image generation program implements the steps of the semantic style separation image generation method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a semantic style separation image generation program, which, when executed by a processor, implements the steps of the semantic style separation image generation method as described in any one of claims 1-7.