Image generation method and device based on low-rank parameter adaptation, equipment and medium
Patent Information
- Application Number
- CN202611147434.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-30
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于低秩参数适配的图像生成方法、装置、设备及存储介质,旨在解决现有文本到图像生成技术在精确布局控制场景下,难以稳定生成同时满足文本语义一致性和预设空间布局关系的图像,并且为适配布局约束通常需要对大型预训练模型进行整体微调,造成训练成本高、参数更新范围大且生成稳定性不足的技术问题
[0010]有益效果:本发明涉及模型构建技术领域,公开了一种基于低秩参数适配的图像生成方法、装置、设备及介质,包括:在包含文本编码器和图像解码器的生成框架中嵌入低秩调节矩阵,将训练文本描述转换为文本语义向量,将训练布局标注转换为布局约束表示;结合当前扩散去噪状态生成中间图像表征,并提取空间注意力图;基于空间注意力图和训练布局标注确定空间相干性损失,基于噪声预测结果确定重建损失,基于重建损失对应的参数梯度确定随机梯度正则化项;基于空间相干性损失、重建损失和随机梯度正则化项更新低秩调节矩阵,在保持预训练权重不变的条件下得到更新后的低秩调节矩阵,并基于目标文本语义向量和目标布局约束表示生成目标图像。本发明可应用于金融科技以及医疗健康等业务场景中,通过在生成框架中引入低秩调节矩阵,并将空间相干性损失、重建损失和随机梯度正则化项共同用于参数优化,使布局约束与语义约束在同一训练路径中协同作用,在保持预训练权重不变的情况下提升图像生成过程中的布局控制能力和生成稳定性。
Smart Images

Figure CN122820871A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of model building technology, and in particular to an image generation method, apparatus, device and medium based on low-rank parameter adaptation. Background Technology
[0002] While existing text-to-image generation technologies can generate images based on natural language descriptions, they still have significant shortcomings in scenarios that require simultaneous satisfaction of semantic content and spatial layout requirements. Models are inconsistent in their responses to positional relationships, regional relationships, and relative relationships between objects in the text, easily leading to issues such as object position shifts, region misalignments, and layout chaos. Furthermore, directly fine-tuning large pre-trained models for layout control often results in high training costs, excessively wide parameter adjustment ranges, and insufficient generation stability.
[0003] In the fintech sector, text-to-image generation is commonly used to create images for financial product displays, return analysis charts, risk warning pages, and information cards. These scenarios typically require titles, curves, amounts, labels, and prompts to be positioned according to preset criteria. While existing technologies can generate image content related to financial keywords, they lack sufficient control over layout requirements such as "curves at the top," "risk warnings at the bottom," and "information cards arranged in sections," easily leading to issues like curves overlapping with text, offset prompt areas, and disordered card order.
[0004] In the healthcare industry, text-to-image generation is commonly used to create diagrams displaying examination indicators, health education materials, treatment flowcharts, and drug information leaflets. These scenarios typically require clearly defined positional relationships between indicator areas, explanation areas, warning areas, and process nodes. Existing technologies, when processing these images, are prone to issues such as area overlap, object offset, and relative layout distortion, resulting in generated images that fail to meet the healthcare industry's requirements for layout accuracy and standardized information expression. Summary of the Invention
[0005] The main objective of this invention is to provide an image generation method, apparatus, device, and storage medium based on low-rank parameter adaptation. This invention aims to solve the technical problems of existing text-to-image generation technologies, which struggle to stably generate images that simultaneously satisfy text semantic consistency and preset spatial layout relationships in scenarios requiring precise layout control. Furthermore, adapting to layout constraints typically requires overall fine-tuning of large pre-trained models, resulting in high training costs, a large parameter update range, and insufficient generation stability.
[0006] To achieve the above objectives, the present invention provides an image generation method based on low-rank parameter adaptation, comprising: A generative framework is constructed, which includes a text encoder and an image decoder, and a low-rank adjustment matrix is embedded at the attention projection position of the generative framework; Obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations; Determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract a spatial attention map from the cross attention layer of the generation framework; The spatial coherence loss is determined based on the spatial attention map and the training layout annotation; the reconstruction loss is determined based on the noise prediction result corresponding to the intermediate image representation; and the stochastic gradient regularization term is determined based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The total loss is determined based on the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term. The low-rank adjustment matrix is updated based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, thus obtaining the updated low-rank adjustment matrix. Using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation.
[0007] Furthermore, to achieve the above objectives, the present invention provides an image generation apparatus based on low-rank parameter adaptation, comprising: A generative framework building module is used to construct a generative framework, which includes a text encoder and an image decoder, and embeds a low-rank adjustment matrix at the attention projection position of the generative framework; The training input processing module is used to obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations. The intermediate representation generation module is used to determine the current diffusion denoising state, generate intermediate image representations based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract spatial attention maps from the cross attention layer of the generation framework. The triple loss calculation module is used to determine the spatial coherence loss based on the spatial attention map and the training layout annotation, determine the reconstruction loss based on the noise prediction result corresponding to the intermediate image representation, and determine the stochastic gradient regularization term based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The low-rank adjustment and update module is used to determine the total loss based on the spatial coherence loss, the reconstruction loss and the stochastic gradient regularization term, and update the low-rank adjustment matrix based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, and obtain the updated low-rank adjustment matrix. The target image generation module is used to generate a target image based on the target text semantic vector and the target layout constraint representation by utilizing the updated low-rank adjustment matrix and the generation framework.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and an image generation program based on low-rank parameter adaptation stored in the memory and executable on the processor, wherein when the image generation program based on low-rank parameter adaptation is executed by the processor, it implements the steps of the image generation method based on low-rank parameter adaptation as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing an image generation program based on low-rank parameter adaptation, wherein the image generation program based on low-rank parameter adaptation, when executed by a processor, implements the steps of the image generation method based on low-rank parameter adaptation as described above.
[0010] Beneficial Effects: This invention relates to the field of model building technology and discloses an image generation method, apparatus, device, and medium based on low-rank parameter adaptation. The method includes: embedding a low-rank adjustment matrix in a generation framework comprising a text encoder and an image decoder; converting training text descriptions into text semantic vectors and training layout annotations into layout constraint representations; generating intermediate image representations by combining the current diffusion denoising state and extracting a spatial attention map; determining spatial coherence loss based on the spatial attention map and training layout annotations, determining reconstruction loss based on noise prediction results, and determining a stochastic gradient regularization term based on the parameter gradients corresponding to the reconstruction loss; updating the low-rank adjustment matrix based on the spatial coherence loss, reconstruction loss, and stochastic gradient regularization term, obtaining the updated low-rank adjustment matrix while maintaining the pre-trained weights unchanged, and generating a target image based on the target text semantic vector and the target layout constraint representation. This invention can be applied to business scenarios such as fintech and healthcare. By introducing a low-rank adjustment matrix into the generation framework and using spatial coherence loss, reconstruction loss and stochastic gradient regularization term together for parameter optimization, the layout constraints and semantic constraints work synergistically in the same training path, thereby improving the layout control capability and generation stability in the image generation process while keeping the pre-trained weights unchanged. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for an image generation method based on low-rank parameter adaptation according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the image generation method based on low-rank parameter adaptation according to the present invention. Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image generation device based on low-rank parameter adaptation of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The image generation method based on low-rank parameter adaptation provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can embed a low-rank adjustment matrix into a generative framework containing a text encoder and an image decoder, converting training text descriptions into text semantic vectors and training layout annotations into layout constraint representations. It then generates intermediate image representations based on the current diffusion denoising state and extracts a spatial attention map. Based on the spatial attention map and training layout annotations, it determines spatial coherence loss, based on noise prediction results, and determines a stochastic gradient regularization term based on the parameter gradients corresponding to the reconstruction loss. The low-rank adjustment matrix is updated based on the spatial coherence loss, reconstruction loss, and stochastic gradient regularization term, resulting in an updated low-rank adjustment matrix while maintaining the pre-trained weights. Finally, it generates the target image based on the target text semantic vector and the target layout constraint representation. This invention can be applied to business scenarios such as fintech and healthcare. By introducing a low-rank adjustment matrix into the generative framework and using spatial coherence loss, reconstruction loss, and stochastic gradient regularization term together for parameter optimization, it enables layout constraints and semantic constraints to work synergistically in the same training path, improving layout control and generation stability during image generation while maintaining the pre-trained weights. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the image generation method based on low-rank parameter adaptation provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the image generation method based on low-rank parameter adaptation proposed in this invention includes the following steps: S10, Construct a generative framework, which includes a text encoder and an image decoder, and embed a low-rank adjustment matrix at the attention projection position of the generative framework; In this embodiment, the generation framework consists of a semantic processing part and an image forming part. The text encoder is responsible for converting the input text into continuous features that can participate in image generation, and the image decoder is responsible for progressively unfolding the continuous features into spatial features and forming image content. This combination is not a simple splicing, but rather allows semantic features to continuously participate in the image forming process, enabling object information, positional relationships, and layout constraints in the text to act on the image space. In actual implementation, the text encoder can be set as a combination of a word embedding layer, a position embedding layer, and a context modeling layer, and the image decoder can be set as a combination of a multi-layer feature transformation unit, an upsampling unit, and a reconstruction unit. The feature dimension output by the text encoder needs to be consistent with the dimension of the conditional information received by the image decoder. When there is a difference in dimension between the two, alignment can be achieved through a linear mapping layer.
[0016] Attention projection locations are crucial entry points within the image decoder responsible for forming attention maps. These locations handle the mapping from input features to query features, key features, and value features, significantly impacting subsequent spatial distribution and object relationships. To concentrate the adjustment on areas most directly affecting layout and object placement, multiple attention projection locations need to be identified within the image decoder, with a unique index created for each. These indexes can be based on hierarchy, resolution, channel segment, or path numbering, ensuring accurate differentiation between locations during subsequent loading and updates.
[0017] The low-rank adjustment matrix acts as an adjustable mapping, employing a combination of low-dimensional mapping and reversion mapping to provide additional corrections to the original attention projection result. This structure controls the scale of the added parameters, concentrating the adjustment capability at the attention projection location rather than spreading it across the entire large set of parameters. In implementation, two sets of matrices can be concatenated: the first set compresses the input features into a lower-dimensional space, and the second set maps the adjustment information from the lower-dimensional space back to the original dimensional space. The resulting adjustment is then fused with the original attention projection result. This fusion can be achieved through additive fusion or gated fusion. Additive fusion preserves the main orientation of the original projection result, while gated fusion allows for hierarchical control of the adjustment amplitude at different locations.
[0018] The embedding relationship is manifested as parallel access rather than replacement. The original attention projection positions retain the original computational path, while the low-rank adjustment matrix participates in the result formation through an additional path, and the two converge at the output. This processing can preserve the existing expressive power of the image decoder while introducing directional correction capabilities for layout control. For layout-style images, chart-style images, card-style images, and explanatory images, the attention projection positions at medium and low resolution levels usually affect the overall partitioning and object arrangement, while the attention projection positions at high resolution levels usually affect the arrangement of local edges and details. Therefore, the loading position of the low-rank adjustment matrix can be set differently according to the level. In the fintech business field, the yield curve area, amount area, risk warning area, and button area, and in the healthcare business field, the indicator area, explanation area, reminder area, and process area are all suitable for forming a more stable spatial separation relationship through this differentiated loading method.
[0019] In one implementation, the text encoder employs a concatenated structure of a word embedding layer, a position embedding layer, and a multi-head context modeling layer, while the image decoder uses a four-level resolution decoding structure, with each level containing a feature transformation unit and an upsampling unit. Each level of the image decoder internally sets multiple attention projection positions, and low-rank adjustment matrices are loaded onto all these attention projection positions. Each low-rank adjustment matrix consists of a dimensionality reduction matrix and an dimensionality increase matrix. The dimensionality reduction matrix compresses the input features into a lower-rank space, while the dimensionality increase matrix maps the compressed adjustment features back to the original dimension. The mapping result is then directly added to the original attention projection result. This implementation is suitable for image generation tasks with complex layout structures, such as multi-section interface images containing title areas, curve areas, prompt areas, and interactive areas.
[0020] In another implementation, the image decoder loads a low-rank adjustment matrix only at the attention projection positions of low- and mid-resolution layers, while the high-resolution layers retain their original attention projection results. The low-rank adjustment matrix is combined with the original attention projection results using a gated fusion method, with the gate value determined by the feature statistics of the corresponding layer. This implementation is more suitable for image generation tasks that require stable large block layouts but need to maintain natural local textures, such as diagrams, flowcharts, or card-based infographics.
[0021] A shared loading approach can also be used. Multiple attention projection positions with similar semantic functions share the same set of low-rank adjustment matrices, and the output correction range is distinguished by position index. This reduces the parameter size while maintaining independent control over different regions. This implementation is suitable for image generation tasks with a lot of text and obvious repetitive layout structures, such as multi-card side-by-side display diagrams, indicator list diagrams, and partitioned explanation diagrams.
[0022] This embodiment concentrates adjustable parameters at the attention projection location, allowing the correction of image spatial distribution to more directly affect object positions, region boundaries, and layout partitioning relationships. A stable correspondence is more easily established between the semantic information provided by the text encoder and the spatial features formed by the image decoder. The parallel loading method preserves the main expressive power of the original projection results, and the low-rank structure compresses the scale of new parameters, resulting in a concentrated adjustment range and a smaller parameter burden. This is suitable for image generation tasks that require control over layout but do not want to significantly alter the existing parameter structure.
[0023] S20, obtain training text description and training layout annotation, convert the training text description into a text semantic vector, and convert the training layout annotation into a layout constraint representation; In this embodiment, the training text description carries the object content, object attributes, and relative relationships between objects in the image to be generated. It is not just a simple sentence input, but can also include location pointers, region allocation, and layout organization information. In implementation, the training text description can first be segmented into words or lexical units, and then the lexical units with independent semantic functions can be organized into a text sequence. Object name lexical units are used to define the entities that need to appear in the image, location lexical units are used to define the distribution area of entities, and relationship lexical units are used to define the relative positions between multiple entities. In fintech business scenarios, yield curves, risk warnings, product tags, subscription buttons, and amount cards can be used as object content, and their location (top, right, near the bottom) can be used as location information. In healthcare business scenarios, indicator panels, reminder icons, explanatory areas, and process nodes can be used as object content, and their centered arrangement, left-right partitioning, and top-bottom distribution can be used as location information. The text semantic vector is output by the text encoder. Its formation method is not simply to preserve the literal order, but rather to compress the object content, location information, and relationship information into continuous features, enabling subsequent image formation stages to directly read object categories, region tendencies, and relative layouts. In the implementation, a word embedding layer, a positional encoding layer, and a context modeling layer can be set in the text encoder. The word embedding layer converts words into initial vectors, the positional encoding layer preserves the preceding and following relationships, and the context modeling layer combines scattered words into feature representations with overall semantics. Finally, the text semantic vector is output.
[0024] Training layout annotations carry spatial distribution information of an image, including at least object categories, target regions, and spatial correspondences between objects. Object categories determine the type of content that needs to occupy space in the image; target regions determine the scope and boundaries of the content within the image; and spatial correspondences determine the vertical, horizontal, containment, parallel, or spaced relationships between multiple objects. In implementation, training layout annotations can be parsed into structured region data, and then layout constraint representations can be generated based on this data. Layout constraint representations are not limited to a single rectangular bounding box; they can also use region masks, partition matrices, or multi-channel position tensors, as long as they can express the correspondence between object categories and image regions in the spatial dimension. To enable layout constraint representations to directly participate in image generation, their spatial dimensions need to match the image's feature resolution. In implementation, the corresponding category can be written into the target region location based on the horizontal and vertical coverage of the target region, and then the region boundaries can be smoothed or interpolated to ensure that the layout constraint representation preserves boundary information while facilitating subsequent feature fusion. In the fintech business, the partitioning information between the revenue chart area, risk warning area, and transaction entry area can be written into different channels; in the healthcare business, the boundary information of the indicator area, explanation area, reminder area, and process area can also be written into the corresponding channels according to the area, thereby forming a layout constraint representation that can directly participate in image space control.
[0025] The training text description is converted into a text semantic vector, and the training layout annotation is converted into a layout constraint representation. These two parts are not independent. The text side provides the object category and relational meaning, while the layout side provides the object target area and spatial correspondence. Both need to maintain semantic consistency in implementation. When the object name and object category are inconsistent, the content emphasized by the text semantic vector will not match the area defined by the layout constraint representation. This can easily lead to situations where the content is correct but the position is incorrect, or the position is correct but the content is chaotic, during subsequent generation. To avoid this problem, an object name set can be established after parsing the training text description, and an object category set can be established after parsing the training layout annotation. Then, alignment can be completed based on the mapping relationship between the object name set and the object category set, ensuring that the object category reflected in the text semantic vector is consistent with the area occupied by the layout constraint representation.
[0026] In one implementation, the training text description is input into a text encoder in the form of a word sequence. The text encoder consists of a word embedding layer, a positional encoding layer, and a multi-layer context modeling unit, outputting a fixed-dimensional text semantic vector. The training layout annotation is input into the layout parsing unit in the form of region boxes plus category labels. The layout parsing unit writes different object target regions into different channels based on the category labels, and then generates a layout constraint representation consistent with the image feature resolution through size transformation. This implementation is suitable for image generation tasks with clearly defined object categories and region boundaries, such as revenue posters, product cards, and risk warning images in the fintech business field, and also suitable for indicator display images, health explanation images, and process reminder images in the healthcare business field.
[0027] In another implementation, after training the text description input text encoder, it simultaneously outputs global semantic features and local labeled features. The global semantic features are used to form the text semantic vector, while the local labeled features are used to preserve the fine-grained semantics of object relationships. The training layout annotation does not use region boxes but rather region masks. Each object category forms an independent region layer in the spatial dimension, and multiple region layers are then combined into a layout constraint representation. This implementation is more suitable for image generation tasks with irregular object boundaries and complex region coverage, such as multi-curve analysis charts, asset allocation charts, and bill interpretation charts in the fintech field, and also suitable for nutrition zoning charts, training action cue charts, and health assessment charts in the healthcare field.
[0028] Another approach is to enhance object relationships. The training text description adds a sequence of relationship tags before the input text encoder, with hierarchical, parallel, containment, and spacing relationships between objects written into these tags. During layout annotation parsing training, distance, overlap, and alignment values between regions are extracted simultaneously, and then combined with object categories to form a layout constraint representation.
[0029] In the fintech business field, training text descriptions can include product names, return ranges, risk levels, prompts, and transaction entry instructions. Training layout annotations can mark the title area, return chart area, risk warning area, and button area. The text semantic vector output by the text encoder preserves the semantic differences of return, risk, and purchase. The layout constraints indicate that the spatial distribution relationship of the four types of areas is preserved, making it easier to generate a screen structure with the title at the top, the chart in the center, the prompts at the bottom, and the entry point at the bottom.
[0030] In the healthcare field, training text descriptions can include indicator names, reference ranges, anomaly alerts, and intervention suggestions. Training layout annotations can mark indicator areas, explanation areas, alert areas, and suggestion areas. Text semantic vectors preserve the content differences between indicators and alerts, and layout constraints preserve the spatial boundaries of each area. This makes it easier to generate image content with concentrated indicator information, prominent alerts, and clearly defined explanation and suggestion sections.
[0031] This embodiment uses semantic encoding to create textual semantic vectors that express object categories, location information, and relative relationships. Layout annotations, after region parsing, form layout constraint representations that express the object target region and its spatial correspondence. The conditional information received by the image side is no longer limited to a single semantic content level, but simultaneously possesses semantic and spatial constraints. When the object category and the object target region remain consistent, image content and layout distribution are more easily synchronized, reducing issues such as region misalignment, object drift, and relationship confusion.
[0032] S30, determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract a spatial attention map from the cross attention layer of the generation framework; In this embodiment, the current diffusion denoising state carries the noise distribution information and feature evolution position of the image generation in the current diffusion round. It is neither simply random noise nor the final image feature, but an intermediate state used to constrain the denoising direction of the current round. In implementation, the time code corresponding to the current diffusion round can be combined with the noise latent representation under the current round to form the current diffusion denoising state. The time code is used to distinguish the evolution position of different denoising rounds, and the noise latent representation is used to preserve the image spatial distribution of the current round. If a discrete round control form is adopted, the current diffusion round can be mapped to a round embedding vector, and then concatenated or added with the noise latent representation; if a continuous time control form is adopted, the current diffusion round can be converted into continuous time features, and then fed into the modulation unit for fusion with the noise latent representation. In fintech businesses, revenue charts, risk warnings, and billing descriptions typically require the title area, chart area, and notification area to remain stably partitioned during denoising. Similarly, in healthcare businesses, indicator display images, reminder descriptions, and health intervention images typically require the indicator area, description area, and reminder area to maintain fixed boundaries during denoising. Therefore, the current diffusion denoising process needs to simultaneously preserve both the overall coarse distribution of the layout and the evolution information of local regions.
[0033] Text semantic vectors and layout constraint representations play different control roles in the generation framework. Text semantic vectors provide object categories, object attributes, and semantic relationships between objects, while layout constraint representations provide object target regions, region boundaries, and relative positional relationships between regions. In implementation, text semantic vectors can be fed into the semantic input of multiple cross-attention layers, and layout constraint representations can be mapped to spatial constraint tensors consistent with the image-side feature resolution, and then fed into the spatial modulation positions of multiple cross-attention layers. With this setup, semantic content and spatial constraints are not fused at a single location all at once, but rather repeatedly participate in image feature updates across multiple levels. For low-resolution levels, text semantic vectors and layout constraint representations are more focused on controlling the general distribution of objects and region division; for high-resolution levels, they are more focused on controlling object edges, text regions, icon regions, and detailed arrangement.
[0034] Intermediate image representations are image-side features gradually formed by the image decoder during the denoising process. They are not equivalent to the final image and are not limited to a single-layer output. In implementation, the image decoder in the generation framework receives the text semantic vector, layout constraint representation, and the current diffusion denoising state. It then sequentially performs feature transformation, attention fusion, and resolution enhancement at multiple decoding levels to obtain multi-layer image features, which are then fused into intermediate image representations. The fusion method can be either summation after progressive upsampling or concatenation followed by mapping. In the fintech business domain, the yield curve region, amount region, and risk warning region often need to maintain both stable block positions and identifiable content; the intermediate image representation needs to retain both region outlines and textual association information. In the healthcare business domain, indicator regions, reminder regions, and explanation regions also need to maintain positional and boundary relationships during image formation; the intermediate image representation needs to retain both object distribution and spatial boundaries.
[0035] The cross-attention layer acts as a mapping between semantic features and image features, while the spatial attention map reflects the distribution of this mapping in the image space. In implementation, attention weights related to text tags can be read from multiple cross-attention layers and rearranged into two-dimensional or three-dimensional attention matrices according to spatial dimensions. Then, the attention matrices of the corresponding text tags are aggregated based on object category to form a spatial attention map. If multiple tags related to the same object exist in the text semantic vector, the attention matrices corresponding to these tags can be summed, averaged, or gated and fused. If the cross-attention layers are distributed across different decoding levels, the spatial attention maps generated at different levels can be size-aligned and then fused between layers to form a spatial attention map that simultaneously reflects the overall layout and local boundaries. The resulting spatial attention map not only reflects the region of interest for the object in the image but also reflects the actual degree of correspondence between text semantics and spatial regions.
[0036] In one implementation, the current diffusion denoising state is composed of a temporal embedding vector and a noise latent representation. The temporal embedding vector is fed into a linear layer after sine and cosine mapping, and the noise latent representation is provided by the diffusion initialization unit. The two are concatenated along the channel dimension and then input into the image decoder. The text semantic vector is copied to multiple cross-attention layers, and the layout constraint representation is converted into a spatial constraint tensor consistent with the image feature resolution of each layer after convolution mapping. The image decoder generates image features at four resolution levels. Each level of cross-attention layer outputs a corresponding attention weight map. The attention weight maps related to the object category are then upsampled to a uniform size and summed to obtain the spatial attention map. This implementation is suitable for revenue curves, product display charts, and risk warning charts in fintech businesses, as well as indicator explanation charts, health reminder charts, and rehabilitation guidance charts in healthcare businesses.
[0037] In another implementation, the current diffusion denoising state consists of round-by-round encoding and multi-scale noise features. Round-by-round encoding operates on different decoding levels, while multi-scale noise features are input independently at each level. Layout constraints are not directly input as the overall image tensor but are split into multiple region constraint sub-tensors, which are injected into the cross-attention layers corresponding to the object regions. Text semantic vectors maintain their label-level representation within the cross-attention layers and are not pre-compressed into a single global vector. The spatial attention map is formed by aggregating the label-level attention matrices output from each cross-attention layer according to object categories, and then integrating the results from different levels through gating fusion. This implementation is more suitable for image generation scenarios with complex regional relationships and a large number of objects, such as multi-card infographics and asset allocation diagrams in fintech businesses, and also suitable for multi-indicator panel charts and process guidance diagrams in healthcare businesses.
[0038] Alternatively, a region-first implementation can be adopted. The layout constraint representation is first parsed into multiple object target region masks. Each object target region mask is applied to different levels of the image decoder, resulting in coarse localization of the object target region at low-resolution levels and then boundary refinement at high-resolution levels. A one-to-one mapping is established between the object markers in the text semantic vector and the object target region masks. The cross-attention layer only enhances the semantic response within the corresponding object region. The spatial attention map is obtained by statistically analyzing the local attention weights within the object region.
[0039] In the fintech business, this task can handle the generation of investment return display charts. The target text includes the product name, return curve description, risk warning, and purchase entry instructions. Layout constraints require the title to be at the top, the return chart in the center, and the risk warning at the bottom. The current diffusion denoising state controls the coarse distribution of the layout in the denoising rounds, the text semantic vector controls the semantic expression of the product name, return description, and warning information, the layout constraints control the positional boundaries of the title area, chart area, and warning area, and the spatial attention map output by the cross-attention layer reflects the correspondence between the return curve text and the chart area, and between the risk warning text and the bottom area.
[0040] In the healthcare business domain, this tool can handle the task of generating health indicator explanatory diagrams. The target text includes the indicator name, anomaly alerts, and improvement suggestions. Layout constraints require the indicator area to be at the top, the explanatory area in the middle, and the alert area at the bottom. The current diffusion denoising state controls the spatial evolution of the image in different rounds, the text semantic vector controls the distinction between indicator content and alert content, the layout constraints control the boundaries and relative positions of the three regions, and the spatial attention map output by the cross-attention layer reflects the distribution of indicator text, explanatory text, and alert text in the image space.
[0041] In this embodiment, by combining the current diffusion denoising state with the text semantic vector and layout constraint representation in the generation framework, the image formation process is no longer controlled solely by semantic content, but also by diffusion round information and spatial constraint information. The spatial attention map output by the cross-attention layer can explicitly preserve the correspondence between object semantics and image regions. The intermediate image representation is more likely to maintain the stability of object region positions, object region boundaries, and relative distribution between objects during the evolution process, thereby reducing the inconsistency between image content and layout requirements.
[0042] S40, determine spatial coherence loss based on the spatial attention map and the training layout annotation, determine reconstruction loss based on the noise prediction result corresponding to the intermediate image representation, and determine stochastic gradient regularization term based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. In this embodiment, the spatial attention map is used to represent the response distribution of text semantics in the image space, and the training layout annotation is used to represent object categories, object target regions, and spatial correspondences between objects. The spatial coherence loss depends on the degree of alignment between the two in the spatial dimension. In implementation, local response regions corresponding to each object category can be extracted from the spatial attention map according to the object categories in the training layout annotation, and then the object target regions in the training layout annotation can be mapped to the spatial attention map. Figure 1 At a consistent resolution, a target layout distribution is formed. The response intensity in the spatial attention map reflects the degree of attention paid to the object region during image generation, while the target layout distribution reflects the range of locations where the object should appear. The difference between the two can be calculated as spatial coherence loss through normalized distribution error. If there are many objects, the error can be calculated separately for each object category before aggregation. In fintech businesses, the distribution of the yield curve area, risk warning area, and transaction entry area can serve as regional information in the training layout annotation; in healthcare businesses, the distribution of the indicator area, explanation area, and reminder area can serve as regional information in the training layout annotation.
[0043] The noise prediction result corresponding to the intermediate image representation is used to measure the deviation between the image formation process and the target denoising direction. The intermediate image representation retains the latent features of the image in the current round, and the noise prediction result is given by the noise prediction branch based on the intermediate image representation. In implementation, the target noise can be read from the current diffusion denoising state, and then the noise prediction result is compared with the target noise element by element to form the reconstruction loss. The reconstruction loss reflects the degree to which the current latent features of the image approximate the denoising target. In fintech business, chart boundaries, numerical regions, and label regions, and in healthcare business, indicator borders, reminder icons, and text description regions, all rely on the noise prediction result to gradually approach the target distribution. The smaller the reconstruction loss, the easier it is to maintain the stability of object boundaries and layout content in the image space.
[0044] The stochastic gradient regularization term acts on the parameter gradient of the low-rank adjustment matrix to suppress overly steep gradient changes during parameter updates. In implementation, the gradient of the low-rank adjustment matrix is first calculated based on the reconstruction loss to obtain the original parameter gradient. Then, a perturbation vector is superimposed on the original parameter gradient to form the perturbation parameter gradient. The difference between the original parameter gradient and the perturbation parameter gradient reflects the sensitivity of the low-rank adjustment matrix to small perturbations. This difference is aggregated to obtain the stochastic gradient regularization term. This constraint reduces excessive shifts in the low-rank adjustment matrix in a few directions, ensuring a relatively stable synergy between layout constraints and image content constraints during parameter updates.
[0045] In one implementation, the spatial attention map is grouped by object category to form multiple object response distributions, and the training layout annotations are projected into regions to form multiple target layout distributions. Both are normalized at a uniform resolution, and the mean square error between each object response distribution and its corresponding target layout distribution is used as the spatial coherence loss. The intermediate image representation is fed into a noise prediction unit to obtain the noise prediction result, which is then subtracted from the target noise in the current diffusion denoising state to obtain the reconstruction loss. The gradient of the original parameters of the low-rank adjustment matrix is obtained through automatic differentiation, the perturbation vector is generated using a zero-mean distribution, and the difference norm between the perturbation parameter gradient and the original parameter gradient is used as a stochastic gradient regularization term. This implementation is suitable for image generation tasks with relatively clear object region boundaries, such as revenue display charts and risk warning charts in fintech businesses, and also suitable for indicator explanation charts and reminder charts in healthcare businesses.
[0046] In another implementation, the spatial coherence loss does not directly use the whole-image error. Instead, it breaks down the training layout annotation into multiple regional sub-images, performs segmented statistical analysis on the response density of corresponding regions in the spatial attention map, and then forms an error term based on the statistical value of each region and the coverage ratio of the regional sub-image. The reconstruction loss uses the weighted error between the noise prediction result and the target noise, with the weights set according to the importance of the image regions. The stochastic gradient regularization term calculates the gradient difference for different parameter subsets within the low-rank adjustment matrix, and then aggregates the multiple difference results. This implementation is suitable for image generation tasks with a large number of objects and complex regional relationships, such as multi-card infographics and asset allocation charts in fintech, and multi-indicator panel charts and flowcharts in healthcare.
[0047] In the fintech business, training layout annotation can identify target regions for the profit curve area, amount area, risk warning area, and transaction entry area. The response regions corresponding to the profit curve text, amount text, and warning text in the spatial attention map are aligned with these target regions, resulting in a spatial coherence loss. The noise prediction result corresponding to the intermediate image representation is compared with the target noise to form the reconstruction loss. The parameter gradient of the low-rank adjustment matrix is superimposed on the perturbation vector to form a stochastic gradient regularization term. This combination of constraints makes it easier to maintain a clear separation between the profit curve, warning box, and button areas.
[0048] In the healthcare field, training layout annotations can identify target regions for indicator, description, and reminder areas. The response regions corresponding to the indicator names, descriptions, and reminders in the spatial attention map are aligned and compared with these target regions, forming a spatial coherence loss. The noise prediction results corresponding to the intermediate image representation are compared with the target noise to form the reconstruction loss, and the difference in parameter gradient perturbation of the low-rank adjustment matrix forms a stochastic gradient regularization term. This combination of constraints makes it easier to maintain a stable regional distribution of indicator, description, and reminder content in the image.
[0049] In this embodiment, by jointly determining the spatial coherence loss through spatial attention maps and training layout annotations, the object attention regions in the image space can maintain a closer distribution relationship with the target regions. After the noise prediction results corresponding to the intermediate image representations are included in the reconstruction loss calculation, the image content is more likely to maintain semantic consistency during denoising. Introducing a stochastic gradient regularization term into the parameter gradient of the low-rank adjustment matrix reduces fluctuations in the parameter update direction. When these three types of constraints work together, the layout position, object boundaries, and semantic content are more easily synchronized and stabilized.
[0050] S50, determine the total loss based on the spatial coherence loss, the reconstruction loss and the stochastic gradient regularization term, update the low-rank adjustment matrix based on the total loss, and keep the pre-trained weights in the generation framework unchanged until the parameters of the low-rank adjustment matrix converge, and obtain the updated low-rank adjustment matrix. In this embodiment, spatial coherence loss, reconstruction loss, and stochastic gradient regularization term jointly contribute to the formation of the total loss, which serves as a unified constraint on the direction and magnitude of parameter updates. Spatial coherence loss reflects the degree of positional fit between the object region and the target region; reconstruction loss reflects the degree to which image content converges towards the target distribution; and stochastic gradient regularization term reflects the smoothness of the low-rank adjustment matrix as parameters change. Directly adding these three terms can easily lead to one term dominating the update. In practice, it is usually necessary to first scale the three losses before combining them. Scale scaling can be achieved through normalization, scaling factor adjustment, or piecewise weighting to ensure that layout-related constraints, content-related constraints, and gradient smoothing constraints remain comparable in magnitude. After the total loss is formed, backpropagation only applies to the low-rank adjustment matrix and not to the pre-trained weights in the generation framework. In implementation, an updatable marker can be established for the low-rank adjustment matrix in the parameter set, and a frozen marker can be established for the pre-trained weights in the generation framework. The differentiation unit then selects gradient preservation or gradient masking based on these markers. This allows the layout-related correction capabilities to be concentrated and compressed onto a low-rank adjustment matrix, avoiding distribution drift caused by simultaneous changes in large-scale parameters.
[0051] The update of the low-rank adjustment matrix is not a simple replacement, but rather involves generating parameter adjustments based on the total loss in each training round, and then accumulating these adjustments to the current parameters of the low-rank adjustment matrix. This can be implemented using gradient descent-based updates or by using a variable term update method. If the low-rank adjustment matrix consists of two concatenated mapping matrices, both matrices can be updated simultaneously or separately with different learning rates, ensuring coordination between the low-dimensional compression and the upscaling / recovery parts. The pre-trained weights in the generative framework remain unchanged, meaning the text encoder, image decoder, and internal attention calculation parameters maintain their existing expressive power, and the low-rank adjustment matrix provides directional adjustments for layout control. Convergence judgment cannot rely solely on the loss value in a single round; it also requires observing the parameter change trend across multiple consecutive training rounds. In implementation, the parameter changes and total loss changes of the low-rank adjustment matrix between adjacent training rounds can be recorded simultaneously. When the parameter changes consistently fall below a preset range, and the total loss fluctuation remains within a preset range, the low-rank adjustment matrix at this point can be considered the updated low-rank adjustment matrix. This can prevent premature stopping caused by accidental decline in losses, and also avoid overtraining that leads to excessive adjustment.
[0052] In fintech applications, revenue charts, asset allocation charts, billing explanation charts, and risk warning charts often simultaneously include numerical areas, chart areas, label areas, and warning areas. Layout constraints require these areas to maintain a stable distribution, while content constraints require consistency in amounts, curves, labels, and warning text. The update process of the low-rank adjustment matrix needs to simultaneously consider both of these requirements. Similarly, in healthcare applications, indicator explanation charts, health intervention charts, dietary recommendation charts, and rehabilitation guidance charts often simultaneously include indicator areas, explanation areas, reminder areas, and process areas. The combination of total losses needs to ensure that the low-rank adjustment matrix can correct for area offsets without disrupting existing content representation. Therefore, freezing the pre-trained weights in the generation framework and updating only the low-rank adjustment matrix is more suitable for image generation tasks with well-defined layouts and sensitive object relationships.
[0053] In one implementation, the spatial coherence loss, reconstruction loss, and stochastic gradient regularization term are multiplied by independent coefficients and then summed to form the total loss. These independent coefficients are given by constants set before training. A separate parameter update table is created for the low-rank adjustment matrix during differentiation, and a frozen table is created for the pre-trained weights in the generation framework during differentiation. During backpropagation, gradients are only retained for parameters in the parameter update table. Parameter updates employ an optimizer with first-order and second-order momentum estimations, making the low-rank adjustment matrix update direction smoother when facing locally sensitive areas in financial images, such as the boundaries of the earnings chart, amount label regions, and risk warning regions. Convergence is determined using a two-condition approach: one condition is that the total loss decreases below a threshold over multiple consecutive training rounds, and the other condition is that the change in the parameters of the low-rank adjustment matrix falls below a threshold over multiple consecutive training rounds. Updates stop when both conditions are met.
[0054] In another implementation, the spatial coherence loss, reconstruction loss, and stochastic gradient regularization term are each normalized by the moving average before combination, and then the total loss is formed by a dynamic scaling factor. The dynamic scaling factor is given by the relative proportion of the three losses in the current training round, ensuring that the loss term with a larger value does not dominate the update for a long time. The two mapping matrices of the low-rank adjustment matrix use different learning rates. The mapping matrix closer to the input side has a smaller update amplitude, while the mapping matrix closer to the output side has a larger update amplitude, so that the spatial distribution correction in the image decoder is more concentrated on the region boundaries and object positions. The freezing strategy adopts the form of gradient masking, directly writing zeros to the gradient positions corresponding to the pre-trained weights in the generation framework, and retaining the original values for the gradient positions corresponding to the low-rank adjustment matrix. The convergence judgment adopts a joint judgment method of parameter change and verification image layout deviation, which is suitable for medical and health image generation tasks where the indicator area, reminder area, and description area are strongly correlated.
[0055] Another approach is to use a grouped update method. The low-rank adjustment matrix is divided into multiple parameter groups based on the attention projection position. The parameter update amount is calculated for each parameter group separately, and then the low-rank adjustment matrix is reconstructed uniformly. The total loss remains in a single form, and different update coefficients are assigned to parameter groups during parameter updates. This allows parameter groups closer to lower resolution levels to handle more overall layout correction, while parameter groups closer to higher resolution levels handle more local boundary correction. This implementation is suitable for image generation tasks with a large number of regions and significant distribution differences, such as multi-card product pages in fintech businesses, and multi-indicator overview maps in healthcare businesses.
[0056] For example, in addition to the text encoder, image decoder, and cross-attention layer, the generative framework also includes a low-rank adjustment matrix loading unit, a noise prediction branch, a loss fusion unit, and a gradient constraint unit. The text encoder receives the training text description and outputs a text semantic vector. The image decoder receives the text semantic vector, layout constraint representation, and the current diffusion denoising state, and generates intermediate image representations layer by layer. The cross-attention layer is located in multiple resolution levels of the image decoder, connected to corresponding attention projection positions. The low-rank adjustment matrix is loaded at these attention projection positions and adds additional corrections to the attention mapping results. The text encoder can employ a multi-layer context modeling structure with 8 to 24 layers. The image decoder can employ a four- to six-level resolution decoding structure, with each level containing a feature transformation unit, a cross-attention unit, and an upsampling unit. The low-rank adjustment matrix can consist of two concatenated matrices, with a rank value of 4 to 64, used to form a more spatially sensitive adjustment amount at a smaller parameter scale. The noise prediction branch is connected to the output of the image decoder and is used to output the noise prediction result based on the intermediate image representation. The loss fusion unit is used to receive the spatial coherence loss, reconstruction loss and stochastic gradient regularization term and output the total loss. The gradient constraint unit is used to shield the gradients corresponding to the pre-trained weights in the generation framework during the backpropagation stage and retain only the gradients corresponding to the low-rank adjustment matrix.
[0057] During training, a batch of training text descriptions and training layout annotations are read from the training dataset. The training text descriptions are processed by a text encoder to obtain text semantic vectors, and the training layout annotations are processed by a layout parsing unit to obtain layout constraint representations. The current diffusion denoising state is jointly determined by the temporal encoding and the latent noise representation corresponding to the current diffusion round. After the text semantic vectors, layout constraint representations, and the current diffusion denoising state are input into the image decoder, a spatial attention map corresponding to the object category is formed in multiple cross-attention layers. Simultaneously, the noise prediction result is output in the noise prediction branch. The spatial coherence loss is determined based on the region alignment error between the spatial attention map and the training layout annotations. The reconstruction loss is determined based on the difference between the noise prediction result and the target noise. The stochastic gradient regularization term is determined based on the difference between the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix and the perturbed gradient. The three losses are normalized and weighted in the loss fusion unit to form the total loss. Then, the differentiation unit updates the parameters of the low-rank adjustment matrix based on the total loss. Training parameters can be adjusted according to the task size. The batch size can be set to 8 to 64, the learning rate can be set to 10^-5 to 5×10^-4, and the adjustment coefficients corresponding to the spatial coherence loss, reconstruction loss and stochastic gradient regularization term can be set to 0.1 to 10 respectively. The training termination criterion can be that the change in parameters and the change in total loss are both below a threshold in multiple consecutive rounds.
[0058] In the fintech business field, training input can include text data such as product name, return range, risk level, subscription prompts, and chart descriptions, as well as training layout annotations for the title area, return curve area, amount area, risk warning area, and button area. The text encoder outputs a text semantic vector with product category, return label, and location relationship. The intermediate image representation and spatial attention map output by the image decoder reflect the attention distribution of the return curve, amount information, and prompt information on the layout. The noise prediction result output by the noise prediction branch is used to constrain the clarity of chart boundaries and text areas. After training, an updated low-rank adjustment matrix suitable for return poster charts, asset allocation charts, and bill description charts is obtained.
[0059] In the healthcare business domain, training input can include text data such as indicator names, anomaly alerts, dietary suggestions, exercise tips, and explanatory statements, as well as training layout annotations corresponding to indicator areas, explanation areas, alert areas, and workflow areas. The text encoder outputs text semantic vectors reflecting the content and positional relationships of the indicators, while the image decoder outputs intermediate image representations and spatial attention maps reflecting the regional distribution of indicator blocks, alert blocks, and explanation blocks. The noise prediction results output by the noise prediction branch are used to constrain the stability of indicator boundaries, icon areas, and text description areas. After training, an updated low-rank adjustment matrix suitable for indicator description diagrams, health tip diagrams, and rehabilitation guidance diagrams is obtained.
[0060] In this embodiment, the total loss is composed of spatial coherence loss, reconstruction loss, and stochastic gradient regularization term. This allows for simultaneous constraints on layout position, image content, and parameter smoothness during the same update process. The update range is limited to the low-rank adjustment matrix, pre-trained weights in the generation framework remain unchanged, parameter changes are more concentrated, and the original image expressiveness is more easily preserved. Convergence judgment considers both parameter and loss changes, reducing biases caused by premature stopping and over-updating, thus maintaining a relatively stable balance between layout control and content generation within the low-rank adjustment matrix.
[0061] S60, using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation.
[0062] In this embodiment, the updated low-rank adjustment matrix serves as the parameter correction tool during the inference phase. After being loaded into the attention projection position in the generation framework, it does not alter the original large-scale parameter distribution of the generation framework, but rather adds a directional correction during the formation of the attention mapping result. The updated low-rank adjustment matrix has already undergone loss constraints during the training phase, and its internal parameters retain the object region distribution, region boundary relationships, and text semantic correspondences. The generation framework is responsible for receiving the target text semantic vector and the target layout constraint representation, and applying both simultaneously to the image formation process. The target text semantic vector includes object category, object attributes, orientation relationships, and layout relationships, while the target layout constraint representation includes the object target region, region boundaries, and relative positions between regions. During image generation, the target text semantic vector provides content constraints, and the target layout constraint representation provides spatial constraints. The updated low-rank adjustment matrix transforms these content and spatial constraints into directional corrections to the attention projection position, making it easier for the image decoder to form an image feature distribution consistent with the target text semantic vector and the target layout constraint representation during the denoising process.
[0063] The target image is not obtained directly from a single semantic mapping, but is gradually formed through multiple rounds of feature evolution within the generation framework. In implementation, a spatial constraint tensor matching the latent space of the image is first generated based on the target layout constraint representation. Then, the target text semantic vector is mapped to a conditional feature space readable by the cross-attention layer. In each round of denoising, the image decoder reads the target text semantic vector and the target layout constraint representation, and adjusts the attention projection results with the participation of an updated low-rank adjustment matrix, ensuring that the response region corresponding to the object category continuously moves closer to the target region. During the target image formation process, if the image content includes multiple regions such as title areas, chart areas, description areas, and prompt areas, the updated low-rank adjustment matrix will prioritize correcting the attention projection positions that have a greater impact on region division. If the image content includes local details such as indicator charts, curve charts, icon areas, and text areas, the updated low-rank adjustment matrix will also refine and correct the attention mapping results near the boundaries. The resulting target image retains both the object meaning in the text semantic vector and the regional relationships in the target layout constraint representation.
[0064] A correspondence between the target text semantic vector and the target layout constraint representation must be maintained. When the object name and object category are inconsistent, the target image may contain correct regions but incorrect content, or correct content but incorrect placement. In implementation, an object category mapping table can be established before inference to ensure that the object markers in the target text semantic vector correspond to the target object regions in the target layout constraint representation. The mapping results are then input into the generation framework. For revenue display charts, risk warning charts, and bill interpretation charts in fintech businesses, the product name, revenue label, and risk warning statements in the target text semantic vector need to correspond to the title area, chart area, and prompt area in the target layout constraint representation. For indicator explanation charts, reminder charts, and health guidance charts in healthcare businesses, the indicator items, reminder items, and suggestion items in the target text semantic vector need to correspond to the indicator area, reminder area, and explanation area in the target layout constraint representation. This makes it easier for the resulting target image to simultaneously meet the requirements in terms of content and space.
[0065] In one implementation, the generation framework employs a latent diffusion generation structure. The updated low-rank adjustment matrix is loaded onto multiple attention projection positions. The target text semantic vector, after linear mapping, is input into a cross-attention layer. The target layout constraint representation, after convolutional mapping, forms a multi-channel spatial constraint tensor, which, along with latent noise features, is then input into the image decoder. The image decoder performs feature transformation and upsampling at multiple resolution levels. The updated low-rank adjustment matrix corrects the attention mapping results at each level, ensuring that object regions such as title areas, chart areas, and prompt areas remain spatially stable. This implementation is suitable for benefit display charts, product card charts, risk warning charts, as well as indicator display charts, reminder and explanation charts, and health guidance charts.
[0066] In another implementation, the updated low-rank adjustment matrix is only loaded at the attention projection positions of the low-to-medium resolution layers, while the high-resolution layers retain the original mapping results. After text encoding, the target text semantic vector outputs both global semantic features and object-level semantic features. The global semantic features control overall content consistency, while the object-level semantic features control the correspondence between object categories and target regions. Target layout constraints are represented using region masks, with different object target regions written into different channels, and then scaled progressively according to resolution before being input into multiple layers. This implementation is more suitable for image generation content with clearly defined object regions and well-defined layout partitions, such as multi-card infographics and asset allocation charts in fintech, and also suitable for multi-indicator overview charts and nutrition advice charts in healthcare.
[0067] Alternatively, a position index-enhanced implementation can be used. The target layout constraint includes not only the target object region but also alignment and spacing information between objects. The updated low-rank adjustment matrix applies different adjustment magnitudes to the attention projection positions based on the position index. Attention projection positions closer to the overall partition handle region position correction, while those closer to local boundaries handle region contour correction. This implementation is suitable for image content requiring both accurate overall layout and clear local boundaries.
[0068] In this embodiment, after the updated low-rank adjustment matrix participates in image formation, the object meaning in the target text semantic vector and the object region in the target layout constraint representation can maintain a continuous correspondence during the generation process. The correction effect of the attention projection position is concentrated on the position that has a more direct impact on the region distribution. The original parameters of the generation framework remain stable, and the object position, region boundary and relative layout in the image content are more likely to be consistent with the target layout constraint representation, while preserving the content expression in the target text semantic vector.
[0069] In one embodiment, step S10 above includes: S101, Establish the connection mapping relationship between the output end of the text encoder and the decoding paths of each level in the image decoder to form the generation frame connection table; S102, based on the generated framework connection table, determine the target decoding path in the image decoder that receives the output of the text encoder, and locate multiple attention projection positions in the self-attention calculation link of each target decoding path; S103, assign a location identifier and a hierarchy identifier to each attention projection location, group the multiple attention projection locations hierarchically according to the hierarchy identifier, and configure a rank control index for each hierarchical group; S104, at each attention projection position, the original self-attention projection calculation path is retained and connected in parallel to a low-rank adjustment matrix. The low-rank adjustment matrix includes a first low-rank matrix and a second low-rank matrix connected in sequence. The first low-rank matrix receives the input features of the corresponding attention projection position, and the second low-rank matrix receives the output features of the first low-rank matrix. The connection dimension between the first low-rank matrix and the second low-rank matrix is determined by the rank control index of the corresponding hierarchical group. S105, based on the position identifier, the output features of the low-rank adjustment matrix and the output features of the original self-attention projection calculation path are merged to generate an updated projection result corresponding to the attention projection position; S106, integrate the updated projection results of each attention projection position into the self-attention computation link corresponding to the generation framework to obtain the generation framework that embeds a low-rank adjustment matrix at the attention projection position.
[0070] In this embodiment, a connection mapping relationship is established between the output of the text encoder and the decoding paths at each level within the image decoder, undertaking the localization work between semantic feature distribution and spatial feature reception. The output of the text encoder forms continuous features in the semantic domain, while the decoding paths at each level within the image decoder receive image features at different resolution levels. If there is no clear mapping between the two, semantic features entering the image side are prone to problems such as scattered injection positions, unclear hierarchical scope, and mixed reception relationships at different levels. In implementation, each decoding path in the image decoder can be assigned an independent path number, and then a correspondence can be established between the output of the text encoder and each path number to form a generative frame connection table. The generative frame connection table records at least the source of text features, decoding path number, feature injection position, and hierarchical affiliation information, so that semantic features can enter the image feature evolution process hierarchically, rather than acting indiscriminately on all positions. After the connection mapping relationship is established, the input position of text features to the image side is searchable and traceable, and subsequent projection position localization and low-rank adjustment matrix loading depend on this relationship.
[0071] The determination of the target decoding path involves filtering out the valid paths that receive the text encoder output from all decoding paths, and then locating adjustable positions within these valid paths. Image decoders typically have multiple resolution levels, and the self-attention computation units in each level have varying degrees of influence on the coarse layout, region boundaries, and local details. Not every position needs to bear the same degree of semantic adjustment. Based on the generative framework connection table, it can be determined which decoding paths receive semantic input from the text encoder, and then multiple attention projection positions can be located within the self-attention computation units of these target decoding paths. These attention projection positions specifically refer to adjustable positions in the query mapping, key mapping, and value mapping. The self-attention projection positions represent the mapping entry points when modeling the internal correlation of image-side features. In implementation, positioning can be achieved using a ternary index of layer number, block number, and mapping type, ensuring that each attention projection position is uniquely identified. After this processing, subsequent adjustments no longer apply to the entire image decoder, but instead focus on key mapping positions that influence spatial distribution.
[0072] Location identifiers and hierarchy identifiers are responsible for accurately distinguishing multiple attention projection locations. Location identifiers distinguish different self-attention projection locations within the same hierarchy, while hierarchy identifiers distinguish projection locations in different resolution hierarchy levels. Having only location identifiers without hierarchy identifiers prevents the representation of differences in control granularity between low-resolution and high-resolution locations; having only hierarchy identifiers without location identifiers prevents independent management of different mapping locations within the same hierarchy. In implementation, hierarchy identifiers can be set to the resolution hierarchy number, and location identifiers can be set to a combination of the block number within the hierarchy and the projection type number. After completing the dual identifiers, multiple attention projection locations are then hierarchically grouped based on the hierarchy identifiers. Hierarchical grouping is not a formal classification, but rather a means to allow different resolution hierarchy levels to carry different intensities of low-rank adjustment. Low-resolution hierarchy levels mainly affect the overall area layout and object partitioning, while high-resolution hierarchy levels mainly affect local boundaries, icon outlines, and text adjacency relationships. Therefore, different hierarchy groupings require different rank control indices. The rank control index is used to determine the connection dimension within the low-rank adjustment matrix, ensuring that the adjustment capacity of different hierarchy levels matches the spatial control tasks undertaken by each level.
[0073] After retaining the original self-attention projection computation path, a low-rank adjustment matrix is connected in parallel. Its purpose is to add trainable corrections without altering the core function of the original mapping path. The original self-attention projection computation path provides the basic projection results, while the low-rank adjustment matrix provides additional adjustment results. Directly replacing the original self-attention projection computation path would completely alter the spatial modeling capabilities of the image decoder; parallel connection maintains the continuity of the original mapping, with the added adjustments only superimposed at the projection result level. The low-rank adjustment matrix is formed by sequentially connecting a first low-rank matrix and a second low-rank matrix. The first low-rank matrix compresses the input features into a lower-dimensional space, while the second low-rank matrix restores the compressed adjustment features to their original dimensions. The connection dimension is given by the rank control index, allowing for separate setting of parameter sizes and adjustment strengths for different level groups. This low-rank structure concentrates the newly added trainable parameters in a smaller-dimensional space, preserving adjustment capabilities while avoiding excessive parameter expansion at the attention projection location. For levels with significant changes in coarse layout distribution, a larger connection dimension can be configured to enhance region rearrangement capabilities; for levels requiring only boundary refinement, a smaller connection dimension can be configured to reduce local perturbations.
[0074] When location identifiers participate in feature merging, their role is to accurately map the output of the low-rank adjustment matrix to the correct position in the original self-attention projection calculation path. Only after the output features of the low-rank adjustment matrix and the output features of the original self-attention projection calculation path are consistent in dimension can element-wise superposition, gated weighting, or proportional fusion be performed. Different merging methods result in different controls on the magnitude of the updated projection result. Element-wise superposition is suitable for maintaining the original projection direction and increasing the correction amount; gated weighting is suitable for changing the correction magnitude according to position or hierarchical differences; and proportional fusion is suitable for limiting the upper bound of the adjustment amount. In implementation, the corresponding fusion parameters can be read based on the location identifier, so that different projection positions within the same level have different adjustment coefficients. The generated updated projection result is no longer simply the original mapping result, nor is it an additional result detached from the original mapping, but rather a comprehensive mapping result that reflects both the output features of the original self-attention projection calculation path and the output features of the low-rank adjustment matrix.
[0075] After the updated projection results are integrated into the self-attention computation link corresponding to the generation framework, the image decoder in the generation framework will directly use these corrected projection results during subsequent spatial feature updates. Integration involves replacing the projection output in the original self-attention computation unit with the updated projection results, so that both the self-attention weight formation process and the feature recombination process are affected by the low-rank adjustment matrix. In this way, after semantic features enter the target decoding path, they not only affect the cross-semantic injection position, but also affect the spatial distribution formation through the projection correction within the self-attention computation. In low-resolution layers, updated projection results are more likely to affect the overall position of the object region and the layout partitioning; in high-resolution layers, updated projection results are more likely to affect boundary sharpness, local adjacency relationships, and fine-grained arrangement. The generation framework constructed in this way, while keeping the original parameters of the text encoder and image decoder unchanged, concentrates the trainable part for spatial layout control across multiple attention projection positions, making it easier to establish stable correspondences between object categories, region positions, and relative distributions during subsequent image formation.
[0076] For example, the low-rank embedding formula for the attention projection position is:
[0077] in, This represents the original training weight matrix corresponding to the attention projection position. This represents the updated weight matrix after embedding the low-rank adjustment matrix. This represents the incremental weights corresponding to the low-rank adjustment matrix. Let represent the first low-rank matrix in the low-rank decomposition. Let represent the second low-rank matrix in the low-rank decomposition. d represents the original feature dimension of the attention projection location. k represents the low-rank dimension, and satisfies k <d。
[0078] In this embodiment, by establishing a connection mapping relationship between the text encoder output and the decoding paths at each level within the image decoder, the semantic feature injection positions have clear hierarchical affiliations and scopes of action, reducing the problem of disordered distribution of semantic input on the image side. After multiple attention projection positions are managed by position identifiers, hierarchical identifiers, and rank control indexes, the adjustment capacity and intensity in different resolution levels can be configured separately according to spatial control tasks. After the original self-attention projection calculation path is retained and connected in parallel with the low-rank adjustment matrix, the original mapping capability is maintained, and the trainable corrections are concentrated on positions more sensitive to spatial distribution, with the scale of new parameters controlled by the connection dimension. After the updated projection results are integrated into the self-attention calculation link, the layout partitions, object positions, and local boundaries are more likely to be simultaneously directionally corrected during image formation, thus reducing object region misalignment, partition mixing, and unstable local arrangement.
[0079] In one embodiment, step S20 above includes: S201, Obtain training text description and training layout annotation, and split and parse the training text description to obtain training text annotation sequence; S202, Extract object words, position words and relation words from the training text tag sequence to form a semantic mapping index table; S203, the training layout annotations are parsed to obtain object categories, object target regions and spatial correspondences between objects, forming layout annotation parsing results; S204, associate and match the semantic mapping index table with the layout annotation parsing results to form a correspondence table between object categories and spatial regions; S205, Input the training text description into the text encoder, and perform feature aggregation on the labeled features output by the text encoder according to the correspondence table between object categories and spatial regions to generate text semantic vectors; S206, Based on the correspondence table between object categories and spatial regions and the layout annotation parsing results, generate a layout constraint representation that matches the image spatial resolution.
[0080] In this embodiment, the training text description carries composite information of image content constraints and spatial relationship constraints. It includes not only object names but also the regions where objects are located, the relative positions between objects, and the organization of multiple objects within the same frame. After obtaining the training text description, it is split and parsed to transform continuous sentences into a training text tag sequence that can participate in subsequent mapping and encoding. The splitting and parsing can be accomplished using a combination of word segmentation, lexical segmentation, phrase boundary recognition, and syntactic fragment separation. The aim is to extract semantic units that affect image content and positional distribution from the continuous text. After the training text tag sequence is formed, object names, location expressions, and relational expressions in the text are no longer mixed in with the surface form of natural language but are transformed into a set of tags that can be independently retrieved and matched. This avoids the situation where only the overall sentence meaning is retained and object-level information is weakened during the text encoding stage, and also avoids the situation where objects exist but their region affiliation is unclear during the subsequent layout mapping stage.
[0081] The extraction of object terms, location terms, and relation terms is responsible for the structuring of semantic constraints. Object terms indicate the entities that need to appear in the image, location terms constrain the spatial distribution direction and region affiliation of objects, and relation terms constrain the adjacency, containment, top-bottom, left-right, front-back, or separation relationships between multiple objects. After extraction, a semantic mapping index table is formed. The semantic mapping index table records at least the label category, the position of the label in the training text label sequence, the association relationship between labels, and the mapping information between labels and potential object categories. The semantic mapping index table is not a simple vocabulary, but an intermediate structure for image generation, used to transform object constraints and spatial constraints in the text into computable index relationships. If there are different names for the same object, different expressions of the same location relationship, or different grammatical structures of the same relationship type in the text, the semantic mapping index table can also perform a normalization mapping function, so that these semantic differences are projected onto a unified object category or a unified relationship category in subsequent encoding.
[0082] Training layout annotations carry spatial distribution information of the image, including not only object categories but also target regions and spatial correspondences between multiple objects. Parsing training layout annotations requires extracting object categories, region boundaries, region sizes, region center positions, and geometric relationships between regions to form the layout annotation parsing results. Object categories are used to establish semantic consistency with object terms in the training text description; target regions indicate the area each object occupies in the image; and spatial correspondences reflect the vertical, horizontal, overlapping, spacing, alignment, and inclusion relationships between object regions. After the layout annotation parsing results are formed, the structural information in the image space no longer remains in the original annotation form but is transformed into regionalized data that can be directly used for subsequent matching, aggregation, and encoding. If the layout annotations use bounding boxes, the region center, width, height, and boundary range can be extracted from the bounding boxes; if the layout annotations use region masks, the region coverage, contour boundaries, and area distribution can be extracted from the masks; if the layout annotations contain hierarchical partitioning information, the hierarchical affiliation of the regions can also be preserved simultaneously.
[0083] The semantic mapping index table and the correspondence matching of layout annotation parsing results are responsible for binding text objects and image regions. Without this matching process, object words in the text and object target regions in the image will exist separately. The text encoding result can only express content constraints, and the layout annotation result can only express spatial constraints. It is difficult for the two to impose common constraints on the same object in subsequent image formation. During matching, a joint judgment can be made based on the consistency of object category names, semantic similarity, relational word constraints, and positional word constraints to establish a one-to-one or one-to-many correspondence between object items in the training text description and spatial regions in the training layout annotation, resulting in a correspondence table between object categories and spatial regions. The correspondence table between object categories and spatial regions can record object category identifiers, corresponding region identifiers, region position parameters, and object relationship indexes. After this correspondence table is formed, the meaning of objects in the text and the region positions in the image are aligned. Subsequent text encoding and layout representation generation are no longer parallel but separate processes, but rather mutually constrained around the same set of objects and the same set of regions.
[0084] After receiving the training text description, the text encoder does not directly output the final control quantity for image generation. Instead, it first generates label-level features. Label-level features preserve the contextual associations of object words, position words, and relation words in the training text, as well as the semantic distance between different labels. Based on the correspondence table between object categories and spatial regions, the label features output by the text encoder are aggregated. The aim is to merge multiple label features related to the same object category, the same target region, and the same relational constraint into a unified representation. Feature aggregation can employ summation aggregation, average aggregation, weighted aggregation, or gated aggregation. If an object category is simultaneously related to multiple labels in the training text, aggregation weights can be assigned to these labels using the correspondence table. If position words and relation words need to strengthen the object representation, higher weights can be applied to position word features and relation word features during aggregation. The resulting text semantic vector is no longer merely a compression of the overall meaning of the training text description, but simultaneously contains object category information, object region tendency information, and inter-object relational information. When the text semantic vector participates in image generation, it is easier to directly apply text content constraints to specific regions.
[0085] Layout constraints are generated jointly by a table of correspondences between object categories and spatial regions, and the results of layout annotation parsing. Their role is to transform discrete layout annotations into continuous spatial control variables. Using layout annotation parsing results alone only yields region boundaries and locations, while using the table of correspondences between object categories and spatial regions alone only yields the mapping relationship between objects and regions. Only by combining these two can a layout constraint representation be formed that includes both object category information and spatial location information. In implementation, based on the object target region from the layout annotation parsing results, the corresponding object category is written into a two-dimensional plane or multi-channel tensor that matches the image spatial resolution. Then, based on the table of correspondences between object categories and spatial regions, each region location is assigned an object category index, region weight, or relational weight. If the image feature resolution is lower than the original layout annotation resolution, size adaptation can be achieved through downsampling, interpolation, or region mapping; if the image feature resolution is higher than the original layout annotation resolution, region continuity can be maintained through upsampling and boundary smoothing. The layout constraint representation generated in this way simultaneously preserves object category attribution, object target region distribution, and relative positional relationships between objects. During the subsequent image generation stage, semantic content constraints and spatial location constraints can function within the same coordinate system.
[0086] In this embodiment, after training the text description by splitting and parsing it to extract object words, location words, and relational words, the object information and spatial relationship information in the text can be preserved in a structured form, and the features output by the text encoder are no longer limited to the overall sentence meaning. After training the layout annotations, the object category, the object target region, and the spatial correspondence between objects can participate in subsequent calculations in a regionalized form. After the semantic mapping index table and the layout annotation parsing results are matched, a clear binding is formed between the object category and the spatial region. The text semantic vector and the layout constraint representation are generated around the same set of objects and regions, and the image content constraints and spatial position constraints are more likely to be consistent. This can reduce the situation where the object category is separated from the target region, the image content is correct but the position is offset, and the position is correct but the object is confused.
[0087] In one embodiment, step S30 above includes: S301, determine the corresponding denoising feature representation based on the current diffusion round, and combine the current diffusion round and the denoising feature representation to form the current diffusion denoising state; S302, the text semantic vector is distributed to the text condition interface corresponding to the multiple cross-attention layers in the generation framework, and the layout constraint representation is mapped to the spatial constraint interface corresponding to the multiple cross-attention layers; S303, the text semantic vector, the layout constraint representation and the current diffusion denoising state are input into the generation framework, so that the image decoder in the generation framework generates hierarchical feature maps sequentially along multiple decoding levels; S304 performs stepwise upsampling and cross-layer feature fusion on the feature maps of each level to generate intermediate image representations; S305, In the process of generating intermediate image representations, the attention response results corresponding to the current diffusion round are extracted from multiple cross-attention layers of the generation framework; S306, based on the object category in the layout constraint representation, collect the text tag attention response corresponding to the object category from the attention response results, and perform inter-layer fusion of the text tag attention responses output by different cross-attention layers to generate a spatial attention map.
[0088] In this embodiment, the current diffusion round is used to identify the position of the image generation in the denoising process. Different diffusion rounds correspond to different image feature states; earlier rounds retain stronger noise components, while later rounds retain more structural and detailed components. The denoising feature representation is used to carry the potential image features under the current diffusion round and can be formed by the noise latent tensor, the round embedding vector, and the modulation parameters related to the round. In implementation, the current diffusion round can be mapped to a temporal embedding vector, and then concatenated, added, or affine modulated with the noise latent tensor corresponding to the current round, so that the round information and noise state are expressed in a unified feature space. After the current diffusion denoising state is obtained by combining the current diffusion round and the denoising feature representation, the image decoder can distinguish between the global structure that should be retained and the local details that should be restored, avoiding overemphasizing boundaries in the early stage of denoising or remaining in the coarse arrangement stage in the later stage of denoising.
[0089] The text semantic vectors are distributed to text conditional interfaces corresponding to multiple cross-attention layers, continuously injecting semantic content into different resolution levels. The text conditional interfaces receive labeled semantic features or aggregated semantic features; their role is not to write all image features at once, but to provide semantic constraints at different levels. Lower resolution levels are better suited for handling object categories and relationships, while higher resolution levels are better suited for handling local details and adjacency relationships. Layout constraint representations are mapped to spatial constraint interfaces corresponding to multiple cross-attention layers, transforming the object target region and its boundary information into spatial control quantities consistent with the image feature resolution of each layer. During mapping, convolutional transformations, linear mappings, or multi-scale interpolation can be applied to the layout constraint representations to ensure the object target region has corresponding dimensions at different levels. With both text conditional interfaces and spatial constraint interfaces present, semantic and spatial information are jointly read in the same cross-attention unit. Image feature updates are no longer driven solely by object meaning but also by the distribution of the target region.
[0090] The image decoder sequentially generates hierarchical feature maps along multiple decoding levels, reflecting the gradual unfolding of image content from coarse to fine distribution. These hierarchical feature maps are not the final image, but rather intermediate visual representations at different resolutions. In implementation, the image decoder performs convolutional transformations, self-attention updates, cross-attention updates, and normalization modulation within each level before feeding the output of that level into the next. Lower-level feature maps are more focused on the arrangement of regions, the number of objects, and their approximate locations, while higher-level feature maps are more focused on boundary contours, adjacent text regions, icon areas, and fine-grained textures. Stepwise upsampling expands low-resolution features to a higher-resolution space, while cross-layer feature fusion combines coarse layout information from the upper layer with detailed information from the lower layer. Fusion can employ additive fusion, post-concatenation mapping fusion, or gated fusion to simultaneously preserve both regional and detailed structural information from different levels. The intermediate image representation is formed by fusing multiple layers of feature maps, which preserves the spatial location of the object, as well as its local boundaries and relative relationships. Subsequent noise prediction and spatial alignment judgment are based on this representation.
[0091] The attention response output of the cross-attention layer reflects the degree of influence of text semantic features on image spatial features. The attention response is typically calculated from query and key features, then mapped back to the image feature space via value features. Here, we need to extract the attention response corresponding to the current diffusion round, as the attention distribution changes with image feature evolution across different diffusion rounds. Directly reading the attention result of a single layer easily yields a spatial response suitable only for a local scale; therefore, multiple cross-attention layers are required for extraction. The object category in the layout constraint representation is used to limit which text tags' corresponding response results need to be read. In implementation, attention responses related to the object category can be collected based on the index position of the object category in the text tag sequence. Then, the responses of multiple text tags are summed, averaged, or weighted according to the object category to obtain the object-level response distribution. The object-level response distributions output by different cross-attention layers differ in resolution, requiring a unified size before inter-layer fusion. Inter-layer fusion can employ upsampling followed by weighted summation or hierarchical gated fusion, allowing low-resolution layers to provide regional generality and high-resolution layers to provide boundary details. Once the spatial attention map is formed, the object attention region, object boundary trend, and the degree of separation between objects in the image space can be preserved in an explicit distribution form, and region-level alignment can be directly performed when comparing with the training layout annotation.
[0092] When object categories are used in text tagging attention response collection, they are not only used to find object names in the text, but also to constrain the distribution of ambiguous semantics in the image space. If multiple object descriptions exist in the same text or the same object contains multiple modifiers, simply reading all attention results will lead to a mixed spatial response. By limiting the scope of text tagging by object categories, content such as profit curves, monetary labels, risk warnings, indicator names, and reminder areas can form independent attention distributions. The resulting spatial attention map is closer to object-level layout information than the average attention distribution of the entire text on the image. When the spatial attention map is subsequently used for layout alignment, it can directly reflect whether object regions are close to target regions, and also whether multiple object regions overlap, shift, or have chaotic boundaries.
[0093] In this embodiment, after the current diffusion round and denoising feature representation are combined to form the current diffusion denoising state, the image decoder can distinguish the focus of coarse distribution restoration and detailed boundary restoration at different rounds, reducing the mutual interference between global layout and local details. Text semantic vectors participate in image feature updates through text conditional interfaces corresponding to multiple cross-attention layers, and layout constraint representations participate in image spatial modulation through spatial constraint interfaces corresponding to multiple cross-attention layers. Semantic content and spatial distribution play a role simultaneously at multiple resolution levels, making it easier for intermediate image representations to maintain consistency between object categories and target regions. The text tag attention responses output by multiple cross-attention layers are filtered by object categories and fused between layers to form a spatial attention map. The spatial response retains both the overall regional trend and boundary detail information. Subsequent region alignment can directly utilize the object-level spatial distribution results, thereby reducing object position offset, region overlap, and relative layout distortion.
[0094] In one embodiment, step S40 above includes: S401, based on the object categories in the training layout annotation, separate the object-level spatial attention distribution corresponding to each object category from the spatial attention map, and convert the object target region in the training layout annotation into a target layout distribution with the same spatial resolution as the spatial attention map; S402, normalize the spatial attention distribution of each object and the layout distribution of each target along the spatial dimension, and determine the category weight coefficient based on the area ratio and region overlap of each object target region; S403, based on the weight coefficients of each category, the alignment error between the normalized object-level spatial attention distribution and the target layout distribution is weighted and converged to obtain the spatial coherence loss; S404, extract the real noise component from the current diffusion denoising state, and determine the reconstruction loss based on the difference between the noise prediction result corresponding to the intermediate image characterization and the real noise component; S405, calculate the original parameter gradient of the low-rank adjustment matrix based on the reconstruction loss, and inject a perturbation vector into the original parameter gradient to obtain the perturbation parameter gradient; S406, the parameters of the low-rank adjustment matrix are divided into multiple parameter subsets, the gradient difference norm between the original parameter gradient and the perturbation parameter gradient corresponding to each parameter subset is calculated, and the gradient difference norms of each parameter subset are aggregated to obtain the stochastic gradient regularization term.
[0095] In this embodiment, when the spatial attention map and training layout annotations jointly participate in the formation of spatial coherence loss, the information in the spatial attention map no longer remains at the level of overall image attention strength, but is broken down to the object category granularity. After the training layout annotations include object categories and object target regions, the corresponding object-level spatial attention distribution can be extracted from the spatial attention map according to object category. The extraction method can employ object category index gating or object category masking to ensure that the attention intensity of each object category in the image plane is independently preserved. The object target regions in the training layout annotations need to be transformed into spatial attention... Figure 1 To achieve the desired spatial resolution, downsampling, interpolation, or region projection can be used during conversion to ensure that the target region maintains its central position, boundary range, and coverage area within the target layout distribution without distortion. Only after the object-level spatial attention distribution and the target layout distribution are unified to the same resolution can they be compared position-by-position in the spatial dimension. Normalization processes place attention intensity and region coverage intensity on a comparable numerical scale, preventing large object response values or large region areas from directly dominating the error results. The area proportion of the target region reflects the spatial importance of the object in the entire image, while the degree of region overlap reflects whether multiple objects spatially occlude each other, are adjacent and crowded, or have overlapping boundaries. After forming class weight coefficients based on area proportion and region overlap, large objects will not suppress small objects simply because of their naturally larger area, and small objects will not be weakened by layout constraints due to fewer pixels. Alignment errors, after being weighted and aggregated by class weight coefficients, form spatial coherence loss. The result simultaneously reflects whether the object's position is close to the target region, whether the object's boundary deviates from the target boundary, and whether the distribution of multiple objects meets the preset region organization method.
[0096] When the noise prediction result corresponding to the intermediate image representation is used in the formation of the reconstruction loss, the focus is on the deviation between the latent content of the image in the current round and the target denoising direction. The current diffusion denoising state contains round information and noise state information. The true noise component, extracted from the current diffusion denoising state, represents the noise component that should be removed in the current round. The intermediate image representation, after passing through the noise prediction branch, yields the noise prediction result. The noise prediction result does not directly represent the image content, but rather represents the noise structure that the image decoder believes still exists in the current round. The smaller the difference between the true noise component and the noise prediction result, the closer the intermediate image representation is to the expected denoising direction; the larger the difference, the stronger the offset in the latent features of the current image. The reconstruction loss can be given in the form of element-wise error, region-weighted error, or multi-scale error. If there are layout partitions, text regions, chart regions, and prompt regions in the image, the reconstruction loss can also be calculated separately on these regions and then synthesized into an overall error, so that the image content restoration is consistent with the spatial partition restoration. The reconstruction loss formed in this way not only constrains the overall visual result of the image, but also constrains the latent representation of the current round to converge towards the correct noise removal direction.
[0097] The stochastic gradient regularization term is formed around the parameter gradients of the low-rank adjustment matrix, aiming to limit excessive fluctuations in the parameter update direction under local perturbations. The original parameter gradient is obtained by differentiating the reconstruction loss relative to the low-rank adjustment matrix. This original parameter gradient reflects the update tendency of each parameter in the low-rank adjustment matrix under the current image content constraints. A perturbation vector is injected into the original parameter gradient to obtain the perturbed parameter gradient. The perturbation vector can be generated using random sampling or by setting the perturbation ratio based on the gradient magnitude, allowing both small and large parameters to accept moderate perturbations. After the low-rank adjustment matrix is divided into multiple parameter subsets, different subsets can correspond to different levels, different projection positions, or different rank components. With this division, the smoothness of parameter updates is no longer measured by a single statistic of the entire matrix, but rather refined to multiple local subspaces. The gradient difference norm is calculated between the original parameter gradient and the perturbed parameter gradient on each parameter subset, yielding the sensitivity of that subset to perturbations. If the gradient difference norm of a certain parameter subset is too large, it indicates that the update direction of these parameters changes significantly under small perturbations, which can easily lead to instability in the training process. If the gradient difference norms of multiple parameter subsets remain at a low level, it indicates that the overall update of the low-rank adjustment matrix is more stable. The gradient difference norms of each parameter subset are aggregated to form a stochastic gradient regularization term. The aggregation result allows spatial layout constraints, image content constraints, and parameter smoothing constraints to simultaneously apply to the updatable part of the low-rank adjustment matrix.
[0098] Spatial coherence loss, reconstruction loss, and stochastic gradient regularization terms constrain object region distribution, image content convergence direction, and parameter update stability, respectively. After compressing the alignment error between the object-level spatial attention distribution and the target layout distribution, the object positions, region boundaries, and multi-object arrangements in the image become closer to the training layout annotations. After compressing the difference between the real noise component and the noise prediction result, the intermediate image representation is more likely to move closer to the target content during denoising. After compressing the difference between the original parameter gradient and the perturbation parameter gradient, the update direction of the low-rank adjustment matrix becomes more stable. These three constraints work together in the same training process, limiting image content shift, layout shift, and the oscillation of the low-rank adjustment matrix update in local directions.
[0099] For example, the formula for object-level spatial alignment loss is:
[0100] in, This represents the spatial coherence loss. 'c' represents the object category index. 'C' represents the total number of object categories. This represents the object-level spatial attention distribution corresponding to object category c. This represents the target layout distribution corresponding to object category c. This represents the result after normalizing the spatial attention distribution of the object-level space in the spatial dimension. This represents the squared Frobenius norm, used to measure the alignment error between the object-level spatial attention distribution and the target layout distribution. This represents the category weight coefficient corresponding to object category c. It is determined by the area ratio of the target region and the degree of regional overlap.
[0101] The formula for reconstructing gradient perturbation regularization is:
[0102] in, This represents the stochastic gradient regularization term. This represents the perturbation vector injected into the parameter gradient calculation. This indicates that the mean is 0 and the variance is 0. The noise distribution. σ represents the disturbance intensity control parameter. This represents the set of parameters for the low-rank adjustment matrix. This indicates the losses incurred during reconstruction. This represents the gradient of the reconstruction loss with respect to the original parameters of the low-rank adjustment matrix. This represents the gradient of the perturbation parameters obtained under perturbation conditions. This represents the squared L2 norm, used to measure the difference between two types of gradients.
[0103] In this embodiment, by aligning the object-level spatial attention distribution with the target layout distribution according to object category, the layout constraints no longer remain at the average level of the entire image but can directly affect the object region location and region boundaries. Spatial coherence loss more effectively suppresses object offset, boundary misalignment, and multi-object crowding. The reconstruction loss formed by the true noise component and the noise prediction result pulls the potential image representation towards the correct denoising direction, and image content restoration and region distribution restoration can converge synchronously. After introducing a perturbation vector into the original parameter gradient of the low-rank adjustment matrix, the gradient difference norm is statistically calculated according to the parameter subset, and local drastic fluctuations in parameter updates are suppressed. As a result, layout accuracy, image content consistency, and parameter update smoothness can be constrained simultaneously. The low-rank adjustment matrix is more likely to maintain stable convergence during training, and the object region distribution and content representation in the image are more likely to remain consistent.
[0104] In one embodiment, step S50 above includes: S501, the low-rank adjustment matrix is divided into multiple parameter groups according to different attention projection positions; S502, based on the alignment error between the spatial attention map extracted from the path where the attention projection position of each parameter group is located and the training layout annotation, determine the layout sensitivity coefficient corresponding to each parameter group. S503, based on each layout sensitivity coefficient, the fusion strength of the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term on each parameter group is adjusted to obtain the group loss corresponding to each parameter group, and each of the group losses is determined as the in-group component of the total loss on the corresponding parameter group, and the total loss is obtained by aggregating the group losses. S504, update the corresponding parameter groups according to the in-group components of the total loss in each parameter group respectively, and block the update transmission of the in-group components of the total loss in each parameter group to the pre-trained weights in the generation framework, and only retain the update transmission of the in-group components of the total loss in each parameter group to the low-rank adjustment matrix. S505, reconstruct the low-rank adjustment matrix based on the updated parameter groups; S506, when the parameter changes of the reconstructed low-rank adjustment matrix in multiple consecutive training rounds satisfy the convergence condition, the reconstructed low-rank adjustment matrix is used as the updated low-rank adjustment matrix.
[0105] In this embodiment, after the low-rank adjustment matrix is divided into multiple parameter groups according to different attention projection positions, parameter updates are no longer treated as a unified object of the entire matrix, but rather multiple locally adjustable units participate in training separately. The attention projection positions within the image decoder are typically distributed across different resolution levels and different self-attention units, and the impact of different positions on region arrangement, boundary formation, and local details varies. When the low-rank adjustment matrix is divided into multiple parameter groups, grouping can be based on the level of the attention projection position, the channel dimension range, and the mapping type, so that each parameter group forms a stable correspondence with one or a group of attention projection positions. This method preserves the overall trainability of the low-rank adjustment matrix while allowing local region control capabilities to be individually identified during the training phase. Low-resolution levels in the image decoder have a greater impact on the overall arrangement and layout of object regions, while high-resolution levels have a greater impact on object boundaries and adjacent region details. After the parameter groups are divided according to attention projection positions, the adjustment capabilities of different levels can be controlled separately. During training, a parameter group index table can be created in the parameter management module. The index table should at least record the parameter group number, attention projection position number, the level to which it belongs, and the updatable state, so that each parameter group has an independent entry point during backpropagation and updating.
[0106] The layout sensitivity coefficient measures the strength of each parameter group's response to layout deviations. The spatial attention map preserves the response distribution of object categories in the image space, and the training layout annotations preserve the target regions of objects and their spatial correspondences. Extracting the spatial attention map based on the path of the attention projection position corresponding to each parameter group means that each parameter group can obtain spatial response results related to its responsible region. Comparing these spatial response results with the target regions in the training layout annotations to determine the alignment error allows us to obtain the sensitivity of each parameter group to layout deviations at the current position. If there is a large deviation between the spatial attention map corresponding to a parameter group and the training layout annotations, the layout sensitivity coefficient will be large, indicating that this parameter group should undertake more region correction tasks; if the deviation is small, the layout sensitivity coefficient will be small, indicating that this parameter group does not currently need to undertake excessive layout correction. In implementation, a layout sensitivity evaluation module can be set up to read the parameter group index table, spatial attention map, and training layout annotations in each training round, and output a set of layout sensitivity coefficients corresponding one-to-one with each parameter group. The layout sensitivity coefficient can be a continuous scalar or a hierarchically normalized proportional value. In fintech businesses, the profit curve area, amount area, risk warning area, and transaction button area serve different display functions on the layout, and the layout sensitivity coefficients of the corresponding parameter groups usually also differ. If the profit curve area and risk warning area are offset, the corresponding coefficients can be increased. In healthcare businesses, if the indicator area, explanation area, reminder area, and process area overlap or are misaligned, the parameter groups corresponding to these areas will also be assigned higher layout sensitivity coefficients.
[0107] When spatial coherence loss, reconstruction loss, and stochastic gradient regularization (SGR) jointly contribute to the formation of the total loss, the strength of these three constraints should not be kept constant across different parameter sets. Spatial coherence loss reflects the alignment between the object region and the target region, reconstruction loss reflects the convergence of the intermediate image representation towards the target denoising direction, and SGR reflects the smoothness of the low-rank adjustment matrix parameter updates. With the introduction of layout sensitivity coefficients, the loss fusion unit can adjust the weight ratio of these three constraints according to the layout sensitivity of different parameter sets. Parameter sets with higher layout sensitivity coefficients can increase the proportion of spatial coherence loss in the group loss, allowing region location and boundary correction to play a greater role; parameter sets with lower layout sensitivity coefficients can increase the relative weights of reconstruction loss and SGR, making image content consistency and parameter update smoothness more stable. The group losses corresponding to these parameter sets are no longer simply a copy of the overall image loss, but rather group losses reconstructed for local responsibilities at different attention projection locations. Each group loss is further defined as the in-group component of the total loss for the corresponding parameter group. The total loss is obtained by aggregating all group losses, ensuring that the total loss maintains a unified global update objective while preserving the local differences in group adjustments. In implementation, a loss fusion module can be set up to receive spatial coherence loss, reconstruction loss, stochastic gradient regularization term, and layout sensitivity coefficient. The fusion strength of these three losses across each parameter group is adjusted before outputting the group loss and the total loss. Training parameters can be configured according to the group update method. The base learning rate can be set from 1e-5 to 5e-4. After the layout sensitivity coefficient participates in the in-group learning rate adjustment, different parameter groups can have different update amplitudes.
[0108] When the in-group components of the total loss update the corresponding parameter groups, the pre-trained weights remain frozen, limiting the update range to within the low-rank adjustment matrix. Freezing not only stops parameter writing but also requires masking the gradients corresponding to the pre-trained weights during backpropagation. In implementation, gradient masking can be established for the pre-trained weights in the generation framework within the gradient routing module, and gradient retention markers can be established for the parameter groups of the low-rank adjustment matrix. After differentiation, the gradient components belonging to the pre-trained weights are directly cleared to zero, while the gradient components belonging to the parameter groups of the low-rank adjustment matrix are retained and updated according to the in-group components of the total loss in each parameter group. This update method keeps the original large-scale parameters in the text encoder, image decoder, and attention calculation unit stable, allowing only the low-rank adjustment matrix to handle layout correction and content adaptation tasks. For revenue display charts, asset allocation charts, and billing explanation charts in fintech businesses, freezing the pre-trained weights preserves the original chart, number, and layout element formation capabilities, while parameter group updates primarily correct the positional relationships between curve areas, title areas, and prompt areas. For indicator illustration charts, reminder charts, and intervention guidance charts in healthcare operations, pre-training weight freezing can preserve the original ability to form icons and text content, while parameter group updates mainly correct the spatial allocation relationship between the indicator area, reminder area, and explanation area.
[0109] The updated parameter groups need to be reassembled back into the low-rank adjustment matrix to continue participating in attention projection adjustment in subsequent training rounds. During reconstruction, parameter blocks cannot be simply pieced together; instead, each parameter group should be written back to the parameter slot corresponding to its original attention projection position according to the parameter group index table. This ensures consistency in the overall dimension, hierarchical mapping, and projection position assignment of the low-rank adjustment matrix. If the reconstruction position is incorrect, the correspondence between the subsequent spatial attention map and the training layout annotation will be disrupted, and the effect of group updates cannot be accumulated. In implementation, a matrix reconstruction module can be set up to write the updated parameter groups back to their corresponding positions in the low-rank adjustment matrix according to the position mapping relationship in the parameter group index table, and then reload the reconstructed low-rank adjustment matrix into the generation framework. The parameter changes in multiple consecutive training rounds are used to determine the convergence state. This determination involves examining whether the magnitude of the parameter group changes after updates continuously decreases, and whether the total loss fluctuates stably within a preset range. If the parameter changes are consistently below a threshold and the total loss fluctuation is also below a threshold, the reconstructed low-rank adjustment matrix can be determined as the updated low-rank adjustment matrix. During training, you can set an upper limit for the number of training epochs, a parameter variation threshold, and a total loss fluctuation threshold. For example, you can set the number of training epochs to 10^4 to 10^5, the parameter variation threshold to 1e-6 to 1e-4, and the total loss fluctuation threshold to 1e-5 to 1e-3. This can prevent premature stopping caused by a random decrease in loss, and also prevent prolonged ineffective training caused by continuous small oscillations in parameters.
[0110] For example, the formula for the total loss of the parameter set is:
[0111] in, This indicates the total loss. This indicates the losses incurred during reconstruction. This represents the stochastic gradient regularization term. This represents the spatial coherence loss. g represents the parameter set index. G represents the total number of parameter sets. This represents the fusion intensity coefficient of the reconstruction loss on parameter group g. This represents the fusion strength coefficient of the stochastic gradient regularization term on the parameter set g. These represent the fusion intensity coefficients that represent the spatial coherence loss on parameter group g. These fusion intensity coefficients are obtained by adjusting the layout sensitivity coefficients.
[0112] In this embodiment, after dividing the parameter groups according to the attention projection position using a low-rank adjustment matrix, the layout bias is no longer corrected uniformly as a whole block of parameters, but is distributed to local update units corresponding to different spatial regions and different resolution levels. After the layout sensitivity coefficient participates in the fusion strength adjustment of the three losses, the parameter group with larger regional bias can obtain stronger layout correction constraints, while the parameter group with smaller regional bias can maintain a stable state of content recovery and gradient smoothing. The intra-group components of the total loss in each parameter group are applied to the corresponding parameter group, the pre-trained weights are kept frozen, and the parameter changes are concentrated within the low-rank adjustment matrix, making it easier to retain the original expressive power of the generated framework. After the updated parameter group is reconstructed into a low-rank adjustment matrix and convergence is judged by the parameter change amount in several consecutive training rounds, the position, boundary relationship and relative arrangement of object regions in the image are more likely to stabilize near the target distribution, while avoiding excessive swings in the parameter update direction in local regions.
[0113] In one embodiment, step S60 above includes: S601, the updated low-rank adjustment matrix is loaded into the attention projection position corresponding to the low-rank adjustment matrix of the generation framework; S602, the target text semantic vector is distributed to the text condition interfaces corresponding to multiple cross-attention layers in the generation framework, and the target layout constraint representation is mapped to the spatial constraint interfaces corresponding to multiple cross-attention layers in the generation framework. S603, determine the initial diffusion denoising state based on the target layout constraint representation; S604, the target text semantic vector, the target layout constraint representation and the initial diffusion denoising state are input into the generation framework. Under the condition that the updated low-rank adjustment matrix participates in the adjustment, the image decoder in the generation framework performs round-by-round denoising and multi-level upsampling along multiple decoding levels to generate the target image representation. S605, based on the target layout constraint representation, perform consistency verification on the object region position, object region range and relative layout between object regions in the target image representation, and determine the target image representation that passes the verification. S606, the verified target image representation is input into the output of the generation framework for image reconstruction to obtain the target image.
[0114] In this embodiment, after the updated low-rank adjustment matrix is loaded into the attention projection position in the generation framework, the attention projection unit no longer simply outputs the original projection result, but instead superimposes the low-rank correction amount that has converged during the training phase during the formation of the original projection result. During loading, the low-rank parameters are written to the corresponding positions according to the attention projection position index, layer index, and channel range, ensuring that the correction magnitude in different resolution layers remains consistent with that during the training phase. The target text semantic vector output by the text encoder is distributed to multiple cross-attention layers via the text conditional interface, enabling object categories, object attributes, orientation relationships, and inter-object relationships to continuously participate in image feature updates at different levels. The target layout constraint representation is mapped to a spatial control tensor consistent with the feature resolution of each layer via the spatial constraint interface. The spatial control tensor retains the target object region, region boundaries, and relative positions between regions. With this setup, semantic constraints are responsible for limiting the generated content, spatial constraints are responsible for limiting the region distribution, and the low-rank adjustment matrix is responsible for writing the layout correction capability learned during the training phase into the attention mapping process.
[0115] The determination of the initial diffusion denoising state needs to consider both noise distribution and target region distribution. In implementation, the target layout constraint representation can be projected onto the latent space to obtain a region prior with the same size as the latent feature. This region prior is then combined with a preset noise distribution to form the initial diffusion denoising state with spatial guidance information. After receiving the target text semantic vector, the target layout constraint representation, and the initial diffusion denoising state, the image decoder performs feature transformation, cross-attention modulation, and resolution upscaling across multiple decoding levels. Attention projection positions in low-resolution layers focus more on controlling the overall layout and object arrangement, while attention projection positions in high-resolution layers focus more on controlling region boundaries, icon outlines, text adjacency, and local details. The updated low-rank adjustment matrix continuously provides directional corrections at these positions, causing the object meaning in the target text semantic vector and the object region in the target layout constraint representation to gradually converge during denoising. With each round of denoising and multi-level upsampling, the generation framework outputs a target image representation that simultaneously preserves object content, region boundaries, and relative layout information between objects.
[0116] Consistency verification involves comparing the generated result with the target layout constraint representation at the region level. During verification, the object region location, object region extent, and relative layout between objects are extracted from the target image representation and then compared item by item with the target region location, target region size, and region relationships in the target layout constraint representation. The object region location can be described using the region center, bounding box, or region centroid; the object region extent can be described using width and height, area, or mask coverage; and the relative layout can be described using top-bottom, left-right, alignment, overlap, and spacing relationships. If the object region location deviates from the target region, the object region extent exceeds the boundary, or the relative layout between multiple objects is inconsistent with the target layout constraint representation, the target image representation fails verification. The verified target image representation is then fed into the output of the generation framework for image reconstruction. The output can consist of a decoding convolutional layer, a color mapping layer, and a pixel recovery unit to restore the latent features to the target image. The resulting target image not only semantically matches the target text semantic vector but also matches the target layout constraint representation in terms of region location and region relationships.
[0117] In fintech scenarios, the target text semantic vector can be formed from data such as product name, benefit description, risk warning, and transaction prompts. The target layout constraint representation can be formed by area markings in the title area, chart area, amount area, risk warning area, and button area. The output data is a benefit display chart, product card chart, or billing explanation chart with regional stability. In healthcare scenarios, the target text semantic vector can be formed from data such as indicator name, anomaly alerts, intervention suggestions, and explanatory statements. The target layout constraint representation can be formed by area markings in the indicator area, alert area, explanation area, and suggestion area. The output data is an indicator explanation chart, health prompt chart, or intervention guidance chart with regional boundary stability. After the input data is set in both semantic and spatial dimensions, the attention projection position in the generation framework can continuously receive content constraints and layout constraints, and the low-rank adjustment matrix retains the layout correction capability formed in the training phase during the inference phase.
[0118] This embodiment loads the updated low-rank adjustment matrix into the attention projection position and allows the target text semantic vector, target layout constraint representation, and initial diffusion denoising state to work together at multiple decoding levels. This ensures that the object content, object regions, and relative layout between objects are simultaneously constrained during the generation process. By performing consistency verification on the target image representation and then sending the verified target image representation to the output to complete image reconstruction, the generated result not only more closely approximates the content distribution defined by the target text semantic vector but also more closely approximates the regional positions and relationships defined by the target layout constraint representation. This reduces object offset, region misalignment, and imbalance in the arrangement of multiple objects.
[0119] In one embodiment, an image generation apparatus based on low-rank parameter adaptation is provided, which corresponds one-to-one with the image generation method based on low-rank parameter adaptation described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the image generation device based on low-rank parameter adaptation of the present invention. The modules include a generation framework construction module 10, a training input processing module 20, an intermediate representation generation module 30, a triple loss calculation module 40, a low-rank adjustment and update module 50, and a target image generation module 60. Detailed descriptions of each functional module are as follows: Generative framework building module 10 is used to build a generative framework, which includes a text encoder and an image decoder, and embeds a low-rank adjustment matrix at the attention projection position of the generative framework; The training input processing module 20 is used to obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations. The intermediate representation generation module 30 is used to determine the current diffusion denoising state, generate intermediate image representations based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract spatial attention maps from the cross attention layer of the generation framework. The triple loss calculation module 40 is used to determine the spatial coherence loss based on the spatial attention map and the training layout annotation, determine the reconstruction loss based on the noise prediction result corresponding to the intermediate image representation, and determine the stochastic gradient regularization term based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The low-rank adjustment update module 50 is used to determine the total loss based on the spatial coherence loss, the reconstruction loss and the stochastic gradient regularization term, update the low-rank adjustment matrix based on the total loss, and keep the pre-trained weights in the generation framework unchanged until the parameters of the low-rank adjustment matrix converge, so as to obtain the updated low-rank adjustment matrix. The target image generation module 60 is used to generate a target image based on the target text semantic vector and the target layout constraint representation using the updated low-rank adjustment matrix and the generation framework.
[0120] In one embodiment, the framework building module 10 is specifically used for: Establish the connection mapping relationship between the output of the text encoder and the decoding paths of each level in the image decoder to form the generation frame connection table; Based on the generated framework connection table, the target decoding path that receives the output of the text encoder in the image decoder is determined, and multiple attention projection positions are located in the self-attention calculation link of each target decoding path; Assign a location identifier and a hierarchy identifier to each attention projection location, group the multiple attention projection locations hierarchically according to the hierarchy identifier, and configure a rank control index for each hierarchical group; At each attention projection location, the original self-attention projection calculation path is retained and connected in parallel to a low-rank adjustment matrix. The low-rank adjustment matrix includes a first low-rank matrix and a second low-rank matrix connected in sequence. The first low-rank matrix receives the input features of the corresponding attention projection location, and the second low-rank matrix receives the output features of the first low-rank matrix. The connection dimension between the first low-rank matrix and the second low-rank matrix is determined by the rank control index of the corresponding hierarchical group. Based on the location identifier, the output features of the low-rank adjustment matrix and the output features of the original self-attention projection calculation path are merged to generate an updated projection result for the corresponding attention projection position. The updated projection results at each attention projection position are integrated into the self-attention computation link corresponding to the generative framework to obtain a generative framework that embeds a low-rank adjustment matrix at the attention projection position.
[0121] In one embodiment, the training input processing module 20 is specifically used for: Obtain training text descriptions and training layout annotations, and split and parse the training text descriptions to obtain training text tag sequences; Extract object words, location words, and relation words from the training text tag sequence to form a semantic mapping index table; The training layout annotations are parsed to obtain object categories, object target regions, and spatial correspondences between objects, forming layout annotation parsing results; The semantic mapping index table is associated and matched with the layout annotation parsing results to form a correspondence table between object categories and spatial regions; The training text description is input into the text encoder. Based on the correspondence table between object categories and spatial regions, the marked features output by the text encoder are aggregated to generate a text semantic vector. Based on the correspondence table between object categories and spatial regions and the layout annotation parsing results, a layout constraint representation that matches the image spatial resolution is generated.
[0122] In one embodiment, the intermediate characterization generation module 30 is specifically used for: Determine the corresponding denoising feature representation based on the current diffusion round, and combine the current diffusion round with the denoising feature representation to form the current diffusion denoising state; The text semantic vector is distributed to the text condition interface corresponding to multiple cross-attention layers in the generation framework, and the layout constraint representation is mapped to the spatial constraint interface corresponding to multiple cross-attention layers. The text semantic vector, the layout constraint representation, and the current diffusion denoising state are input into the generation framework, so that the image decoder in the generation framework generates hierarchical feature maps sequentially along multiple decoding levels. Perform stepwise upsampling and cross-layer feature fusion on the feature maps at each level to generate intermediate image representations; During the generation of intermediate image representations, attention response results corresponding to the current diffusion round are extracted from multiple cross-attention layers of the generation framework; Based on the object category in the layout constraint representation, text tag attention responses corresponding to the object category are collected from the attention response results, and inter-layer fusion is performed on the text tag attention responses output by different cross-attention layers to generate a spatial attention map.
[0123] In one embodiment, the triple loss calculation module 40 is specifically used for: Based on the object categories in the training layout annotations, object-level spatial attention distributions corresponding to each object category are separated from the spatial attention map, and the object target regions in the training layout annotations are converted into target layout distributions with the same spatial resolution as the spatial attention map. The spatial attention distribution of each object and the layout distribution of each target are normalized along the spatial dimension, and the category weight coefficient is determined based on the area ratio and region overlap of each object target region. Based on the weight coefficients of each category, the alignment error between the normalized object-level spatial attention distribution and the target layout distribution is weighted and converged to obtain the spatial coherence loss. Extract the real noise component from the current diffusion denoising state, and determine the reconstruction loss based on the difference between the noise prediction result corresponding to the intermediate image characterization and the real noise component; The original parameter gradient of the low-rank adjustment matrix is calculated based on the reconstruction loss, and a perturbation vector is injected into the original parameter gradient to obtain the perturbation parameter gradient; The parameters of the low-rank adjustment matrix are divided into multiple parameter subsets. The gradient difference norm between the original parameter gradient and the perturbation parameter gradient corresponding to each parameter subset is calculated. The gradient difference norms of each parameter subset are aggregated to obtain the stochastic gradient regularization term.
[0124] In one embodiment, the low-rank adjustment update module 50 is specifically used for: The low-rank adjustment matrix is divided into multiple parameter groups according to different attention projection positions; Based on the alignment error between the spatial attention map extracted from the path where the attention projection position of each parameter group is located and the training layout annotation, the layout sensitivity coefficient corresponding to each parameter group is determined. Based on each layout sensitivity coefficient, the fusion strength of the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term on each parameter group is adjusted to obtain the group loss corresponding to each parameter group. Each group loss is then determined as the in-group component of the total loss on the corresponding parameter group, and the total loss is obtained by aggregating the group losses. The corresponding parameter groups are updated according to the in-group components of the total loss in each parameter group, and the update transmission of the in-group components of the total loss in each parameter group to the pre-trained weights in the generation framework is blocked. Only the update transmission of the in-group components of the total loss in each parameter group to the low-rank adjustment matrix is retained. Reconstruct the low-rank adjustment matrix based on the updated parameter sets; When the parameter changes of the reconstructed low-rank adjustment matrix satisfy the convergence condition in multiple consecutive training rounds, the reconstructed low-rank adjustment matrix is used as the updated low-rank adjustment matrix.
[0125] In one embodiment, the target image generation module 60 is specifically used for: The updated low-rank adjustment matrix is loaded into the attention projection position corresponding to the low-rank adjustment matrix in the generation framework; The target text semantic vector is distributed to the text condition interfaces corresponding to multiple cross-attention layers in the generation framework, and the target layout constraint representation is mapped to the spatial constraint interfaces corresponding to multiple cross-attention layers in the generation framework. The initial diffusion denoising state is determined based on the target layout constraints. The target text semantic vector, the target layout constraint representation, and the initial diffusion denoising state are input into the generation framework. Under the condition that the updated low-rank adjustment matrix participates in the adjustment, the image decoder in the generation framework performs round-by-round denoising and multi-level upsampling along multiple decoding levels to generate the target image representation. Based on the target layout constraint representation, the consistency of the object region position, object region range, and relative layout between object regions in the target image representation is checked, and the target image representation that passes the check is determined. The verified target image representation is input into the output of the generation framework for image reconstruction to obtain the target image.
[0126] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a server-side image generation method based on low-rank parameter adaptation.
[0127] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a low-rank parameter adaptation-based image generation method.
[0128] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: A generative framework is constructed, which includes a text encoder and an image decoder, and a low-rank adjustment matrix is embedded at the attention projection position of the generative framework; Obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations; Determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract a spatial attention map from the cross attention layer of the generation framework; The spatial coherence loss is determined based on the spatial attention map and the training layout annotation; the reconstruction loss is determined based on the noise prediction result corresponding to the intermediate image representation; and the stochastic gradient regularization term is determined based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The total loss is determined based on the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term. The low-rank adjustment matrix is updated based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, thus obtaining the updated low-rank adjustment matrix. Using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation.
[0129] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: A generative framework is constructed, which includes a text encoder and an image decoder, and a low-rank adjustment matrix is embedded at the attention projection position of the generative framework; Obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations; Determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract a spatial attention map from the cross attention layer of the generation framework; The spatial coherence loss is determined based on the spatial attention map and the training layout annotation; the reconstruction loss is determined based on the noise prediction result corresponding to the intermediate image representation; and the stochastic gradient regularization term is determined based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The total loss is determined based on the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term. The low-rank adjustment matrix is updated based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, thus obtaining the updated low-rank adjustment matrix. Using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation.
[0130] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0131] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0132] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0133] It should be noted that any software tools or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0134] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An image generation method based on low-rank parameter adaptation, characterized in that, Includes the following steps: A generative framework is constructed, which includes a text encoder and an image decoder, and a low-rank adjustment matrix is embedded at the attention projection position of the generative framework; Obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations; Determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract a spatial attention map from the cross attention layer of the generation framework; The spatial coherence loss is determined based on the spatial attention map and the training layout annotation; the reconstruction loss is determined based on the noise prediction result corresponding to the intermediate image representation; and the stochastic gradient regularization term is determined based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The total loss is determined based on the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term. The low-rank adjustment matrix is updated based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, thus obtaining the updated low-rank adjustment matrix. Using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation.
2. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, Constructing a generative framework, the generative framework comprising a text encoder and an image decoder, and embedding a low-rank adjustment matrix at the attention projection position of the generative framework, including: Establish the connection mapping relationship between the output of the text encoder and the decoding paths of each level in the image decoder to form the generation frame connection table; Based on the generated framework connection table, the target decoding path that receives the output of the text encoder in the image decoder is determined, and multiple attention projection positions are located in the self-attention calculation link of each target decoding path; Assign a location identifier and a hierarchy identifier to each attention projection location, group the multiple attention projection locations hierarchically according to the hierarchy identifier, and configure a rank control index for each hierarchical group; At each attention projection location, the original self-attention projection calculation path is retained and connected in parallel to a low-rank adjustment matrix. The low-rank adjustment matrix includes a first low-rank matrix and a second low-rank matrix connected in sequence. The first low-rank matrix receives the input features of the corresponding attention projection location, and the second low-rank matrix receives the output features of the first low-rank matrix. The connection dimension between the first low-rank matrix and the second low-rank matrix is determined by the rank control index of the corresponding hierarchical group. Based on the location identifier, the output features of the low-rank adjustment matrix and the output features of the original self-attention projection calculation path are merged to generate an updated projection result for the corresponding attention projection position. The updated projection results at each attention projection position are integrated into the self-attention computation link corresponding to the generative framework to obtain a generative framework that embeds a low-rank adjustment matrix at the attention projection position.
3. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, Obtaining training text descriptions and training layout annotations, converting the training text descriptions into text semantic vectors, and converting the training layout annotations into layout constraint representations, including: Obtain training text descriptions and training layout annotations, and split and parse the training text descriptions to obtain training text tag sequences; Extract object words, location words, and relation words from the training text tag sequence to form a semantic mapping index table; The training layout annotations are parsed to obtain object categories, object target regions, and spatial correspondences between objects, forming layout annotation parsing results; The semantic mapping index table is associated and matched with the layout annotation parsing results to form a correspondence table between object categories and spatial regions; The training text description is input into the text encoder. Based on the correspondence table between object categories and spatial regions, the marked features output by the text encoder are aggregated to generate a text semantic vector. Based on the correspondence table between object categories and spatial regions and the layout annotation parsing results, a layout constraint representation that matches the image spatial resolution is generated.
4. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, Determine the current diffusion denoising state, generate an intermediate image representation based on the text semantic vector, the layout constraint representation, and the current diffusion denoising state, and extract a spatial attention map from the cross-attention layer of the generation framework, including: Determine the corresponding denoising feature representation based on the current diffusion round, and combine the current diffusion round with the denoising feature representation to form the current diffusion denoising state; The text semantic vector is distributed to the text condition interface corresponding to multiple cross-attention layers in the generation framework, and the layout constraint representation is mapped to the spatial constraint interface corresponding to multiple cross-attention layers. The text semantic vector, the layout constraint representation, and the current diffusion denoising state are input into the generation framework, so that the image decoder in the generation framework generates hierarchical feature maps sequentially along multiple decoding levels. Perform stepwise upsampling and cross-layer feature fusion on the feature maps at each level to generate intermediate image representations; During the generation of intermediate image representations, attention response results corresponding to the current diffusion round are extracted from multiple cross-attention layers of the generation framework; Based on the object category in the layout constraint representation, text tag attention responses corresponding to the object category are collected from the attention response results, and inter-layer fusion is performed on the text tag attention responses output by different cross-attention layers to generate a spatial attention map.
5. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, Spatial coherence loss is determined based on the spatial attention map and the training layout annotations; reconstruction loss is determined based on the noise prediction results corresponding to the intermediate image representations; and stochastic gradient regularization is determined based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix, including: Based on the object categories in the training layout annotations, object-level spatial attention distributions corresponding to each object category are separated from the spatial attention map, and the object target regions in the training layout annotations are converted into target layout distributions with the same spatial resolution as the spatial attention map. The spatial attention distribution of each object and the layout distribution of each target are normalized along the spatial dimension, and the category weight coefficient is determined based on the area ratio and region overlap of each object target region. Based on the weight coefficients of each category, the alignment error between the normalized object-level spatial attention distribution and the target layout distribution is weighted and converged to obtain the spatial coherence loss. Extract the real noise component from the current diffusion denoising state, and determine the reconstruction loss based on the difference between the noise prediction result corresponding to the intermediate image characterization and the real noise component; The original parameter gradient of the low-rank adjustment matrix is calculated based on the reconstruction loss, and a perturbation vector is injected into the original parameter gradient to obtain the perturbation parameter gradient; The parameters of the low-rank adjustment matrix are divided into multiple parameter subsets. The gradient difference norm between the original parameter gradient and the perturbation parameter gradient corresponding to each parameter subset is calculated. The gradient difference norms of each parameter subset are aggregated to obtain the stochastic gradient regularization term.
6. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, The total loss is determined based on the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term. The low-rank adjustment matrix is then updated based on this total loss, while maintaining the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge. This results in the updated low-rank adjustment matrix, which includes: The low-rank adjustment matrix is divided into multiple parameter groups according to different attention projection positions; Based on the alignment error between the spatial attention map extracted from the path where the attention projection position of each parameter group is located and the training layout annotation, the layout sensitivity coefficient corresponding to each parameter group is determined. Based on each layout sensitivity coefficient, the fusion strength of the spatial coherence loss, the reconstruction loss, and the stochastic gradient regularization term on each parameter group is adjusted to obtain the group loss corresponding to each parameter group. Each group loss is then determined as the in-group component of the total loss on the corresponding parameter group, and the total loss is obtained by aggregating the group losses. The corresponding parameter groups are updated according to the in-group components of the total loss in each parameter group, and the update transmission of the in-group components of the total loss in each parameter group to the pre-trained weights in the generation framework is blocked. Only the update transmission of the in-group components of the total loss in each parameter group to the low-rank adjustment matrix is retained. Reconstruct the low-rank adjustment matrix based on the updated parameter sets; When the parameter changes of the reconstructed low-rank adjustment matrix satisfy the convergence condition in multiple consecutive training rounds, the reconstructed low-rank adjustment matrix is used as the updated low-rank adjustment matrix.
7. The image generation method based on low-rank parameter adaptation as described in claim 1, characterized in that, Using the updated low-rank adjustment matrix and the generation framework, a target image is generated based on the target text semantic vector and the target layout constraint representation, including: The updated low-rank adjustment matrix is loaded into the attention projection position corresponding to the low-rank adjustment matrix in the generation framework; The target text semantic vector is distributed to the text condition interfaces corresponding to multiple cross-attention layers in the generation framework, and the target layout constraint representation is mapped to the spatial constraint interfaces corresponding to multiple cross-attention layers in the generation framework. The initial diffusion denoising state is determined based on the target layout constraints. The target text semantic vector, the target layout constraint representation, and the initial diffusion denoising state are input into the generation framework. Under the condition that the updated low-rank adjustment matrix participates in the adjustment, the image decoder in the generation framework performs round-by-round denoising and multi-level upsampling along multiple decoding levels to generate the target image representation. Based on the target layout constraint representation, the consistency of the object region position, object region range, and relative layout between object regions in the target image representation is checked, and the target image representation that passes the check is determined. The verified target image representation is input into the output of the generation framework for image reconstruction to obtain the target image.
8. An image generation apparatus based on low-rank parameter adaptation, characterized in that, The image generation device based on low-rank parameter adaptation includes: A generative framework building module is used to construct a generative framework, which includes a text encoder and an image decoder, and embeds a low-rank adjustment matrix at the attention projection position of the generative framework; The training input processing module is used to obtain training text descriptions and training layout annotations, convert the training text descriptions into text semantic vectors, and convert the training layout annotations into layout constraint representations. The intermediate representation generation module is used to determine the current diffusion denoising state, generate intermediate image representations based on the text semantic vector, the layout constraint representation and the current diffusion denoising state, and extract spatial attention maps from the cross attention layer of the generation framework. The triple loss calculation module is used to determine the spatial coherence loss based on the spatial attention map and the training layout annotation, determine the reconstruction loss based on the noise prediction result corresponding to the intermediate image representation, and determine the stochastic gradient regularization term based on the parameter gradient of the reconstruction loss relative to the low-rank adjustment matrix. The low-rank adjustment and update module is used to determine the total loss based on the spatial coherence loss, the reconstruction loss and the stochastic gradient regularization term, and update the low-rank adjustment matrix based on the total loss, while keeping the pre-trained weights in the generation framework unchanged, until the parameters of the low-rank adjustment matrix converge, and obtain the updated low-rank adjustment matrix. The target image generation module is used to generate a target image based on the target text semantic vector and the target layout constraint representation by utilizing the updated low-rank adjustment matrix and the generation framework.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and an image generation program based on low-rank parameter adaptation stored in the memory and executable on the processor. When executed by the processor, the image generation program based on low-rank parameter adaptation implements the steps of the image generation method based on low-rank parameter adaptation as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores an image generation program based on low-rank parameter adaptation, which, when executed by a processor, implements the steps of the image generation method based on low-rank parameter adaptation as described in any one of claims 1-7.