Vision generation method and device based on semantic association modeling, equipment and medium

By constructing a dialogue data set guided by interactive requirements and a low-rank adaptation matrix fine-tuning pre-trained language model, the method of generating visual content solves the consistency and personalized adaptation problems of visual content generation in the prior art, and realizes efficient personalized visual content generation in the fields of financial technology, medical health and poster design.

CN120542428APending Publication Date: 2025-08-26SHENZHEN PINGAN COMM TECH CO LTD
View PDF 0 Cites 25 Cited by

Patent Information

Application Number
CN202510620689.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The existing visual content generation methods are difficult to meet the structured semantic expression, personalized style control and spatial layout optimization, and cannot accurately understand user input and generate consistent visual content. Especially in the fields of financial technology, medical health and poster design, there are problems such as loose content generation, misalignment of elements, and style fragmentation.

Method used

Construct an interactive requirements-guided dialogue dataset, generate perturbation training samples through semantic perturbation operations, fine-tune the pre-trained language model for low-rank adaptation matrix, extract element semantic features and determine association weights, generate layout optimization functions in combination with spatial distribution constraints, and iteratively generate target visual content.

Benefits of technology

It realizes structured response and spatial mapping to user semantic needs, improves the expression consistency and personalized adaptability of visual content generation, and is suitable for business scenarios such as financial technology, medical health and poster design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120542428A_ABST
    Figure CN120542428A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health, poster design and the like, and discloses a visual sense generation method and device based on semantic association modeling, equipment and a medium. Generating a demand text containing theme and style parameters; semantic features in the demand text are extracted, semantic association weights are constructed, and element layout coordinates are optimized in combination with spatial distribution constraints; and encoding the layout information into a control matrix, fusing the control matrix with the initial noise, adjusting a noise reduction process through an encoding and decoding network, and generating target visual content highly matched with the semantic meaning of the user instruction. According to the method, the layout optimization function is constructed, the diffusion model is guided to focus the semantic salient region in space, language model output and the visual generation process are closely combined, structured response and space mapping of user semantic requirements are achieved, and the expression consistency and personalized adaptation capacity of visual content generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech semantics technology, and in particular to a visual generation method, device, equipment and storage medium based on semantic association modeling. Background Art

[0002] With the development of multimodal intelligent generation technology, the role of visual content in assisting information expression, enhancing user perception, and supporting decision-making communication is becoming increasingly significant. However, existing visual content generation methods still have many technical bottlenecks and are unable to meet multiple requirements such as structured semantic expression, personalized style control, and spatial layout optimization. On the one hand, the natural language, structured parameters, or semantic labels input by users are often not understood by the system with high quality and accurately mapped into visual expression content; on the other hand, existing models lack the ability to model the logical relationship between semantic elements during the generation process, and the generated content often has problems such as loose layout, misplaced elements, and style fragmentation. In addition, after the user proposes modification suggestions, the current system lacks an effective parameter adjustment and style maintenance mechanism, making it difficult to retain the original semantic structure and visual consistency while optimizing the expression of user preferences.

[0003] In the fintech sector, companies often need to generate visual content for marketing, compliance demonstrations, or risk management explanations based on information such as customer characteristics, product structure, and risk preferences. These visual content, such as investment advisory product maps, profit path diagrams, and financial plan cards, typically requires precise semantic expression, a standardized and consistent style, high information density, and clear categorization logic. Existing methods commonly suffer from loose content structure, lack of focus, and inconsistent formatting with compliance standards, severely impacting their practicality in promoting financial products, educating clients, and providing professional explanations.

[0004] In the healthcare sector, systems must generate clear, professional visual aids based on inputs such as diagnosis and treatment processes, patient labels, medical service packages, or popular science topics. These include personalized health reminders, medical knowledge diagrams, and interactive consultation results. This type of content often contains a large number of medical terms and professional concepts, requiring the model to possess strong contextual understanding and controllable visual expression when processing semantic content. However, current visual generation systems generally lack modeling support for professional language structures, making it difficult to achieve accurate information mapping, clear visual hierarchy, and controllable visual risks, limiting their application effectiveness in real-world medical interaction scenarios.

[0005] In the field of poster design and visual communication, posters, as the core carrier for rapid information transmission and visual appeal, play a key role in scenarios such as commercial promotion, cultural communication, and event notifications. Traditional poster creation relies on manual completion of processes such as semantic interpretation, creative composition, and image and text arrangement. The cycle is long, the efficiency is low, and the style expression is highly dependent on the designer's experience. Although some automated poster generation methods use template matching or style transfer mechanisms to improve generation efficiency, they are often difficult to adapt to the requirements of multiple themes, personal expression, and structural differences. Especially when faced with complex semantic instructions and multi-level design elements, existing solutions find it difficult to achieve the unity of high-precision semantic control, layout logic optimization, and personalized style maintenance. Summary of the Invention

[0006] The main purpose of the present invention is to provide a visual generation method, device, equipment and storage medium based on semantic association modeling, aiming to solve the technical problem that the existing technology cannot generate structured design parameters based on real-time interactive semantics, and integrate semantic association weights and spatial constraints to accurately drive the generation model to complete the construction of personalized visual content.

[0007] To achieve the above objectives, the present invention provides a visual generation method based on semantic association modeling, comprising:

[0008] Constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing semantic perturbation operations on the original interaction instructions to generate perturbation training samples;

[0009] Inputting the perturbed training sample into a pre-trained language model, and fine-tuning the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model;

[0010] receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters;

[0011] Extracting semantic features of elements in the requirement text, and determining association weights between the semantic features of each element as semantic association weights;

[0012] Combining the semantic association weight with a preset spatial distribution constraint term to construct a layout optimization function, and optimizing the layout optimization function through backpropagation to generate element coordinate data;

[0013] Encoding the element coordinate data into a spatial position vector, and fusing the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0014] The layout control matrix and the initial noise tensor are input into a codec network, and the layout control matrix and the initial noise tensor are fused through the codec network to generate a regional focus weight matrix. The noise reduction process of the codec network is adjusted based on the regional focus weight matrix to iteratively generate target visual content.

[0015] Furthermore, to achieve the above-mentioned purpose, the present invention provides a visual generation device based on semantic association modeling, comprising:

[0016] An interaction data construction module is used to construct a dialogue dataset containing interaction demand guidance samples, extract original interaction instructions from the dialogue dataset, and perform semantic perturbation operations on the original interaction instructions to generate perturbation training samples;

[0017] A language model fine-tuning module is used to input the perturbation training sample into a pre-trained language model and fine-tune the parameters of the pre-trained language model through a low-rank adaptation matrix to generate a fine-tuned language model;

[0018] A demand text generation module, configured to receive a real-time interaction instruction, input the real-time interaction instruction into the fine-tuned language model, and generate a demand text including a topic identification field and style level parameters;

[0019] A semantic association modeling module is used to extract semantic features of elements in the requirement text and determine the association weights between the semantic features of each element as semantic association weights;

[0020] a layout optimization analysis module, configured to combine the semantic association weights with preset spatial distribution constraints to construct a layout optimization function, and generate element coordinate data by optimizing the layout optimization function through backpropagation;

[0021] a layout control fusion module, configured to encode the element coordinate data into a spatial position vector, and fuse the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0022] The visual content generation module is used to input the layout control matrix and the initial noise tensor into the codec network, fuse the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjust the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generate target visual content.

[0023] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a visual generation program based on semantic association modeling stored in the memory and runnable on the processor. When the visual generation program based on semantic association modeling is executed by the processor, the steps of the visual generation method based on semantic association modeling as described above are implemented.

[0024] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a visual generation program based on semantic association modeling is stored. When the visual generation program based on semantic association modeling is executed by a processor, the steps of the visual generation method based on semantic association modeling as described above are implemented.

[0025] Beneficial effects: The present invention relates to the field of speech semantic technology and can be applied to business scenarios such as financial technology, medical health, and poster design. It discloses a visual generation method based on semantic association modeling, including: constructing a dialogue data set and generating perturbation training samples, fine-tuning the parameters of the pre-trained language model, using the fine-tuned language model to process real-time interactive instructions to generate demand text, extracting element semantic features and determining semantic association weights, constructing a layout optimization function to generate element coordinate data, encoding the coordinate data into a spatial position vector and fusing it with the semantic features to generate a layout control matrix, inputting the layout control matrix and the initial noise tensor into the encoding and decoding network, generating a regional focus weight matrix, and adjusting the noise reduction process to iteratively generate visual content. The present invention constructs a layout optimization function and guides the diffusion model to focus on semantically significant areas in space, closely combining the language model output with the visual generation process, thereby achieving structured response and spatial mapping to user semantic needs, and improving the expression consistency and personalized adaptation capabilities of visual content generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0027] Figure 1 A schematic diagram of an application environment of a visual generation method based on semantic association modeling in an embodiment of the present invention;

[0028] Figure 2 Schematic diagram of a flow chart of an embodiment of a visual generation method based on semantic association modeling according to the present invention;

[0029] Figure 3 Schematic diagram of functional modules of a preferred embodiment of a visual generation device based on semantic association modeling of the present invention;

[0030] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0031] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0032] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0033] The visual generation method based on semantic association modeling provided by the embodiment of the present invention can be applied in the following fields: Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can construct a dialogue data set and generate perturbation training samples through the user end, fine-tune the parameters of the pre-trained language model, use the fine-tuned language model to process real-time interactive instructions to generate demand text, extract element semantic features and determine semantic association weights, construct a layout optimization function to generate element coordinate data, encode the coordinate data into a spatial position vector and fuse it with the semantic features to generate a layout control matrix, input the layout control matrix and the initial noise tensor into the encoding and decoding network, generate a regional focus weight matrix, and adjust the noise reduction process to iteratively generate visual content. The present invention constructs a layout optimization function and guides the diffusion model to focus on semantically significant areas in space, closely combines the language model output with the visual generation process, realizes structured response and spatial mapping to user semantic needs, and improves the expression consistency and personalized adaptation ability of visual content generation. Among them, the user end can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server end can be implemented by an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0034] See also Figure 2 , Figure 2 This is a flowchart of an embodiment of a method for visual generation based on semantic association modeling provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0035] like Figure 2 As shown, the visual generation method based on semantic association modeling proposed in the present invention includes the following steps:

[0036] S10, constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing a semantic perturbation operation on the original interaction instructions to generate perturbation training samples;

[0037] In this embodiment, the process of constructing a conversation dataset containing interactive demand guidance samples aims to provide a stable, clearly structured, and semantically diverse language input resource for the multimodal generative model, thereby improving the model's response accuracy and scenario generalization capabilities to user natural language requests. "Interaction demand guidance samples" refer to user sentence examples with clear visual instruction intent and labeled with dimensions such as specific scene elements, style preferences, and layout requirements. Such samples should not only be parseable but also reflect task-driven attributes. Constructing the conversation dataset involves integrating multiple data sources and performing manual or semi-automatic annotation. The data sources can come from historical records of actual online design dialogue platforms, questions and answers collected from user surveys, and synthesized virtual instruction dialogues. The conversation content must cover a wide range of topic categories, style options, scenario settings, and other dimensions to meet the model's generalization requirements for different tasks.

[0038] Extracting original interaction commands from the constructed guided sample dataset involves extracting the natural language command body through a data parsing mechanism, eliminating redundant background descriptions, annotations, or the influence of multiple rounds of context, and retaining only the direct request statements expressed by the user in a single round. For example, the original interaction command extracted from a query like "I need a promotional image suitable for summer children's activities, with a playful theme and refreshing colors" should be "I need a promotional image suitable for summer children's activities" to reduce the impact of interfering information on language model training.

[0039] Semantic perturbations are performed on the extracted original interaction commands to improve the language model's adaptability to semantic variation and expression diversity, thereby enhancing the model's robustness to real user commands. Semantic perturbations include two core methods: synonym replacement and fuzzy generalization. Synonymous replacement involves replacing keywords or expressions in the original command with other words with similar meanings while maintaining semantic similarity. This process can be implemented based on a structured synonym library or contextual embedding model. For example, "blue technological background" can be replaced with "cold-toned smart background" to create a different semantic expression path. Fuzzy generalization generalizes precise descriptive terms, including relativizing time descriptions (e.g., replacing "May 2024" with "recent") and linguistically defining spatial coordinates (e.g., replacing "upper left corner" with "upper area"). These perturbations ensure that the target semantics do not drift from their core meaning, but introduce more expressions and fuzzy understanding space to simulate the uncertainty and non-standard expressions that may occur in real user interactions.

[0040] These training samples generated after the perturbation operation are used as the input of the language model to fine-tune a pre-trained language model that already has basic language understanding capabilities. The introduction of a low-rank adaptation matrix in the fine-tuning process refers to the use of efficient parameter adjustment mechanisms such as LoRA (Low-Rank Adaptation) to insert small-scale trainable parameter modules into the original model structure to avoid stability problems and overfitting risks caused by the overall model parameter update. The low-rank matrix is ​​added between specific layers of the model, and the language representation space is fine-grainedly adjusted through multiplicative or additive adjustments, so that the model retains the original expressive ability while learning the generalized response structure under the semantic perturbation of interactive instructions. After training is completed, the resulting model has stronger interactive understanding, adaptability to fuzzy and variant expressions, and can accurately extract task themes and target control parameters in generation tasks.

[0041] Real user demand expressions collected from conversational scenarios can be used as the basic data source. Manual labeling systems can be used to annotate visual instruction types, target topics, style preferences, and other fields to construct a structured set of interaction guidance samples. Synthetic data can also be generated by expanding rule templates or language models to generate diverse expression instructions as supplementary samples. The semantic perturbation process can use the WordNet synonym set, design industry vocabulary, and word similarity analysis based on BERT word embeddings to screen replacement options. A rule-based expression generalization engine is also used to implement temporal and spatial fuzzy expression replacement. During fine-tuning, LoRA is used to perform interpolation training on the Transformer-like pre-trained language model, limiting the trainable parameters to a matrix form of rank 4 or 8, achieving efficient and stable model semantic transfer and adaptation. Training data uses perturbation instructions as input and the parsed standard task target structure as output for aligned training.

[0042] Furthermore, during interactive instruction perturbations, in order to improve the diversity of training samples and the model's robustness to non-standard inputs, in addition to utilizing explicit perturbations such as synonym replacement and temporal / spatial fuzzy generalization, a gradient-guided adversarial sample enhancement method can be introduced. This method simulates the perturbation direction to which the model is most sensitive, introducing weak but aggressive perturbations to the original input to further construct challenging training samples.

[0043] Specifically, let x represent the original input text sample, and generate the adversarial perturbation sample by the following formula

[0044]

[0045] Where: ∈ is the perturbation intensity factor, which is used to control the semantic acceptability of the perturbation; f θ(x) represents the predicted output of the current language model for input x, with parameter θ; y is the true semantic label or target output corresponding to the input; L(·) represents the loss function (such as cross entropy loss); It represents the gradient of the loss function with respect to the input x, which is used to measure the sensitive direction of the input under the current model; sign(·) is the sign function that normalizes the gradient direction to +1 or -1.

[0046] Example: In a healthcare data application scenario, a user might issue a command such as "Design a promotional graphic to remind seniors to have regular physical examinations." This command might be expressed as "Help me make a graphic with physical examination reminders suitable for seniors." The difference between the original and perturbed commands lies in the ambiguity and variation in semantic expression. The model needs to identify keywords such as "physical examination," "senior citizens," and "reminder" and construct a standardized semantic structure.

[0047] In financial scenarios, a command like "Create a diagram showing the loan application process" might be followed by an expression like "Draw a flowchart showing how to apply for a loan." The model must accurately understand the logical relationship between "process" and "loan." In these scenarios, building high-quality guidance data and introducing a semantic perturbation training mechanism enables language models to handle unstructured expressions, ensuring accurate and professional content generation. This offers significant advantages when navigating input from non-expert users.

[0048] By interactively guiding the construction of samples and the generation of semantic perturbations, the language model's ability to understand user instructions can be significantly expanded, so that task objectives can still be accurately identified when faced with ambiguous, variant or non-standard expressions.

[0049] S20, inputting the perturbed training sample into a pre-trained language model, and fine-tuning the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model;

[0050] In this embodiment, inputting perturbed training samples into a pre-trained language model further guides the model to learn the semantic recognition and command comprehension capabilities required for specific domains or tasks, building upon the existing large-scale language knowledge structure. Perturbed training samples serve two key purposes: first, they provide expression variations of the original user interaction commands, simulating the diversity, ambiguity, and stylistic differences found in real conversations; second, they help the language model robustly map these variations, enabling the model to generate structured outputs with greater consistency and generalization capabilities.

[0051] Pretrained language models typically refer to deep neural network models that have been trained on large-scale general-purpose corpora, such as BERT, T5, and GPT. These models possess powerful language modeling and contextual understanding capabilities, but they may not be precisely adaptable to specific instruction styles or domain tasks. Therefore, they require further targeted fine-tuning to enhance their task-specific performance. While directly fine-tuning the entire model parameters on a large scale can be effective, it often faces challenges such as high computational resources, a high risk of overfitting, and difficulty migrating.

[0052] To address the above issues, a low-rank adaptation matrix is ​​introduced to fine-tune the parameters of the pre-trained language model. This process is essentially a structural parameter insertion training strategy. Representative methods include LoRA (Low-Rank Adaptation). The core idea is to insert two sets of low-rank matrices between specific weight layers of the original model to construct an expression space that is similar to the fully connected parameter update. The dimensions of these low-rank matrices are much smaller than the number of parameters in the original model. During training, only the trainable parts of the inserted matrices are updated, and the other original model parameters remain frozen. In this way, fast targeted adaptation can be achieved at a very low computational cost without destroying the original language modeling capabilities. It is especially suitable for generative model systems that are sensitive to resource constraints or have strong multi-task deployment requirements.

[0053] After fine-tuning, a diversified mapping pathway from perturbation expressions to standard semantic structures is established within the language model. When faced with real interactive input, it can identify the core elements of user intent and generate semantic output that matches the structured visual control parameters, laying a solid foundation for subsequent feature extraction, layout optimization, and image generation.

[0054] Pre-trained language models with an encoder-decoder architecture, such as T5 and BART, can be used as the basic framework to ensure a balance between generation and comprehension capabilities. Perturbed training samples serve as model input, and the output is a pre-structured visual requirement description format, such as a set of key-value pairs containing keywords, style fields, and dimension annotations. Fine-tuning utilizes the LoRa framework, inserting low-rank matrices into the query and value transformation paths of the attention module in the Transformer layer. Matrix dimensions are chosen to be rank 4 or 8, adjusted based on task complexity. During training, all weights of the pre-trained model are frozen, and only the parameters of the LoRA insertion layer are optimized, enabling lightweight and efficient semantic instruction adaptation. Training data is constructed using supervised learning, with the goal of ensuring that the perturbation instructions converge semantically to a standardized expression. Fine-tuning performance can be evaluated using metrics such as BLEU, ROUGE, and structural consistency accuracy to ensure output controllability and clarity.

[0055] For example, after constructing perturbation training samples and inputting them into the language model, training optimization based on conversational data is required to achieve the transfer of semantic expression and command comprehension capabilities. During fine-tuning, the standard objective function of autoregressive language modeling is used to drive the model to accurately predict the next word and improve the consistency and logic of text generation. This objective loss function can be defined as follows:

[0056]

[0057] Where: x represents the input sequence, which may be the interaction instruction after perturbation; y t is the tth word in the target output sequence; y <t represents all the words that the model has generated before the t-th round of prediction; P(y t |y <t ,x) is given input x and context y <t Under the condition of , the model predicts the next word is y t The probability of ; T is the total length of the target sentence.

[0058] This objective function, based on maximum likelihood estimation, guides the model to gradually learn semantic structure and contextual consistency by minimizing the negative log-likelihood loss. During training, this loss function works together with the Low-Rank Adaptation (LoRA) mechanism to fine-tune the model only within a specific weight subspace, improving parameter efficiency and reducing computational resource consumption.

[0059] Example: In healthcare applications, users may express similar needs by typing "make a notification to remind the elderly to avoid heatstroke when going out in the summer" or "generate a poster with instructions for preventing heatstroke in the summer for the elderly." The fine-tuned language model must be able to identify the shared semantic cores between these different expressions, such as "elderly," "heatstroke reminder," and "summer," and output the demand text in a unified format, including keywords such as "elderly health" and "high temperature protection," style preferences such as "refreshing, warm," and attributes such as the image-to-text ratio.

[0060] In the financial field, instructions such as "Give me a diagram illustrating the loan process" or "Make a diagram to show my parents how to apply for a loan" have similar semantic structures, although the styles are quite different. The fine-tuned language model can recognize keywords such as "loan application" and "operation steps" and output clear structured demand text as the input basis for the visual content generation process.

[0061] By combining the input of perturbed training samples with low-rank parameter fine-tuning of the language model, we achieve a significant improvement in its ability to understand and adapt to diverse user input while retaining the original language capabilities of the pre-trained language model. The low-rank adaptation mechanism avoids the overfitting and generalization failure issues that can arise from full parameter fine-tuning, significantly reduces training resource overhead, and enables multi-task and multi-scenario deployment. Furthermore, this mechanism provides semantically clear, expressive, and structurally unified instruction input for subsequent modules such as structured semantic extraction and layout control generation, enhancing the stability and controllability of the multimodal interaction process.

[0062] S30, receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters;

[0063] In this embodiment, receiving real-time interactive instructions refers to dynamically obtaining user input expressed in natural language while the system is running. These inputs can come from text input boxes, speech-to-text modules, API calls, and other methods. These interactive instructions are unstructured and semantically variable, and often contain a mixture of multi-dimensional information such as task objectives, style preferences, and presentation requirements. The system must be able to effectively parse and extract the structure of these naturally expressed instructions in order to provide reliable control parameters for subsequent generation processes.

[0064] Real-time interaction commands are fed into the fine-tuned language model, indicating that the model is now capable of extracting and generating specific structured expressions based on user intent. The model's core task is no longer open-ended natural language generation, but rather is geared towards generating specific visual content, performing semantic understanding, key element extraction, parameter classification, and structural reconstruction on the input language.

[0065] Generating a demand text containing a subject identification field and style level parameters means that the output of the model is no longer a natural language text, but a structured information set. The subject identification field is the core semantic representation extracted from the user interaction instructions, such as "energy-saving lamp promotion", "middle-aged and elderly health care tips", "intelligent investment advisory platform introduction", etc. These fields can be keyword groups, concept labels or descriptive phrases, which are used to clarify the main direction of the visual content. The style level parameters are the quantitative results of the user's subjective preference or suggestive expression of the visual style, usually including color tones (such as "warm colors" and "minimalist"), typesetting rhythm (such as "dense layout" and "partitions and columns"), and the proportion of pictures and texts, element hierarchy and other dimensions, which are used to guide the style expression of subsequent visual elements in the layout and graphics generation stages.

[0066] This requirement text is structured as a key-value pair, semantically normalized and formatted to ensure accurate parsing and use by subsequent modules. For example, the "Theme Identifier Field" key might contain values ​​like ["Fintech," "Robo-Advisor"], while the "Style Level Parameters" key might contain values ​​like {"Main Color": "Blue and Gold," "Layout": "Symmetrical," "Font Style": "Formal and Clear"}. This format transforms subjective and ambiguous expressions into machine-controllable generation parameters, serving as a critical intermediate semantic bridge in building an end-to-end closed-loop generation loop.

[0067] For example, in the field of medical health, the interactive end can input the real-time command "Please design a soft-style promotional picture for the elderly health lecture." The language model will extract "elderly health lecture" as the theme identification field, identify "soft style" as part of the style hierarchy parameters, and generate layout prompts with the main color of "blue and white" and the font type of "rounded sans serif", providing a parameter basis for the subsequent generation of gentle and friendly visual content that conforms to the emotions of medical scenarios.

[0068] In the field of financial services, users can use voice or text input commands such as "I want a cover image with a strong sense of stability for corporate annual reports." The model will automatically parse "corporate annual report" as the core theme, and convert "strong sense of stability" into an attribute description in the style hierarchy. It will further expand the generation of recommended values ​​for font thickness, background pattern texture, and color saturation, making the generated image both authoritative and aesthetically consistent.

[0069] In the poster creation scenario, the user inputs "Please design a minimalist style poster for the spring new product launch". The language model will extract "Spring New Product Launch" as the theme identification field, and at the same time identify "minimalist style" as the control parameter in the style hierarchy, and then generate content style suggestions characterized by "white background + line drawing + single point highlight", driving the visual content generation process to strictly adhere to the minimalist design principles and meet brand communication needs.

[0070] Through real-time interactive instructions, the system parses and generates structured demand text containing subject identification fields and style-level parameters, enabling users to express themselves efficiently through natural language without having to master specialized graphic design terminology. The fine-tuned language model is capable of transforming fuzzy, subjective expressions into controllable generation parameters, significantly improving the system's ability to understand and respond to diverse user needs. This mechanism provides stable, clear, and uniformly formatted input control conditions for subsequent modules, serving as a critical bridge supporting the automatic generation of personalized visual content.

[0071] S40, extracting semantic features of elements in the requirement text, and determining association weights between the semantic features of the elements as semantic association weights;

[0072] In this embodiment, after the system receives text content processed by language modeling, it needs to extract key expression units from the text and establish semantic connections between these expression units to support the subsequent control of visual generation logic. The core of this process is to transform abstract language information into a representation with computable properties and further construct a relevance graph that reflects the semantic structure.

[0073] The expression units in the text can be divided into two categories. One category is phrases that directly describe the content theme, such as "brand release", "medical consultation", "investment products", etc. This type of information reflects the main semantic direction of the image content; the other category is modifiers or adjective phrases that describe the appearance, style or context, such as "minimalist", "warm", "stable", "futuristic", etc. These contents are usually used to propose preference constraints on non-content dimensions such as layout, color matching, and graphic structure.

[0074] When processing these expression units, it is necessary to construct an embedded vector representation for each unit in the semantic space. This vector must capture its semantic properties, contextual dependencies, and directional relationships. This can be achieved by using semantic models pre-trained on large-scale text (such as the Transformer series of models) to generate context-aware vectors. Alternatively, nested fine-tuning can be performed within specific domains to enhance the differentiation of professional expressions. For example, the word "stable" may have different meanings in finance and healthcare, and the vector representation should reflect these subtle differences in domain context.

[0075] To construct connections between vectors, a similarity calculation method based on vector angles, such as cosine similarity, can be introduced to measure whether two expression units are semantically close. The result of the calculation is a symmetric matrix, where each value corresponds to the semantic closeness between a pair of expression units. To make this information usable for subsequent control strategies, the original similarity matrix needs to be normalized to prevent semantic evaluation bias caused by different scales. Normalization can be achieved through Min-Max scaling or Softmax transformation, so that all values ​​are within a uniform scale range.

[0076] Furthermore, the constructed connection weight matrix may contain a lot of noise or low-value associations, so it is necessary to introduce a sparsification mechanism, that is, to set a minimum semantic connection strength threshold and remove items below the threshold. In this way, word pair information with clear semantic orientation and high correlation can be retained. For example, when a user expression involves two units, "diabetic diet management" and "customized exercise plan", the semantic distance between them is relatively close and should be retained with a high weight, while "diabetic diet management" and "image compression ratio" can be eliminated due to the lack of logical commonality.

[0077] In practice, a pre-trained semantic model such as BERT or RoBERTa can be used as a base extractor. The structured requirement text input is expanded using key-value pairs, encoding fields such as "title," "primary visual elements," and "background mood." The text corresponding to each field is fed into the model to extract a fixed-length semantic vector representation.

[0078] Subsequently, the cosine similarity between any two semantic vectors is calculated, generating a symmetric semantic similarity matrix. Each entry in the matrix represents the degree of correlation between two semantic elements. A Softmax function can be used to normalize the similarity, weighting it between 0 and 1 to enhance comparability. A weight threshold (e.g., 0.2) is set, and entries below this threshold are set to 0. This yields a sparse semantic association graph structure that can be used to drive spatial layout optimization.

[0079] This process can be deployed in the background semantic analysis module to run in real time, or a weight set can be pre-generated in the model preprocessing stage for control guidance of batch image and text tasks.

[0080] By extracting semantic features and calculating semantic association weights, the system transforms unstructured user textual intent into a structured semantic relationship graph. This enables the system to control element layout and content matching during visual generation in a semantically driven manner, achieving a fusion of semantic and spatial consistency. This improves the quality of automatically generated visual logic while ensuring accurate thematic expression, significantly enhancing the stylistic coherence and clarity of the generated content.

[0081] S50, combining the semantic association weight with a preset spatial distribution constraint item to construct a layout optimization function, and optimizing the layout optimization function by backpropagation to generate element coordinate data;

[0082] In this example, after obtaining the semantic association weights between expression units, this semantic structure needs to be mapped into a two-dimensional space to guide the positional arrangement of elements in the subsequent image synthesis task. At this point, abstract linguistic similarity must be converted into actual coordinate parameters. The key to this conversion lies in constructing an optimization objective that ensures that highly semantically related elements are spatially adjacent while maintaining the structural constraints of the visual arrangement.

[0083] Semantic association weights are used to express the strength of the semantic coupling between any two expression units. These weights are derived from the high-weight connection terms retained from the previously generated semantic similarity matrix. When modeling space, these weights can be considered "attraction" parameters; higher values ​​indicate a stronger "desire for proximity" between elements. Based on this, a loss function can be constructed to minimize the distance in space between expression units with strong semantic relationships, thereby encouraging consistency between spatial and semantic structures.

[0084] However, in a layout based purely on semantic optimization, elements may cluster together, ignoring the overall layout rules, resulting in inefficient space usage or visual clutter. Therefore, a spatial distribution constraint mechanism needs to be introduced. This mechanism is built based on prior knowledge or historical layout models, and typically uses a Gaussian distribution function to model the degree of deviation between the spatial distance between elements and the expected layout spacing. For example, when two elements should maintain a certain distance according to the preset strategy, the layout optimization function will generate a larger penalty gradient when these two points are too close, pushing them to regress to a more reasonable spatial structure.

[0085] This optimization function can be modeled as a combined objective consisting of multiple loss terms, where the semantic attraction term is defined as the weighted sum of the squared distances between multiple element pairs multiplied by the semantic association weight, and the spatial arrangement term is defined as the weighted sum of the degree of deviation of all element pairs in space. Both can be controlled by a balance parameter to take into account both semantic aggregation and visual distribution balance.

[0086] To solve this optimization function, all expression units must be initialized with two-dimensional coordinate values ​​as a starting point and a gradient descent-based optimization process is introduced. In each iteration, the system calculates the optimization target value for the current layout, takes partial derivatives of all coordinate variables, determines the gradient direction, and updates the current element position accordingly. Because the objective function is continuously differentiable, the backpropagation mechanism enables efficient batch updates.

[0087] This entire process continues over multiple iterations until the change in the optimization objective falls below the set convergence threshold for multiple consecutive rounds, or the maximum number of iterations is reached. Finally, a set of 2D coordinates for all expression units is output. This set serves as a spatial prior for the subsequent image manipulation phase, determining the initial placement of each element and directly constraining the attention mechanism of the content generation network.

[0088] Spatial optimization is not only applicable to visual content generated from text but can also be extended to a variety of tasks, including mixed text and image typesetting, chart generation, and interactive interface design. Under complex input conditions such as multi-target input, multi-language structures, and sets of elements of variable length, flexible adaptation can be achieved simply by adjusting the initial layout strategy or adaptively adjusting the weights between semantic and spatial terms. Different optimization structures (such as neural network structures and energy propagation network structures) can also serve as the inference engine for this optimization process, further improving convergence efficiency and structural fidelity for large-scale element sets.

[0089] For example, when constructing the spatial distribution of semantic elements, it is important to consider the spatial proximity between semantically related elements. At the same time, it is also necessary to guide the spatial arrangement of the overall elements. To this end, a layout optimization function is established to optimize the arrangement of semantic elements in the image in two-dimensional coordinates. This optimization function can be shown as follows:

[0090]

[0091] This optimization objective function consists of two main parts, corresponding to semantic space constraints and layout-guided matching:

[0092] The first part is the semantic space collaboration item:

[0093]

[0094] Among them, c i ,c j Represents the semantic category encoding or semantic embedding of semantic elements i and j; Represents a function that measures the semantic similarity between two elements, which can be calculated based on cosine similarity or sentence vector similarity; p i , p j is the actual coordinate of semantic element i, j on the two-dimensional canvas; ||p i -p j || 2 Represents the square of the spatial distance between elements; σ represents the bandwidth parameter of the Gaussian kernel function, which is used to control the attenuation degree of "greater impact when the distance is close and smaller impact when the distance is far".

[0095] This item is used to encourage semantically similar elements to be closer together, forming semantic clusters in space, thereby enhancing visual logical consistency and clarity of expression.

[0096] The second part is the layout guide matching item:

[0097] λ||p i -p guide,i || 2

[0098] Where: p i Represents the current two-dimensional spatial coordinate of the i-th semantic element. This coordinate is a variable that is continuously updated during the back-propagation optimization process and represents the actual position of the semantic element in the final layout. guide,i Represents the guidance target coordinates generated by the layout guidance module (such as the layout planning agent) for the i-th semantic element. These coordinates are usually determined by preset rules, templates, or historical layout patterns, and are used to guide the arrangement of semantic elements on the canvas to be aesthetically pleasing, standardized, or semantically clear. ||p i -p guide,i ||2 Represents the squared Euclidean distance between the current semantic element coordinates and their target coordinates, used to measure the degree of deviation. The smaller this term, the closer the semantic element is to the guided target position. λ is a preset balancing factor used to adjust the weight of the guidance loss in the overall layout optimization function. Larger λ values ​​favor an exact match of the guided coordinates, while smaller values ​​allow for more adaptive layout deviations.

[0099] This formula ensures semantically coordinated arrangement of elements while guiding each element as close to its intended visual position as possible, thereby improving the overall readability and structural aesthetics of the layout. This formula can significantly improve the effectiveness of directing user attention in visual tasks involving important prompts (such as "Warning," "Price," and "Main Title").

[0100] Example: In the healthcare sector, an organization wants to generate a visual health checkup report cover. Users interactively provide semantic elements such as "home checkup," "elderly-friendly," and "gentle tones." System analysis reveals a high degree of semantic correlation between "home," "elderly," and "gentle." Using an optimization function, the visual units corresponding to these keywords are automatically arranged near the center of the image. Elements with weaker coupling to the main semantics, such as "hospital" and "testing facility," are placed at the edges, highlighting the core semantics and enhancing thematic focus.

[0101] In the financial sector, a platform wanted to generate a product guide based on user-provided keywords: "risk warning," "regular savings," and "smart recommendations." Using semantic encoding, the system discovered that "regular" and "savings" had the strongest correlation, focusing them where the eye would focus. Meanwhile, "smart recommendations" was positioned in the second highest user attention area. Gaussian distribution constraints were used to control the contrast between different elements, maintaining a clear visual hierarchy while preserving the consistent style of the financial page.

[0102] In the poster generation scenario, the user inputs "environmental protection exhibition", "low-carbon life" and "green technology". The system analyzes the high semantic coupling between "green technology" and "low-carbon life" by optimizing the objective function, and spatially clusters these two expression units in the middle of the main image area. "Environmental protection exhibition" is placed as the keyword in the title area. Through spatial guidance, the information transmission path is strengthened, the overall layout structure is coordinated, and the information focus is highlighted.

[0103] By jointly modeling semantic association weights and spatial distribution constraints, and constructing a layout optimization function to perform backpropagation optimization, we can balance semantic consistency and visual structural rationality when generating element spatial layouts. Highly semantically related expression units are adaptively moved closer together in space, reducing layout fragmentation or loss of focus. Furthermore, spatial distribution no longer relies on static templates but instead has dynamic adjustment capabilities. The entire layout process is driven by a learnable mechanism that can adapt to different input structures and expression goals, improving the structural consistency and diversity of the generated content.

[0104] S60, encoding the element coordinate data into a spatial position vector, and fusing the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0105] In this embodiment, element coordinate data is a structure containing the positioning information of each element in two-dimensional space, usually composed of horizontal and vertical coordinate values, and accompanied by a unique identifier to distinguish different elements. The spatial position vector is a vectorized encoding of the spatial position based on the coordinate data, so that it has a numerical structure suitable for neural network input. For example, the horizontal and vertical coordinate values ​​can be standardized to the interval [0, 1] and expanded into a fixed-length vector form, consistent with other semantic features in terms of dimension.

[0106] The fusion of spatial position vectors with the semantic features of elements in the requirement text aims to jointly model the visual layout relationship with the semantic relationship reflected in the text description. The semantic features of the elements themselves are encoded by the language model and contain unstructured feature information such as design theme, visual style, and emotional attributes. Fusion can be achieved through vector concatenation, weighted feature superposition, or feature mapping via a fully connected layer, creating a joint representation of the two different types of feature information in the same representation space. The key to this fusion is maintaining the correspondence between spatial and semantic information, ensuring that the subsequent visual generation stage can accurately understand "where" and "what" are presented.

[0107] After fusion is complete, the resulting structure is the layout control matrix. This matrix is ​​essentially a set of conditional vectors used to guide regional control during image generation. Each row represents a comprehensive control vector for an element, including its position, semantic attributes, and identification information. This control matrix can serve as input for subsequent visual generation stages, driving the image generation model to produce output with clear semantics and spatial structure.

[0108] One implementation involves normalizing the two-dimensional coordinate data (x, y) to the range [0, 1] and encoding it into a 128-dimensional spatial position vector through a small fully connected network. This vector is then concatenated with the element-wise semantic features output by the language model and then mapped to a shared feature space through a fusion layer (such as an MLP).

[0109] Alternatively, we can use the position encoding technique in the multi-head attention mechanism to embed coordinate values ​​directly into the vector representation of semantic features to achieve spatial semantic alignment. Another approach is to use a two-stream structure to encode position and semantic information separately and integrate them using a cross-attention mechanism in the fusion stage.

[0110] When generating the layout control matrix, it can be uniformly output as a tensor of shape N × D, where N is the number of elements and D is the dimension of the fused features. The control matrix retains the element sequence to ensure accurate mapping between position and semantics.

[0111] In some computing resource-sensitive scenarios, low-rank projection can also be used to reduce the spatial vector dimension, thereby maintaining overall performance stability and improving model processing efficiency.

[0112] Example: In the healthcare field, doctors can use natural language descriptions to generate patient physiological structure diagrams or data report layouts. For example, "a lung scan is displayed in the upper left corner, and the CT numerical trend is marked on the right." At this time, the spatial position vector is integrated with semantic labels such as "lung" and "CT trend" into a control matrix, which is used to guide the medical image generation model to present specified content in a specified area.

[0113] In the financial field, users can input "the upper part shows the earnings trend chart of this quarter, and the lower part lists the risk score summary." The model can extract "earnings trend chart" and "risk score" as semantic elements, and fuse them with the coordinate vectors of "upper" and "lower" to generate a control matrix, driving the report generation interface to accurately arrange key indicators.

[0114] In the field of visual content design, for example, in poster design, when the user inputs "place the brand logo in the center and display product information at the bottom", the system can fuse keywords such as "brand logo" and "product information" with the vectors encoded with the center / bottom coordinates to generate a control matrix to guide the image model to accurately place the content elements in the specified design position, thereby improving the aesthetics of the layout and the efficiency of information communication.

[0115] By fusing spatial position vectors with semantic features to generate a layout control matrix, a coupling relationship between position and semantics can be established in the early stages of image generation. This allows subsequent models to not only identify "what content to generate" but also precisely control "where it is generated," thereby effectively enhancing the structural consistency and semantic expressiveness of the visual content.

[0116] S70, inputting the layout control matrix and the initial noise tensor into the codec network, fusing the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjusting the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generating the target visual content.

[0117] In this embodiment, the layout control matrix is ​​a high-dimensional structured tensor that contains the fusion of the spatial position vectors of the elements and their semantic features. This matrix forms a sequence of control vectors that are spatially and semantically bound, at the element-by-element granularity. The initial noise tensor is a tensor with the same shape as the final image but completely random values. It serves as the starting input for a diffusion model or image generation model, guiding the image's gradual evolution from pure noise to a clearly recognizable visual output.

[0118] Feeding the layout control matrix and the initial noise tensor into the encoder-decoder network together introduces structured control information into the initial stages of image generation, contributing to the spatial control and semantic regulation of the generation process. This input can be integrated through channel-wise concatenation, conditional embedding, or attention gating, enabling the generation model to possess semantic position guidance capabilities from the outset.

[0119] The regional focus weight matrix is ​​generated by the network's internal attention calculation module based on the fused features. It is a spatial distribution map, with each value representing the degree of attention paid to the corresponding location during the image generation process. This matrix guides the denoising process to focus computing resources on semantically dense areas or spatially critical regions to ensure that key content in the image is clearly presented.

[0120] The noise reduction process is regulated by using a regional focus weight matrix to achieve dynamic mask control or noise amplitude adjustment. For example, the regional focus weights act as filter weights or amplification factors on the noise estimation output. During each model update, the noise reduction amplitude is selectively strengthened or weakened in different regions based on the weights. This process is repeated over multiple iterations, gradually recovering the image from noise, ultimately yielding visual content with clear structure and accurate semantics.

[0121] One implementation approach is to concatenate the layout control matrix along the channel dimension to the initial noise tensor in a U-Net diffusion model to form a fused input tensor. The input is passed through a downsampling encoder to extract multi-scale features, which are then restored to the original image size by a decoder. Multiple attention gating modules are inserted into the decoder, taking in the fused high-level semantic information and low-level spatial information, and outputting a matrix of regional focus weights.

[0122] The region focus weight matrix can be calculated using a multi-head attention mechanism, generating a spatial attention map using the control matrix as the query and the encoded feature map as the key and value. This matrix is ​​further used to adjust the output of the noise prediction module. For example, the region focus weights are element-wise multiplied by the noise estimation result to serve as the basis for noise reduction at the current time step. The weight matrix can also be used for rescaling or region fusion when updating the intermediate image in each iterative round.

[0123] The denoising process is based on the reverse sampling algorithm in the diffusion model. DDPM (Denoising Diffusion Probabilistic Models) or its accelerated variants such as DDIM can be used, combined with regional focus weights to achieve noise step updates, and a dynamic step size control strategy can be set to adapt to differences in content complexity.

[0124] For example, in the process of visual content generation based on layout control, in order to enable the model to reasonably introduce spatial guidance information into the attention mechanism of image generation, it is first necessary to fuse text features with layout information. To this end, when calculating the attention distribution, not only is the standard attention calculation logic based on the text query vector and image key-value pairs used, but the spatial guidance weight matrix mapped by the layout control matrix is ​​also introduced into the original attention score. This step constitutes an extension of the attention mechanism that integrates layout priors in the image generation process, and its mathematical expression is:

[0125]

[0126] Among them, Q, K, and V are query matrix, key matrix, and value matrix respectively, which can be formed by fusing text and image multimodal features. k is a scaling factor that controls the scale of the attention weight. layout This is a spatial guidance module generated by the layout control matrix, where each element reflects the spatial importance of that position relative to other elements. By incorporating this term into the attention mechanism, the image decoding module prioritizes important spatial regions suggested by the layout plan, thereby improving the spatial structural consistency and semantic coordination of the generated content.

[0127] By using the layout control matrix as a guide for the image generation process and adjusting the noise reduction path through the regional focus weight matrix, we can achieve spatially targeted generation of target content in the image. Dynamically focusing on regions effectively suppresses noise interference from irrelevant areas, making the generated results more structured and semantically aligned, improving the clarity and expressiveness of the image content.

[0128] The present invention relates to the field of speech semantic technology and can be applied to business scenarios such as financial technology, medical health and poster design. A visual generation method based on semantic association modeling is disclosed, comprising: constructing a dialogue data set and generating perturbation training samples, fine-tuning parameters of a pre-trained language model, using the fine-tuned language model to process real-time interactive instructions to generate demand text, extracting element semantic features and determining semantic association weights, constructing a layout optimization function to generate element coordinate data, encoding the coordinate data into a spatial position vector and fusing it with semantic features to generate a layout control matrix, inputting the layout control matrix and an initial noise tensor into an encoding and decoding network, generating a regional focus weight matrix, and adjusting the noise reduction process to iteratively generate visual content. The present invention constructs a layout optimization function and guides the diffusion model to focus on semantically significant areas in space, closely integrating the language model output with the visual generation process, achieving structured response and spatial mapping to user semantic needs, and improving the expression consistency and personalized adaptation capabilities of visual content generation.

[0129] In one embodiment, the above step S10 includes:

[0130] S101, collecting user initial conversation records in a preset scenario, marking interaction intention labels and design requirement fields in the user initial conversation records, and generating interaction requirement guidance samples;

[0131] S102, extracting natural language instructions without semantic modification from the interaction demand guidance sample as original interaction instructions;

[0132] S103, constructing a domain-related synonym database, wherein the synonym database includes synonyms for design style descriptors, color coding parameters, and typesetting terms;

[0133] S104, identifying a target replacement word in the original interaction instruction, and generating a synonym replacement candidate set of the target replacement word based on the synonym library;

[0134] S105, screening replacement items of the target replacement vocabulary from the synonymous replacement candidate set according to a preset replacement probability threshold, and generating a synonymous replacement perturbation sample;

[0135] S106, detecting time description sentences in the original interaction instruction, and converting the absolute time description in the time description sentence into a relative time period description, to generate a time fuzzy generalized perturbation sample;

[0136] S107, identifying a spatial position description statement in the original interaction instruction, and replacing the coordinate parameters in the spatial position description statement with a directional word description to generate a spatial fuzzy generalized perturbation sample;

[0137] S108 , mixing the synonymous replacement disturbance sample, the temporal fuzzy generalization disturbance sample, the spatial fuzzy generalization disturbance sample, and the original interaction instruction in a preset ratio to generate the disturbance training sample.

[0138] In this embodiment, when constructing a corpus dataset containing interactive demand guidance samples, we should not only focus on whether the original user expression is complete, but more importantly, identify its potential driving factors in the visual content construction task. These factors include both explicit needs (such as "color", "style", and "theme") and implicit preferences (such as "emphasis rhetoric" and "spatial modification"). In the setting of guided samples, it is usually necessary to establish subsets for multiple different interaction types, such as interaction templates for title generation, graphic layout control, and visual style adaptation. Different subsets correspond to different field annotation strategies, such as whether they contain numerical parameters, whether they have directional words, etc.

[0139] To meet the needs of subsequent language model development for understanding semantically equivalent input, finer-grained pragmatic feature control is required within the perturbation generation mechanism. For example, when generating synonym replacement candidate sets, in addition to filtering based on semantic proximity of word vectors, a contextual matching evaluation metric can be introduced. This constructs a local dependency graph structure using a context window to filter out candidate words that semantically conflict within the current context. This mechanism ensures that the semantically perturbed expression maintains task consistency while also embracing linguistic diversity.

[0140] In fuzzy generalization operations, different types of expressions have different perturbation sensitivities. For example, the mapping of absolute time to relative time periods should not rely solely on rule substitution, but rather require the introduction of a time category division model (e.g., logical dimensions such as weekends / weekdays and holidays / weekdays) to ensure pragmatic rationality in the conversion results. Furthermore, the generalization of spatial location can be further combined with visualization preference rules for mixed text and image layout areas. For example, replacing "center-right" with "right of the golden section" can better align with the layout learning structure of the design perception model.

[0141] Absolute time descriptions directly indicate a specific point in time or time period, typically including a specific date, time, or time interval. For example, "December 1, 2024," "9:00 AM," and "this Wednesday afternoon" are all examples of absolute time descriptions. This type of time expression is common in user interactions, particularly in design tasks, where it is used to specify release dates, display times, holiday themes, and other scenarios. However, this description style is highly time-sensitive during model generalization, making it difficult to reuse templates or perform transfer learning.

[0142] Relative time descriptions are non-specific time expressions derived from the current time or a reference time point, such as "yesterday," "last weekend," "in the next two days," or "one hour after the event ends." Relative descriptions are more adaptable, enabling language models to learn temporal logic and content temporal relationships, thereby maintaining content rationality in tasks spanning multiple time periods. For example, the original sentence "Please make a poster for the meeting on December 1, 2024" can be transformed into "Please make a poster for the meeting at the beginning of next month."

[0143] Coordinate parameters are a precise way of expressing spatial locations, typically using numerical or directional quantities to represent a point in a layout. For example, "(120,320)" represents the horizontal and vertical coordinates of a pixel in an image, or "50 pixels from the left edge, 100 pixels from the top edge." Coordinate parameters are commonly found in graphics and annotation systems, controlling the precise placement of elements in space.

[0144] Positional descriptions convert spatial coordinates into relative positional expressions commonly found in natural language. For example, descriptions like "upper left corner," "right center," "upper right," and "center of the screen" more intuitively reflect how humans understand spatial relationships. Generalizing coordinate parameters into positional descriptions helps create more readable and versatile semantic input and improves user comprehension of interactive content. For example, the original sentence "The button is located at (300, 150)" can be generalized to "The button is located slightly above the screen."

[0145] In addition, the perturbation training samples are not only the corpus for training the language model, but also undertake the task of early input adaptation of the model structure and behavior. Different distribution forms of semantic perturbations will form different gradient responses in the adjustment of pre-training model parameters. Therefore, when designing perturbation samples, it is also necessary to balance their language complexity, degree of semantic deviation and training sensitivity. For example, the perturbation level distribution can be set, from "light perturbation" (such as synonym replacement) to "medium perturbation" (such as fuzzy time replacement), and then to "strong perturbation" (such as multi-field composite replacement) to build a layered perturbation sample pool, so that the model can gradually establish a robust representation capability for semantic diversity.

[0146] Finally, when fusing perturbed samples with original instructions to generate a training set, a perturbation source field can be added to the semantic label layer to record whether each sample contains semantic perturbations, the type of perturbation, and the location range. This structural annotation not only helps regulate losses during training but also provides conditional guidance during model prediction, further improving the model's ability to recognize and respond to input semantic structures.

[0147] This embodiment constructs guidance samples with real-world interaction semantic labels and design requirement fields, and performs multiple types of semantic perturbations on the original expressions. This effectively improves the language model's generalization capabilities in multimodal generation tasks, enhancing its understanding of synonyms, ambiguous sentences, and spatial metaphors. This mechanism not only reduces the risk of the model overfitting to specific expressions but also significantly improves its tolerance to user interaction errors in actual deployment scenarios.

[0148] In one embodiment, the above step S30 includes:

[0149] S301, sending an initial guidance instruction to the interactive terminal;

[0150] S302, receiving a first interaction instruction returned by the interaction terminal, and parsing a topic keyword in the first interaction instruction;

[0151] S303, generating an initial candidate set of topic identification fields based on the topic keywords;

[0152] S304, calling a preset style guidance template to send a style level selection instruction to the interactive terminal, wherein the style level selection instruction includes a plurality of preset design style options;

[0153] S305: Receive a second interaction instruction returned by the interaction terminal, and verify whether the design style type in the second interaction instruction belongs to the preset design style option;

[0154] S306, extracting associated attribute parameters from the verified design style type, and generating a style level parameter set based on the associated attribute parameters;

[0155] S307: Integrate the initial candidate set of the subject identification field and the style level parameter set into a structured requirement text.

[0156] In this embodiment, during the user interaction process, the user is first guided to clarify the theme and key elements of the visual content they wish to generate by sending an initial guidance instruction. This initial guidance instruction should not be regarded as a static template, but should be dynamically generated according to the scenario, with a semantic framework that adapts to different fields or task requirements. The first interaction instruction returned by the interactive end is usually a natural language description, such as "I want to design a promotional image about low-carbon life". The system needs to extract keywords such as "low-carbon life" from it, and further input them as theme keywords into the analysis module in the next stage.

[0157] By parsing topic keywords, we can establish a preliminary set of candidate topic identifier fields. This set includes not only the original topic but also semantically related extension candidates obtained by querying a pre-built semantic extension knowledge base, synonym graph, near-topic mapping, or graph neural network inference models. For example, "environmental protection," "green travel," and "energy-saving living" can all be candidates for "low-carbon living." This initial set provides a semantic foundation for subsequent generation of personalized visual content.

[0158] Based on the style guidance template, the system also sends a style hierarchy selection request to the user, containing several visual design style options. The pre-set style categories in the template must be adaptable across scenarios, such as "Technology," "Minimalist," "Flat," "Retro," "Illustration," and "Visual Illustration." The style hierarchy design considers not only the overall style type but also sub-attributes such as color scheme, typography density, and background texture.

[0159] The second interaction instruction returned by the user is verified to ensure that its content falls within the recognizable style range. Once verified, attribute parameters are extracted from the style instruction, which will form the key components of a structured style description. Unlike simple label classification, this extraction process can be achieved by parsing style dictionaries, hierarchical graphs, and style parameter configuration files, such as "illustration style + soft color scheme + high white space ratio."

[0160] After completing the dual structuring of theme and style, the system merges the two parts into a structured demand text. This text is expressed in a key-value pair format, such as {"Theme":"Low-Carbon Life","Style":{"Main Style":"Illustration Style","Color":"Soft","Layout":"White Space Priority"}}. This structure provides a unified, parseable format that is compatible with downstream semantic processing modules and generation engines.

[0161] Before outputting this structured text, the system uses a fine-tuned language model to perform a semantic integrity check. This verification goes beyond formatting and focuses on semantic logic consistency, the harmony of theme and style, and the absence of key semantic elements. The language model makes a comprehensive assessment based on the structural dependencies and contextual semantic compatibility captured in the training data.

[0162] The final output results must meet the requirements of structural standardization, semantic integrity and clear expression. All field values ​​are normalized to facilitate accurate use in subsequent semantic feature extraction, layout optimization and image generation.

[0163] This semantic recognition and structure generation process can be implemented through a natural language processing module based on multiple rounds of interaction. Initial guidance instructions are dynamically generated by the template retrieval module based on the current user profile or domain label. Topic extraction uses a combination of a named entity recognition model and a topic classifier. Semantic expansion uses a BERT-based vector matching model combined with graph neural reasoning to generate a candidate set of topic expansions. Style analysis relies on a combination of a style classification model and a rule graph to perform structured analysis of style statements. In the verification phase, a trained multi-task language understanding model is used to score topic consistency, style rationality, and expression completeness, with structured output only completed when the set threshold is reached.

[0164] In different implementations, the design scope of the style guidance template can be adjusted to adapt to different usage scenarios such as visual advertising, popular science explanations, and data presentations; the subject knowledge base used for semantic expansion can also be replaced according to the target field, such as medical and health vocabulary, financial information graphs, etc.; the degree of fine-tuning of the language model can also be adjusted to control the breadth of understanding of new semantic expressions.

[0165] This embodiment establishes a structured demand text that complies with standards, is complete, and is parseable through multiple rounds of interactive understanding, semantic keyword extraction, and style structure generation based on user natural language input. This process significantly improves the system's ability to understand complex, ambiguous, and personalized user input, providing structured semantic support for subsequent layout optimization and image generation. This avoids the problems of missing expressions, misunderstandings, and content conflicts that exist in traditional approaches based on keyword retrieval or static template mapping.

[0166] In one embodiment, the above step S40 includes:

[0167] S401, parsing all semantic elements from the demand text, and inputting the semantic elements into a pre-trained semantic encoder to generate a semantic feature vector of each semantic element;

[0168] S402, determining the cosine similarity value between every two semantic feature vectors to generate a semantic similarity matrix;

[0169] S403, performing normalization processing on the semantic similarity matrix to generate a semantic association weight matrix;

[0170] S404: Filter association weight items in the semantic association weight matrix that are lower than a preset weight threshold to generate a filtered semantic association weight set.

[0171] In this embodiment, to measure the strength of association between semantic elements, all semantic elements must first be extracted from the structured demand text. These semantic elements include both the keyword set extracted from the subject field and descriptive items in the style parameters, such as "low-carbon living," "green energy," "illustration style," "light green," and "white space composition." These semantic units may have multi-dimensional semantic expressions and require representation using a unified vector encoding format.

[0172] By encoding each semantic element through a semantic encoder, it can be converted into a real-number vector representation with uniform dimensionality. The encoder used can adopt the embedding layer of a pre-trained language model or a Transformer structure fine-tuned through a multimodal alignment task. This ensures that the generated semantic vector not only captures the static semantic relationships at the lexical level but also incorporates pragmatic features expressed in the context.

[0173] After obtaining all semantic feature vectors, the system calculates the cosine similarity between all pairs of semantic elements one by one to construct a complete semantic similarity matrix. Each item in this matrix represents the degree of similarity between two elements in the semantic space, and the value range is between -1 and 1, where 1 represents complete semantic consistency and 0 represents irrelevance. Negative values ​​can be used to represent semantic opposition but are usually set to 0 and ignored in this scenario. To achieve numerical uniformity and stability of subsequent calculations, the system normalizes the semantic similarity matrix, using Min-Max normalization or Softmax normalization strategies to compress all associated values ​​to between 0 and 1, facilitating coordinated adjustment in subsequent weight calculations and visual position fusion processes.

[0174] Since there are a large number of weakly correlated or irrelevant combinations between semantic elements, in order to avoid low-confidence information interfering with the layout and content generation process, the system will perform filtering operations on the normalized weight matrix. The specific approach is to introduce a weight threshold, retaining only semantic associations above the threshold, treating them as strong semantic connections, and incorporating them into the subsequent layout optimization function construction. This threshold can be set to a fixed value (such as 0.3), or it can be dynamically adjusted based on the mean and standard deviation of the weight distribution in the current task through an adaptive mechanism. The system can set an initial static threshold, and in actual deployment, it allows dynamic recalculation through hyperparameter tuning or user feedback, so that the filtering strategy is flexible and task-relevant.

[0175] The final retained weight set constitutes the semantic association weight set, which serves as the input for the subsequent construction of semantic space constraints and visual layout optimization models, ensuring that the generated content fully reflects the aggregation and differentiation of semantic relevance in spatial configuration.

[0176] In implementation, BERT or RoBERTa can be used as the base semantic encoder. By freezing its parameters or performing moderate fine-tuning on small-scale tasks, it can achieve good semantic differentiation in the vectorized representation of structured semantic fragments. Cosine similarity calculations use standard formulas, improving computational efficiency through matrix operations, allowing for rapid comparisons of thousands of element pairs in multi-threaded or GPU environments. Softmax normalization is recommended for normalization, ensuring that the sum of all associated values ​​for each row vector element is 1, making it suitable for weighted attention mechanisms.

[0177] Regarding threshold settings, an initial value of 0.3 can be set in the system settings and adjusted based on the application. For example, in medical illustration scenarios where stricter distinctions between terms are required, the value can be set to 0.4 or higher. In style-aware posters, where more extensive collaboration of style items is required, the value can be set to 0.2 to include more connected items. The adaptive threshold mechanism automatically determines the cutoff point based on statistical characteristics, such as using a value greater than the mean plus twice the standard deviation as the threshold boundary.

[0178] Example: In a healthcare scenario, the user provides a structured semantic instruction: "Theme: Diabetes Management, Style: Illustrated, Tone: Cool Colors." The system extracts semantic elements such as "diabetes," "management strategy," "illustration," and "blue," and generates vectors through the encoder. Analysis shows that the correlation between "diabetes" and "management strategy" is as high as 0.91, and the correlation between "blue" and "cool colors" is 0.88, both exceeding the set threshold of 0.3. These elements are retained as the semantic association weight set.

[0179] In a financial scenario, the user requested to create a graphic and text content introducing a pension insurance product. The system extracted semantic elements such as "pension", "retirement plan", "safe and stable", and "light gold". After analysis, the highly correlated element pairs of "pension" and "retirement plan" (0.94) and "safe and stable" and "light gold" (0.82) were retained for the subsequent layout constraint model.

[0180] In poster creation, the user provided the theme of "Environmentally Friendly Life" and the style of "Illustration Style + Green Main Tone". The system discovered through the semantic encoder that the semantic correlation between "Illustration Style" and "Green Main Tone" is 0.77, which is much higher than the set threshold. Based on this, the system strengthens the spatial aggregation of these two elements, so that the image presents a consistent style and content orientation.

[0181] This embodiment quantifies and filters the similarities between structured semantic elements to construct a set of association weights with reasonable semantic strength, sparse structure, and high task relevance, thereby enhancing the semantic coordination ability in downstream visual layout, having higher semantic sensitivity and adaptability, and effectively avoiding meaningless connections that interfere with the arrangement logic, thereby improving the quality of the final visual content in terms of thematic expression and visual consistency.

[0182] In one embodiment, the above step S50 includes:

[0183] S501, generating layout guiding coordinates by a layout planning agent module, wherein the layout guiding coordinates are determined based on a preset layout strategy or a historical layout pattern;

[0184] S502: Obtain a Gaussian kernel function of a spatial distribution constraint term, wherein the Gaussian kernel function determines the spatial correlation strength between the two semantic elements based on a ratio of the square of an actual distance between the two semantic elements to a preset bandwidth parameter;

[0185] S503, constructing a semantic space coordination term of the layout optimization function, wherein the semantic space coordination term is obtained by multiplying the semantic association weight of each semantic element pair by the output value of the corresponding Gaussian kernel function and then summing the results;

[0186] S504, constructing a layout guide matching item of the layout optimization function, wherein the layout guide matching item is obtained by determining the sum of squared Euclidean distances between the coordinates of all semantic elements and the layout guide coordinates and multiplying the sum by a preset balance weight parameter;

[0187] S505, adding the semantic space collaboration item and the layout guidance matching item to generate a final layout optimization function;

[0188] S506 , randomly initializing two-dimensional coordinate values ​​for all semantic elements to generate an initial layout coordinate set;

[0189] S507, performing multiple rounds of iterative optimization on the initial layout coordinate set using a back-propagation algorithm, performing a gradient calculation operation in each round of iteration, wherein the gradient calculation operation is to calculate the partial derivative value of the layout optimization function with respect to the horizontal and vertical coordinates of each semantic element;

[0190] S508, updating the coordinate values ​​of all semantic elements along the reverse direction of the gradient according to a preset step size according to the output result of the gradient calculation operation;

[0191] S509, determining whether the updated layout optimization function value satisfies a preset convergence condition;

[0192] S510 , when the updated layout optimization function value satisfies a preset convergence condition, outputting optimized semantic element coordinate data.

[0193] In this embodiment, when generating the spatial layout of semantic elements, the optimization function constructed by fusing semantic association weights with spatial distribution constraints has two core goals: one is to drive semantically closely related elements closer together in two-dimensional space, thereby maintaining visual semantic coherence; the other is to maintain the spatial balance of the overall layout, so that the generated results have good readability and visual aesthetics. To achieve the above goals, preliminary layout guidance coordinates are first generated by the layout planning agent module. This module can calculate recommended positions based on preset structural templates, user preferences, or historical sample experience, or an adaptive position model formed based on multi-task learning. These guidance coordinates serve as spatial prior information to limit the degree of deviation during the layout generation process and ensure the stability and aesthetics of the output structure.

[0194] To express the semantic dependencies between elements, a semantic association weight mechanism is introduced to quantify the intrinsic semantic connections between semantic elements. This semantic association weight can be calculated based on the similarity matrix output by a pre-trained language model or graph neural network. A higher value indicates that the two semantic elements are spatially closer. To measure the spatial proximity between elements, a spatial distribution constraint based on a Gaussian kernel function is introduced. Its functional form is determined by the ratio of the square of the element spacing to the bandwidth parameter. This constraint is used to characterize the impact of actual spatial distance on layout quality. The closer the output value of the Gaussian kernel function is to 1, the closer the actual spatial positions of the two elements are.

[0195] Combining the two aforementioned aspects, a layout optimization function is constructed, which includes a semantic space coordination term and a layout guidance matching term. The semantic space coordination term spatially clusters semantically related elements by multiplying the semantic weights by the Gaussian kernel function outputs pairwise. The layout guidance matching term calculates the sum of the squared Euclidean distances between each semantic element's current position and the guidance coordinates, and applies a balancing factor to regulate its influence. The weighted sum of the two terms constitutes the overall optimization objective. In the early stages of training, the weight of the guidance term can be appropriately increased to ensure stability, while the weight of the coordination term can be gradually increased during the convergence phase to enhance the semantic clustering effect.

[0196] During the optimization process, a two-dimensional coordinate vector is initialized for each semantic element, forming an initial layout coordinate set. Multiple rounds of iterations are performed on the current layout coordinate set using the backpropagation algorithm. Each round requires calculating the partial derivatives of the optimization function with respect to each element's horizontal and vertical coordinates to guide the coordinate update. During the update process, the step size of each coordinate change can be adjusted based on a dynamic learning rate strategy to balance convergence speed and avoidance of local optimality.

[0197] To improve optimization robustness and adapt to termination requirements in various scenarios, various convergence criteria can be set. These include, but are not limited to: the change in the optimization target value for several consecutive rounds falling below a threshold; the optimization target value failing to significantly improve within a certain number of rounds; the coordinate update amplitude falling below a set threshold; or reaching the maximum number of iterations. The system automatically selects or combines appropriate convergence strategies based on the model training status, balancing computational overhead and generation quality.

[0198] The final coordinate output contains the location of each semantic element in two-dimensional space and its unique identification information. This identification information can be a text description, image label, or structure ID, facilitating data alignment and fusion processing with the visual generation module in subsequent steps. The entire process ensures personalized layout, controllability, and semantic consistency.

[0199] This embodiment constructs a layout optimization function by combining semantic association weights with spatial distribution constraints. This allows for both semantic aggregation requirements and spatial structural rationality during the optimization process, allowing semantically related visual elements to be naturally close together in the layout, thereby enhancing the semantic coherence and visual readability of the generated content. Layout guide matching items are constructed with the help of guide coordinates to provide a solid reference for the overall layout structure, avoid excessively discrete or chaotic element distribution, and effectively improve the structural stability and layout aesthetics of the generated results. Further, through backpropagation optimization and iterative updating, dynamic balance control of both semantic and spatial objectives is achieved, and multiple convergence conditions are introduced to give the optimization process good adaptability and termination robustness, thereby outputting high-quality, clearly structured, and semantically consistent element coordinate data in different application scenarios.

[0200] In one embodiment, the above step S70 includes:

[0201] S701, concatenating the layout control matrix and the initial noise tensor along the channel dimension to generate a fused input tensor, and inputting the fused input tensor into the encoding and decoding network;

[0202] S702, performing a multi-scale downsampling operation on the fused input tensor through the encoder of the codec network to generate a multi-level encoding feature map;

[0203] S703: Input the multi-level encoding feature map into the attention gating module of the decoder of the encoding and decoding network to perform cross-layer feature aggregation to generate a regional focus weight matrix, where each element value of the regional focus weight matrix represents the attention intensity of the corresponding spatial position in the noise reduction process;

[0204] S704, initializing the intermediate image to the initial noise tensor;

[0205] S705, performing multi-step denoising iterations to generate target visual content through a diffusion model sampling strategy;

[0206] S706, in each round of noise reduction iteration, generating a noise prediction output based on the intermediate image at the current time step and the conditional input of the layout control matrix, multiplying the noise prediction output by the regional focus weight matrix element-by-element to generate an adjusted noise prediction result, updating the intermediate image according to the adjusted noise prediction result, and scaling the regional focus weight matrix according to the noise level at the current time step;

[0207] S707: When the number of iterations reaches the preset number of sampling steps, the target visual content is output.

[0208] In this embodiment, the layout control matrix is ​​a structural representation of the relationship between the semantics of the encoding element and its spatial position, which contains the positioning information of each semantic element in two-dimensional space and the semantic feature vector associated with the element. The initial noise tensor is the Gaussian noise data used for the initial state of image generation in the diffusion model generation task, usually in the form of a multi-channel tensor. By splicing the layout control matrix and the initial noise tensor along the channel dimension, cross-modal feature fusion can be achieved without destroying the original spatial distribution, effectively coupling the semantic constraints with the initial state of image generation.

[0209] The concatenated fused input tensor is fed into the encoder-decoder network, which consists of two main components: an encoder and a decoder. The encoder extracts feature representations at different resolutions from the fused input tensor through a multi-scale downsampling mechanism. A common implementation involves successive convolutional layers followed by pooling with a stride of 2, ultimately resulting in a multi-level encoded feature map. The multi-level feature map is then fed into the attention gating module within the decoder for cross-layer aggregation. This process can be based on spatial attention, channel attention, or a dual-branch gating mechanism, aiming to preserve semantic guidance while restoring image resolution. The output of cross-layer feature fusion is used to generate a regional focus weight matrix. The value at each position in this matrix indicates the level of attention that spatial region should receive during denoising. Higher values ​​indicate more semantically driven content at that position.

[0210] The initial state of the intermediate image is the initial noise tensor, which is then gradually de-noised under the control of the diffusion model. The diffusion model uses a sampling strategy to perform a multi-step image restoration operation, and each sampling iteration is dynamically updated with the current time step as a reference. In each iteration, the model generates a noise prediction output for the current moment based on the current intermediate image and the layout control matrix. This output is multiplied element-by-element by the regional focus weight matrix to obtain the adjusted noise prediction result. This operation dynamically adjusts the regional noise reduction intensity, prioritizing noise removal in semantically critical areas while retaining detailed information in non-critical areas, thereby improving the semantic consistency of the overall image.

[0211] The region focus weight matrix is ​​scaled at each iteration based on the noise level at the current time step. Common scaling strategies include linear decay, exponential decay, or a cosine function decay associated with the time step. Ultimately, after reaching the preset number of sampling steps, the model outputs a denoised target visual content image. The spatial positions of each semantic element in this image are consistent with the coordinate data defined in the layout control matrix, thus achieving explicit control over the spatial structure of the generated result.

[0212] In practical applications, the U-Net structure can be used as the basic architecture of the encoder-decoder network. Its encoder consists of four layers of convolutional blocks and a downsampling module. The decoder is equipped with four layers of upsampling operations and a cross-layer connection mechanism. A spatial attention mechanism is embedded to generate a regional focus weight matrix. The regional focus weight matrix is ​​generated through Softmax normalization to form a probabilistic distribution for all spatial locations.

[0213] The initial noise tensor's dimensions match the target image's, and the number of channels can be 3 or 4, depending on whether a transparency channel is included. The layout control matrix concatenates the semantic feature encoding of each element with its coordinate position information through a mapping function, repeatedly expanding the matrix to match the spatial dimensions of the noise tensor. During the denoising iterations, the noise prediction module can utilize a residual block structure supplemented by a conditional attention mechanism to further enhance modeling of key regions.

[0214] The scaling of the region-focusing weight matrix is ​​dynamically adjusted based on the diffusion time step t, and can be calculated using the following function: Weight scaling factor = cos(π·t / T), where T is the total number of sampling steps. The final output image is evaluated for quality using a discriminator or perceptual consistency network to ensure that it meets semantic constraints while maintaining a high degree of naturalness.

[0215] This embodiment can achieve the fusion of semantic spatial information and image generation process by splicing the layout control matrix and the initial noise tensor in the channel dimension and inputting them into the codec network. The regional focus weight matrix introduces a spatial attention mechanism, which enables the model to dynamically adjust the restoration strength of each region during the denoising process, thereby improving semantic consistency and spatial accuracy. The denoising process is guided by the regional focus matrix and performs multiple rounds of iterative optimization, which can gradually enhance the image quality of semantically key areas and control the texture diffusion of non-key areas. Through the controllable sampling mechanism of the diffusion model, it ultimately achieves structurally aligned and style-unified visual content output.

[0216] In one embodiment, after the above step S70, the method further includes:

[0217] S801, receiving adjustment parameters returned by a feedback end, where the adjustment parameters are generated based on a user's layout score and semantic relevance annotation results for the target visual content;

[0218] S802, extracting the historical task contributions of the language model parameters from a pre-stored parameter importance matrix, and constructing an elastic weight solidification loss function in combination with the adjustment parameters;

[0219] S803, calculating the gradient direction and update amount of the parameters of the language model using a back propagation algorithm based on the elastic weight solidification loss function to generate an updated language model;

[0220] S804: Process the real-time interaction instruction using the updated language model to generate optimized visual content aligned with the user feedback.

[0221] In this embodiment, after the layout control matrix and the initial noise tensor are input into the codec network, generating the target visual content after multiple rounds of iterations only generates the initial result. In actual use, users usually provide specific adjustment feedback based on the visual content. Therefore, the system introduces a feedback adjustment mechanism. First, the front-end interaction module collects the user's evaluation data on the current visual content. These evaluations may include multiple dimensions such as spatial structure rationality, semantic element relevance, and style consistency. The feedback information structurally contains two key sub-items: layout score and semantic relevance annotation. The former measures the spatial rationality of the visual arrangement, and the latter marks whether the semantic combination conforms to the expected logic.

[0222] To prevent the model from forgetting its original multimodal knowledge during fine-tuning, a parameter importance matrix is ​​introduced as a constraint. This matrix records the model's sensitive parameters in various multimodal tasks over the course of historical training, reflecting the importance of each parameter to the original task performance. This parameter importance is typically estimated based on the Fisher information matrix, emphasizing the parameter's responsiveness to changes in the original loss.

[0223] By integrating the current user's adjustment intent with the need to maintain historical tasks, we construct an elastic weighted fixed loss function. This function consists of two parts: a target loss term for the new task, which captures the feature adjustment direction driven by current user feedback; and a regularization term for the old task, which applies a weighted penalty to important parameters to prevent the model from forgetting prior knowledge. This loss structure is reflected in the language model by applying different update strengths to different parameters during the update process: important parameters are strongly constrained, while low-weight parameters are allowed a wider range of updates.

[0224] Guided by this loss function, the language model's gradient direction is calculated through standard backpropagation, and weight updates are generated based on the learning rate parameter, resulting in an updated model that is compatible with the new task while retaining the capabilities of the old task. This model is reused to parse real-time interactive commands, automatically responding to user-indicated deviations in visual content generation, such as adjusting layout density, strengthening semantic aggregation, or replacing specific color expressions. At the same time, the parts protected by regularization still maintain the continuity of the original style, ensuring a stable and consistent overall visual language.

[0225] The structure of feedback information can be adjusted in different usage scenarios. For example, in medical scenarios requiring strong semantic interpretation capabilities, feedback content can focus on the logical rationality of medical term combinations, while in advertising and marketing scenarios, more emphasis can be placed on color matching or composition tension scores. The parameter importance matrix can be generated using different calculation methods, such as training error sensitivity, feature entropy-based interpretability scores, etc. The construction of elastic weight solidification loss can also set the regularization term weight coefficient based on the risk tolerance of different scenarios.

[0226] Backpropagation calculations can use optimizers such as AdamW to enhance learning of historical gradients. Update strategies support multiple rounds of iteration at the batch or mini-batch level. Updated language models can be temporarily stored in a candidate model pool, and the optimal version can be selected for online deployment through a rapid A / B testing mechanism to accommodate the frequent iteration requirements of user changes.

[0227] For example, after performing visual content generation, in order to respond to user feedback and adjust the language model parameters so that it better retains the style capabilities of the original task while adapting to the optimization requirements of the new task, it is necessary to introduce an elastic weight solidification mechanism. This mechanism is based on the regularization concept in multi-task learning and incorporates the difference between the current language model parameters and the historical model parameters into the loss function as a constraint. The importance of the parameters is represented by the parameter importance matrix pre-calculated during task training. The logical process of parameter update can be expressed by the following formula:

[0228]

[0229] Among them, ρ i Represents the i-th optimizable parameter in the current language model; represents the optimal estimated value of the i-th parameter in the t-th historical task; F i represents the importance weight of the i-th parameter evaluated in the historical task, which is usually calculated based on the Fisher information matrix; L(ρ) represents the main loss function constructed based on the target feedback in the current task, such as prediction error or cross entropy loss; γ represents the adjustment coefficient used to balance the current task goal and the historical task knowledge retention, controlling the strength of the regularization term.

[0230] This embodiment introduces a feedback adjustment mechanism and a flexible weight solidification structure, enabling the language model to quickly and adaptively modify the semantics and style of user-specified content without losing its original multimodal understanding capabilities. This mechanism not only strengthens the dominance of user intent in the model update process, but also preserves the model's ability to express feature distribution in the original task. This achieves a dynamic balance between stability and flexibility in the generated results between the new and old tasks, thereby improving the personalization and overall visual consistency of the generated content.

[0231] In one embodiment, a visual generation device based on semantic association modeling is provided, and the visual generation device based on semantic association modeling corresponds one-to-one to the visual generation method based on semantic association modeling in the above embodiment. Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of a visual generation device based on semantic association modeling according to the present invention. These modules include an interactive data construction module 10, a language model fine-tuning module 20, a demand text generation module 30, a semantic association modeling module 40, a layout optimization and analysis module 50, a layout control fusion module 60, and a visual content generation module 70. Each functional module is described in detail below:

[0232] An interaction data construction module 10 is configured to construct a dialogue dataset containing interaction demand guidance samples, extract original interaction instructions from the dialogue dataset, and perform semantic perturbation operations on the original interaction instructions to generate perturbation training samples;

[0233] A language model fine-tuning module 20 is configured to input the perturbed training sample into a pre-trained language model and fine-tune the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model;

[0234] The demand text generation module 30 is configured to receive a real-time interaction instruction, input the real-time interaction instruction into the fine-tuned language model, and generate a demand text including a topic identification field and style level parameters;

[0235] A semantic association modeling module 40 is configured to extract semantic features of elements in the requirement text and determine association weights between the semantic features of the elements as semantic association weights;

[0236] a layout optimization analysis module 50 for combining the semantic association weights with preset spatial distribution constraints to construct a layout optimization function, and generating element coordinate data by optimizing the layout optimization function through backpropagation;

[0237] A layout control fusion module 60 is configured to encode the element coordinate data into a spatial position vector, and fuse the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0238] The visual content generation module 70 is used to input the layout control matrix and the initial noise tensor into the codec network, fuse the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjust the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generate the target visual content.

[0239] In one embodiment, the interactive data construction module 10 is specifically configured to:

[0240] Collecting user initial conversation records under preset scenarios, marking interaction intention labels and design requirement fields in the user initial conversation records, and generating interaction requirement guidance samples;

[0241] extracting natural language instructions without semantic modification from the interaction demand guidance sample as original interaction instructions;

[0242] Building a domain-related synonym library, wherein the synonym library includes synonyms of design style descriptors, color coding parameters and typography terms;

[0243] identifying a target replacement vocabulary in the original interaction instruction, and generating a synonym replacement candidate set of the target replacement vocabulary based on the synonym library;

[0244] According to a preset replacement probability threshold, selecting replacement items of the target replacement vocabulary from the synonymous replacement candidate set to generate a synonymous replacement perturbation sample;

[0245] Detecting time description sentences in the original interaction instructions, and converting the absolute time description in the time description sentences into relative time period descriptions, to generate time fuzzy generalized perturbation samples;

[0246] Identifying spatial position description sentences in the original interaction instructions, and replacing coordinate parameters in the spatial position description sentences with directional word descriptions to generate spatial fuzzy generalized perturbation samples;

[0247] The synonymous replacement disturbance sample, the time fuzzy generalization disturbance sample, the space fuzzy generalization disturbance sample and the original interaction instruction are mixed in a preset ratio to generate the disturbance training sample.

[0248] In one embodiment, the requirement text generation module 30 is specifically configured to:

[0249] Send initial guidance instructions to the interactive end;

[0250] receiving a first interaction instruction returned by the interaction terminal, and parsing a subject keyword in the first interaction instruction;

[0251] Generating an initial candidate set of subject identification fields according to the subject keywords;

[0252] Calling a preset style guidance template to send a style level selection instruction to the interactive terminal, wherein the style level selection instruction includes a plurality of preset design style options;

[0253] receiving a second interaction instruction returned by the interaction terminal, and verifying whether the design style type in the second interaction instruction belongs to the preset design style option;

[0254] Extracting associated attribute parameters from the verified design style type, and generating a style level parameter set based on the associated attribute parameters;

[0255] The initial candidate set of the subject identification field and the style level parameter set are integrated into a structured requirement text.

[0256] In one embodiment, the semantic association modeling module 40 is specifically configured to:

[0257] Parsing all semantic elements from the demand text, and inputting the semantic elements into a pre-trained semantic encoder to generate a semantic feature vector for each semantic element;

[0258] Determine the cosine similarity value between each two semantic feature vectors and generate a semantic similarity matrix;

[0259] performing normalization processing on the semantic similarity matrix to generate a semantic association weight matrix;

[0260] The association weight items in the semantic association weight matrix that are lower than a preset weight threshold are filtered to generate a filtered semantic association weight set.

[0261] In one embodiment, the layout optimization analysis module 50 is specifically configured to:

[0262] generating layout guiding coordinates by a layout planning agent module, wherein the layout guiding coordinates are determined based on a preset layout strategy or a historical layout pattern;

[0263] Obtaining a Gaussian kernel function of a spatial distribution constraint term, wherein the Gaussian kernel function determines a spatial correlation strength between the two semantic elements based on a ratio of a square of an actual distance between the two semantic elements and a preset bandwidth parameter;

[0264] Constructing a semantic space coordination term of the layout optimization function, wherein the semantic space coordination term is obtained by multiplying the semantic association weight of each semantic element pair by the output value of the corresponding Gaussian kernel function and then summing the results;

[0265] Constructing a layout guidance matching item of the layout optimization function, wherein the layout guidance matching item is obtained by determining the sum of squared Euclidean distances between the coordinates of all semantic elements and the layout guidance coordinates and multiplying the sum by a preset balance weight parameter;

[0266] Adding the semantic space coordination item and the layout guidance matching item to generate a final layout optimization function;

[0267] Randomly initialize the two-dimensional coordinate values ​​of all semantic elements to generate an initial layout coordinate set;

[0268] Perform multiple rounds of iterative optimization on the initial layout coordinate set using a back-propagation algorithm, performing a gradient calculation operation in each iteration, wherein the gradient calculation operation is to calculate the partial derivative value of the layout optimization function with respect to the horizontal and vertical coordinates of each semantic element;

[0269] According to the output result of the gradient calculation operation, the coordinate values ​​of all semantic elements are updated along the reverse direction of the gradient according to a preset step size;

[0270] Determine whether the updated layout optimization function value meets the preset convergence condition;

[0271] When the updated layout optimization function value meets the preset convergence condition, the optimized semantic element coordinate data is output.

[0272] In one embodiment, the visual content generation module 70 is specifically configured to:

[0273] Concatenate the layout control matrix and the initial noise tensor along the channel dimension to generate a fused input tensor, and input the fused input tensor into the encoding and decoding network;

[0274] Performing a multi-scale downsampling operation on the fused input tensor through the encoder of the codec network to generate a multi-level encoding feature map;

[0275] Inputting the multi-level encoding feature map into the attention gating module of the decoder of the encoding and decoding network for cross-layer feature aggregation to generate a regional focus weight matrix, wherein each element value of the regional focus weight matrix represents the attention intensity of the corresponding spatial position in the noise reduction process;

[0276] Initializing the intermediate image to the initial noise tensor;

[0277] Target visual content is generated by performing multi-step denoising iterations through a diffusion model sampling strategy;

[0278] In each round of denoising iteration, a noise prediction output is generated based on the intermediate image at the current time step and the conditional input of the layout control matrix, the regional focus weight matrix is ​​element-wise multiplied by the noise prediction output to generate an adjusted noise prediction result, the intermediate image is updated according to the adjusted noise prediction result, and the regional focus weight matrix is ​​scaled according to the noise level at the current time step;

[0279] When the number of iterations reaches the preset number of sampling steps, the target visual content is output.

[0280] In one embodiment, the visual content generation module 70 is specifically configured to:

[0281] receiving adjustment parameters returned by the feedback end, wherein the adjustment parameters are generated based on the layout score and semantic relevance annotation result of the target visual content by the user;

[0282] Extracting the historical task contributions of the language model parameters from a pre-stored parameter importance matrix, and constructing an elastic weight curing loss function in combination with the adjustment parameters;

[0283] Based on the elastic weight solidification loss function, the gradient direction and update amount of the parameters of the language model are calculated by a back propagation algorithm to generate an updated language model;

[0284] The real-time interaction instructions are processed through the updated language model to generate optimized visual content aligned with the user feedback.

[0285] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a visual generation method based on semantic association modeling.

[0286] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the user side of a visual generation method based on semantic association modeling.

[0287] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0288] Constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing semantic perturbation operations on the original interaction instructions to generate perturbation training samples;

[0289] Inputting the perturbed training sample into a pre-trained language model, and fine-tuning the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model;

[0290] receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters;

[0291] Extracting semantic features of elements in the requirement text, and determining association weights between the semantic features of each element as semantic association weights;

[0292] Combining the semantic association weight with a preset spatial distribution constraint term to construct a layout optimization function, and optimizing the layout optimization function through backpropagation to generate element coordinate data;

[0293] Encoding the element coordinate data into a spatial position vector, and fusing the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0294] The layout control matrix and the initial noise tensor are input into a codec network, and the layout control matrix and the initial noise tensor are fused through the codec network to generate a regional focus weight matrix. The noise reduction process of the codec network is adjusted based on the regional focus weight matrix to iteratively generate target visual content.

[0295] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0296] Constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing semantic perturbation operations on the original interaction instructions to generate perturbation training samples;

[0297] Inputting the perturbed training sample into a pre-trained language model, and fine-tuning the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model;

[0298] receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters;

[0299] Extracting semantic features of elements in the requirement text, and determining association weights between the semantic features of each element as semantic association weights;

[0300] Combining the semantic association weight with a preset spatial distribution constraint term to construct a layout optimization function, and optimizing the layout optimization function through backpropagation to generate element coordinate data;

[0301] Encoding the element coordinate data into a spatial position vector, and fusing the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix;

[0302] The layout control matrix and the initial noise tensor are input into a codec network, and the layout control matrix and the initial noise tensor are fused through the codec network to generate a regional focus weight matrix. The noise reduction process of the codec network is adjusted based on the regional focus weight matrix to iteratively generate target visual content.

[0303] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0304] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0305] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0306] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A visual generation method based on semantic association modeling, characterized in that: The following steps are involved: Constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing semantic perturbation operations on the original interaction instructions to generate perturbation training samples; Inputting the perturbed training sample into a pre-trained language model, and fine-tuning the parameters of the pre-trained language model using a low-rank adaptation matrix to generate a fine-tuned language model; receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters; Extracting semantic features of elements in the requirement text, and determining association weights between the semantic features of each element as semantic association weights; Combining the semantic association weight with a preset spatial distribution constraint term to construct a layout optimization function, and optimizing the layout optimization function through backpropagation to generate element coordinate data; Encoding the element coordinate data into a spatial position vector, and fusing the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix; The layout control matrix and the initial noise tensor are input into a codec network, and the layout control matrix and the initial noise tensor are fused through the codec network to generate a regional focus weight matrix. The noise reduction process of the codec network is adjusted based on the regional focus weight matrix to iteratively generate target visual content.

2. The visual generation method based on semantic association modeling according to claim 1, characterized in that: Constructing a dialogue dataset containing interaction demand guidance samples, extracting original interaction instructions from the dialogue dataset, and performing semantic perturbation operations on the original interaction instructions to generate perturbation training samples, including: Collecting user initial conversation records under preset scenarios, marking interaction intention labels and design requirement fields in the user initial conversation records, and generating interaction requirement guidance samples; extracting natural language instructions without semantic modification from the interaction demand guidance sample as original interaction instructions; Building a domain-related synonym library, wherein the synonym library includes synonyms of design style descriptors, color coding parameters and typography terms; identifying a target replacement vocabulary in the original interaction instruction, and generating a synonym replacement candidate set of the target replacement vocabulary based on the synonym library; According to a preset replacement probability threshold, selecting replacement items of the target replacement vocabulary from the synonymous replacement candidate set to generate a synonymous replacement perturbation sample; Detecting time description sentences in the original interaction instructions, and converting the absolute time description in the time description sentences into relative time period descriptions, to generate time fuzzy generalized perturbation samples; Identifying spatial position description sentences in the original interaction instructions, and replacing coordinate parameters in the spatial position description sentences with directional word descriptions to generate spatial fuzzy generalized perturbation samples; The synonymous replacement disturbance sample, the time fuzzy generalization disturbance sample, the space fuzzy generalization disturbance sample and the original interaction instruction are mixed in a preset ratio to generate the disturbance training sample.

3. The visual generation method based on semantic association modeling according to claim 1, characterized in that: Receiving a real-time interaction instruction, inputting the real-time interaction instruction into the fine-tuned language model, and generating a demand text including a topic identification field and style level parameters, including: Send initial guidance instructions to the interactive end; receiving a first interaction instruction returned by the interaction terminal, and parsing a subject keyword in the first interaction instruction; Generating an initial candidate set of subject identification fields according to the subject keywords; Calling a preset style guidance template to send a style level selection instruction to the interactive terminal, wherein the style level selection instruction includes a plurality of preset design style options; receiving a second interaction instruction returned by the interaction terminal, and verifying whether the design style type in the second interaction instruction belongs to the preset design style option; Extracting associated attribute parameters from the verified design style type, and generating a style level parameter set based on the associated attribute parameters; The initial candidate set of the subject identification field and the style level parameter set are integrated into a structured requirement text.

4. The visual generation method based on semantic association modeling according to claim 1, characterized in that: Extracting semantic features of elements in the requirement text and determining association weights between the semantic features of the elements as semantic association weights include: Parsing all semantic elements from the demand text, and inputting the semantic elements into a pre-trained semantic encoder to generate a semantic feature vector for each semantic element; Determine the cosine similarity value between each two semantic feature vectors and generate a semantic similarity matrix; performing normalization processing on the semantic similarity matrix to generate a semantic association weight matrix; The association weight items in the semantic association weight matrix that are lower than a preset weight threshold are filtered to generate a filtered semantic association weight set.

5. The visual generation method based on semantic association modeling according to claim 1, characterized in that: Combining the semantic association weight with a preset spatial distribution constraint item to construct a layout optimization function, and optimizing the layout optimization function by backpropagation to generate element coordinate data, including: generating layout guiding coordinates by a layout planning agent module, wherein the layout guiding coordinates are determined based on a preset layout strategy or a historical layout pattern; Obtaining a Gaussian kernel function of a spatial distribution constraint term, wherein the Gaussian kernel function determines a spatial correlation strength between the two semantic elements based on a ratio of a square of an actual distance between the two semantic elements and a preset bandwidth parameter; Constructing a semantic space coordination term of the layout optimization function, wherein the semantic space coordination term is obtained by multiplying the semantic association weight of each semantic element pair by the output value of the corresponding Gaussian kernel function and then summing the results; Constructing a layout guidance matching item of the layout optimization function, wherein the layout guidance matching item is obtained by determining the sum of squared Euclidean distances between the coordinates of all semantic elements and the layout guidance coordinates and multiplying the sum by a preset balance weight parameter; Adding the semantic space coordination item and the layout guidance matching item to generate a final layout optimization function; Randomly initialize the two-dimensional coordinate values ​​of all semantic elements to generate an initial layout coordinate set; Perform multiple rounds of iterative optimization on the initial layout coordinate set using a back-propagation algorithm, performing a gradient calculation operation in each iteration, wherein the gradient calculation operation is to calculate the partial derivative value of the layout optimization function with respect to the horizontal and vertical coordinates of each semantic element; According to the output result of the gradient calculation operation, the coordinate values ​​of all semantic elements are updated along the reverse direction of the gradient according to a preset step size; Determine whether the updated layout optimization function value meets the preset convergence condition; When the updated layout optimization function value meets the preset convergence condition, the optimized semantic element coordinate data is output.

6. The method for visual generation based on semantic association modeling according to claim 1, wherein: Inputting the layout control matrix and the initial noise tensor into a codec network, fusing the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjusting the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generating target visual content, including: Concatenate the layout control matrix and the initial noise tensor along the channel dimension to generate a fused input tensor, and input the fused input tensor into the encoding and decoding network; Performing a multi-scale downsampling operation on the fused input tensor through the encoder of the codec network to generate a multi-level encoding feature map; Inputting the multi-level encoding feature map into the attention gating module of the decoder of the encoding and decoding network for cross-layer feature aggregation to generate a regional focus weight matrix, wherein each element value of the regional focus weight matrix represents the attention intensity of the corresponding spatial position in the noise reduction process; Initializing the intermediate image to the initial noise tensor; Target visual content is generated by performing multi-step denoising iterations through a diffusion model sampling strategy; In each round of denoising iteration, a noise prediction output is generated based on the intermediate image at the current time step and the conditional input of the layout control matrix, the regional focus weight matrix is ​​element-wise multiplied by the noise prediction output to generate an adjusted noise prediction result, the intermediate image is updated according to the adjusted noise prediction result, and the regional focus weight matrix is ​​scaled according to the noise level at the current time step; When the number of iterations reaches the preset number of sampling steps, the target visual content is output.

7. The visual generation method based on semantic association modeling according to claim 1, characterized in that: Inputting the layout control matrix and the initial noise tensor into a codec network, fusing the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjusting the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generating the target visual content, further comprising: receiving adjustment parameters returned by the feedback end, wherein the adjustment parameters are generated based on the layout score and semantic relevance annotation result of the target visual content by the user; Extracting the historical task contributions of the language model parameters from a pre-stored parameter importance matrix, and constructing an elastic weight curing loss function in combination with the adjustment parameters; Based on the elastic weight solidification loss function, the gradient direction and update amount of the parameters of the language model are calculated by a back propagation algorithm to generate an updated language model; The real-time interaction instructions are processed through the updated language model to generate optimized visual content aligned with the user feedback.

8. A visual generation device based on semantic association modeling, characterized in that: The visual generation device based on semantic association modeling includes: An interaction data construction module is used to construct a dialogue dataset containing interaction demand guidance samples, extract original interaction instructions from the dialogue dataset, and perform semantic perturbation operations on the original interaction instructions to generate perturbation training samples; A language model fine-tuning module is used to input the perturbation training sample into a pre-trained language model and fine-tune the parameters of the pre-trained language model through a low-rank adaptation matrix to generate a fine-tuned language model; A demand text generation module, configured to receive a real-time interaction instruction, input the real-time interaction instruction into the fine-tuned language model, and generate a demand text including a topic identification field and style level parameters; A semantic association modeling module is used to extract semantic features of elements in the requirement text and determine the association weights between the semantic features of each element as semantic association weights; a layout optimization analysis module, configured to combine the semantic association weights with preset spatial distribution constraints to construct a layout optimization function, and generate element coordinate data by optimizing the layout optimization function through backpropagation; a layout control fusion module, configured to encode the element coordinate data into a spatial position vector, and fuse the spatial position vector with the element semantic features of the requirement text to generate a layout control matrix; The visual content generation module is used to input the layout control matrix and the initial noise tensor into the codec network, fuse the layout control matrix and the initial noise tensor through the codec network to generate a regional focus weight matrix, adjust the noise reduction process of the codec network based on the regional focus weight matrix, and iteratively generate target visual content.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a visual generation program based on semantic association modeling stored in the memory and capable of running on the processor. When the visual generation program based on semantic association modeling is executed by the processor, the steps of the visual generation method based on semantic association modeling as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a visual generation program based on semantic association modeling, and when the visual generation program based on semantic association modeling is executed by a processor, the steps of the visual generation method based on semantic association modeling as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Content generation method and system based on multi-modal large model

    CN120832885A

  • IPTV interface element differentiation display method and system

    CN120935418A

  • Object interaction analysis method and device based on visual features, equipment and medium

    CN120997743A

  • Object interaction analysis method and device based on visual features, equipment and medium

    CN120997743B

  • Fire control room graph display device and automatic graph importing method and device thereof

    CN121188020A