Model training method, scene graph generation method, equipment, storage medium and program product
By introducing an artificial intelligence evaluation model into the scene image generation process and optimizing the parameters of the text generation model, the problems of high cost and poor quality in generating product main images in existing technologies have been solved, and efficient and aesthetically pleasing artificial intelligence-generated scene images have been achieved.
Patent Information
- Application Number
- CN202511645261.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-10
- Publication Date
- 2026-02-13
AI Technical Summary
In existing technologies, the process of generating product main images requires the participation of a professional team, which is labor-intensive and inefficient. Furthermore, the quality of existing text prompts varies, making it difficult to generate high-quality, aesthetically pleasing images.
By introducing artificial intelligence technology and using an evaluation model to assess scene images from a spatial aesthetic perspective, the parameters of the text generation model are optimized in reverse to generate scene prompts that meet spatial aesthetic requirements, thereby improving the quality of scene images.
It reduces labor costs and improves generation efficiency during scene graph generation, while generating high-quality, aesthetically pleasing scene graphs that conform to human aesthetic preferences.
Smart Images

Figure CN121527593A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, a scene graph generation method, a device, a storage medium, and a program product. Background Technology
[0002] In e-commerce, product images play a crucial role as the first visual point of contact for users browsing products. High-quality product images, such as clear pictures, simple designs, and distinctive visual effects, can significantly increase click-through rates, thereby improving product conversion rates.
[0003] In traditional methods, merchants design scenes and props, then photograph the products within those scenes as the main product images. This approach requires a professional team, resulting in high labor costs, low efficiency, and the final shooting effect may not meet expectations due to limitations in shooting conditions.
[0004] With the development of artificial intelligence technology, models capable of generating images based on text have emerged. By inputting text prompts into the model, it can efficiently generate product main images guided by those prompts. While these models improve the efficiency of product main image generation, they require the creation of text prompts, and the quality of these prompts affects the quality of the product main image. Therefore, how to generate high-quality text prompts is a problem that needs to be solved. Summary of the Invention
[0005] This application provides a model training method, a scene graph generation method, a device, a storage medium, and a program product, which can generate scene prompts that meet spatial aesthetic requirements, thereby improving the quality of scene graphs.
[0006] This application provides a model training method, including: inputting a sample image containing a sample object into a text generation model, generating scene prompts using the current model parameters to obtain sample scene prompts corresponding to the sample object; inputting the sample image and sample scene prompts into an image generation model to generate a sample scene image containing the sample object; evaluating the sample scene image from a spatial aesthetic perspective using an evaluation model, and generating a reward signal based on the evaluation results; and updating the model parameters of the text generation model based on the reward signal so that the text generation model generates scene prompts that meet the requirements of spatial aesthetics.
[0007] This application provides a scene graph generation method, including: acquiring a target image including a target object; inputting the target image into a text generation model to generate scene prompts, so as to obtain target scene prompts corresponding to the target image, wherein the target scene prompts describe the target scene information to be generated by the model in text form; inputting the target image and the target scene prompts into an image generation model to generate a target scene graph containing the target object and the target scene information.
[0008] This application embodiment also provides a scene graph generation method, including: displaying an image generation interface, the image generation interface including an image upload control; in response to a trigger operation on the image upload control, obtaining a target image uploaded by a user, the target image including a target object; inputting the target image into a text generation model to generate scene prompt words, so as to obtain target scene prompt words corresponding to the target image, the target scene prompt words describing the target scene information to be generated by the model in text form; inputting the target image and the target scene prompt words into the image generation model to generate a target scene graph containing target object and target scene information.
[0009] This application also provides a computing device, including: a memory and a processor; wherein, the memory stores executable code, and when the executable code is executed by the processor, the processor performs the steps in the model training method or the steps in the scene graph generation method.
[0010] This application also provides a computer-readable storage medium storing executable code. When the executable code is executed by a processor of a computing device, the processor performs steps as described in the model training method or the scene graph generation method.
[0011] This application also provides a computer program product, including: a computer program / instructions, which, when executed by a processor, enable the processor to implement the steps in the model training method or the steps in the scene graph generation method.
[0012] In this embodiment, sample scene prompts are generated based on sample images containing sample objects. Based on the sample images and prompts, an image generation model generates a sample scene image containing the sample objects. The sample scene image is then evaluated from a spatial aesthetic perspective. The model parameters of the text generation model are updated based on the evaluation results. Specifically, by using the evaluation results of the final generated sample scene image to back-optimize the model parameters of the text generation model, the model generates scene prompts that meet spatial aesthetic requirements, thus improving the quality of the scene prompts. Furthermore, during the inference phase, guided by the scene prompts, the image generation model can generate scene images that better meet spatial aesthetic requirements, further improving the quality of the scene images. Attached Figure Description
[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A schematic flowchart of a model training method provided for an exemplary embodiment of this application; Figure 2 A schematic flowchart of a model training method provided in yet another exemplary embodiment of this application; Figure 3 A schematic diagram of the architecture of a text generation model provided for an exemplary embodiment of this application; Figure 4 A schematic diagram of the architecture of an image generation model provided for an exemplary embodiment of this application; Figure 5 A schematic diagram of the architecture of a first evaluation model provided for an exemplary embodiment of this application; Figure 6 A schematic diagram of the architecture of a second evaluation model provided for an exemplary embodiment of this application; Figure 7 A schematic diagram of the architecture of a third evaluation model provided for an exemplary embodiment of this application; Figure 8 A scene diagram illustrating a scene graph generation method provided as an exemplary embodiment of this application; Figure 9 A schematic diagram of the structure of a model training apparatus provided in an exemplary embodiment of this application; Figure 10 A schematic diagram of the structure of a scene graph generation apparatus provided in an exemplary embodiment of this application; Figure 11 This is a schematic diagram of the structure of a computing device provided for an exemplary embodiment of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0016] In e-commerce, the main product image plays a crucial role as the first visual point of contact for users browsing products. High-quality main product images (or scene images), such as clear pictures, simple designs, and differentiated visual effects, can significantly increase click-through rates, thereby improving product conversion rates. Traditionally, merchants design scenes and props, then photograph the product within those scenes as the main product image. This method requires a professional team, resulting in high labor costs, low efficiency, and the final shooting effect may not meet expectations due to limitations in shooting conditions.
[0017] With the development of artificial intelligence (AI) technology, models capable of generating images based on text have emerged. By inputting text prompts into the model, it can efficiently generate product main images guided by the prompts. While these models improve the efficiency of main image generation, they require the construction of text prompts, and the quality of the text prompts affects the quality of the product main image.
[0018] Current methods for constructing prompts mainly include the following: One approach is to provide different preset prompt templates for merchants to choose from. However, these preset templates often suffer from inconsistent quality and a lack of specificity, making it difficult to meet diverse product display needs. Another approach allows merchants to write their own text prompts. However, manually writing high-quality prompts requires deep design knowledge and a profound understanding of AI models, which is too high a barrier for ordinary merchants. Merchants cannot independently complete the prompt writing and need to invest additional time and operational costs to seek professional support. A third approach uses a text model to automatically expand the text input by merchants into long prompts. In this method, the text model is obtained through supervised learning based on a pre-built "query-prompt" dataset. The model learns to imitate the ability of "good human-written prompts," and the effectiveness of long prompts is limited by the quality and size of the dataset. Furthermore, since its optimization target is text-level prompts and does not incorporate image feedback or human preference signals, the generated images, while conforming to the prompt description, still differ from real human subjective preferences in an aesthetic dimension. This requires a complex "fine-tuning" process for calibration, which is complex and inefficient.
[0019] In summary, the current problem to be solved is how to construct a text prompt that matches the characteristics of the target object and guides the image generation model to produce high-quality, aesthetically pleasing images.
[0020] To address this technical problem, in this embodiment, during the scene image generation process, on the one hand, artificial intelligence (AI) technology is introduced to help optimize the scene image generation process, quickly generate scene images, reduce shooting costs, and improve scene image generation efficiency. On the other hand, the text generation model used in the scene image generation process is trained. During model training, the evaluation results of the scene image are used as a reward signal to reverse-optimize the process of generating text prompts, thereby constructing a text prompt that matches the characteristics of the target object and guides the image generation model to produce high-quality, aesthetically pleasing images.
[0021] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0022] Figure 1 A schematic flowchart of a model training method provided for an exemplary embodiment of this application is shown below. Figure 1 As shown, the method includes: Step 11: Input the sample image containing the sample object into the text generation model, and use the current model parameters to generate scene prompts to obtain the sample scene prompts corresponding to the sample object.
[0023] Step 12: Input the sample image and sample scene prompts into the image generation model to generate a sample scene image containing the sample objects.
[0024] Step 13: Use the evaluation model to evaluate the sample scene image from the perspective of spatial aesthetics, and generate a reward signal based on the evaluation results.
[0025] Step 14: Update the model parameters of the text generation model according to the reward signal so that the text generation model generates scene prompts that meet the requirements of spatial aesthetics.
[0026] In this embodiment, a sample image refers to an image containing a sample object. The sample object can be any object, such as an animal, plant, building, person, etc., or various goods that can be sold or displayed on e-commerce platforms, such as vehicles, mobile phones, tables, sofas, clothing, sporting goods, etc. This embodiment does not impose any limitations.
[0027] The sample scene prompts describe the sample scene information that the model needs to generate in text form. The sample scene diagram includes sample objects and sample scene information. The sample objects are the main objects in the sample scene diagram. The scene diagram is an image that visually displays a product in scenarios such as online e-commerce, product display, or related visual tasks; it can also be called a main scene image, display image, etc. Scene information can be understood as the elements in the image other than the object (such as the product), which together constitute the image's context and background. The scene diagram can be of any style and type. The style can be modern, minimalist, pastoral, or industrial; the functional type can be a living room, bedroom, kitchen, or office. For example, the scene diagram can be a modern bedroom, a pastoral kitchen, or an industrial office, etc., and this embodiment is not limited. This embodiment can generate scene diagrams of diverse home decoration scenes and styles, enhancing the visual richness of product displays and user experience.
[0028] Before training the model, a sample dataset can be constructed, which includes multiple sample images. The electronic device responsible for model training can obtain the sample dataset. This embodiment does not limit the specific implementation of the acquisition operation. Optionally, the electronic device can directly obtain the uploaded sample dataset in response to an upload operation for the sample dataset, or the electronic device can obtain the sample dataset from a cloud server, or the electronic device can download the sample dataset via a URL link.
[0029] After obtaining the sample dataset, the text generation model can be trained based on the sample dataset. For example, such as... Figure 2As shown, a text generation model can be trained using reinforcement learning based on a sample dataset. The text generation model acts as the agent, while the image generation and evaluation models act as the environment. Through interaction between the agent and the environment, reward signals are obtained by evaluating sample scene images. Based on these reward signals, better model parameters are learned, with the goal of maximizing long-term cumulative rewards or reaching a set threshold. During training, the model parameters of the image generation and evaluation models remain unchanged; only the parameters of the text generation model are optimized.
[0030] The text generation model can be any neural network model capable of simultaneously understanding images and text. Optionally, the text generation model can be a deep learning model with a relatively large number of parameters. However, "large model" is merely an example, and this application does not limit the number of parameters supported by the deep learning model, aiming to meet actual needs. For example, the text generation model can be a Vision-Language Model (VLM) or a Multimodal Large Language Model (MLLM), etc. In this application embodiment, a sample image is input into the text generation model, and the text generation model can output sample scene prompts corresponding to the sample object based on the current model parameters.
[0031] The image generation model can be a deep generation model that generates high-quality images based on a denoising process. Guided by input generation conditions, it can generate images that meet the requirements of those conditions. In this embodiment, a sample image and sample scene prompts are input into the image generation model. Guided by the sample image and the sample scene prompts, the image generation model can generate a sample scene image containing the sample object.
[0032] The evaluation model can assess multiple quantifiable dimensions, such as aesthetic fit, semantic alignment, and scene richness, from a spatial aesthetic perspective. The model has undergone expert knowledge calibration, making it more aesthetically pleasing and closer to real human subjective preferences, ensuring the accuracy of the evaluation and that the results align with human aesthetics. Furthermore, a reward signal is generated based on the evaluation results, and the model parameters of the text generation model are updated according to this signal. This allows the trained text generation model to possess design sense and aesthetic ability, generating scene prompts that not only match product characteristics but also align with human aesthetics. These scene prompts meet spatial aesthetic requirements and demonstrate creativity and professionalism. The scene prompts are used as input to the image generation model to guide the generation of scene images. Guided by scene prompts aligned with human aesthetics, the scene images generated by the image generation model also align with human aesthetics, making them more aesthetically pleasing and closer to real human subjective preferences.
[0033] In the model training process of this application embodiment, instead of evaluating the quality of sample scene prompts, the model uses sample images and sample scene prompts as guidance to generate sample scene graphs containing sample objects using an image generation model. Then, the concrete sample scene graphs are evaluated from a spatial aesthetic perspective. Evaluating whether a sample scene graph meets spatial aesthetic requirements is far more objective and accurate than judging whether a sample scene prompt "can generate a beautiful image." By using the evaluation results of the sample scene graphs to back-optimize the model parameters of the text generation model, the model can not only automatically generate scene prompts based on the input image, but also generate scene prompts that meet spatial aesthetic requirements, thus improving the quality of the scene prompts. This "what you see is what you get" reward mechanism, combined with multi-dimensional aesthetic evaluations such as semantic alignment, scene richness, and aesthetic fit, can significantly improve the spatial aesthetic performance of the text generation model. Furthermore, guided by scene prompts that meet spatial aesthetic requirements, the scene graphs generated by the image generation model also meet spatial aesthetic requirements.
[0034] This application does not limit the specific implementation of the text generation model. In some optional embodiments, step 11 above can be implemented in the following way: inputting a sample image containing the sample object into the text generation model, using the multimodal sequence generated based on the sample image as a guiding condition, performing scene prompt word generation processing on the sample image to obtain the sample scene prompt words corresponding to the sample object.
[0035] In this embodiment of the application, the text generation model has the following capabilities: it can perform semantic parsing on the input image, generate a multimodal sequence as a guiding condition, and then perform deep understanding and linguistic reconstruction of the image content based on the multimodal sequence, and output scene prompt words that match the objects in the image.
[0036] After inputting sample images into the text generation model, the model first perceives and encodes the images, extracting deep semantic features, including category, appearance attributes, pose, scale, and local contextual information. Then, based on these deep semantic features, the model generates a multimodal sequence representation containing object ontology features and scene context. This multimodal sequence serves as a guiding condition, integrating visual structure and latent language concepts, to regulate the subsequent language generation process. Guided by the multimodal sequence, the model dynamically focuses on key visual elements in the image and correlates them with scene common sense and semantics learned by the model, gradually generating accurate, coherent, and scene-aware sample scene prompts. This ensures that the output text not only describes the explicit objects in the image but also reflects their environment, interaction relationships, and overall semantic context.
[0037] In some optional embodiments, the aforementioned embodiment's "inputting a sample image containing the sample object into a text generation model, using a multimodal sequence generated based on the sample image as a guiding condition, and performing scene cue word generation processing on the sample image to obtain the sample scene cue words corresponding to the sample object" can be implemented based on the following steps S1-S4: S1, the first image encoder of the text generation model is used to extract semantic features from the sample image to obtain the first visual features.
[0038] In this text generation model, the first image encoder is used to convert the input image into a representation that can be understood and processed by the machine, namely, a feature map or feature vector. This application does not limit the specific implementation of the first image encoder. For example, after inputting a sample image into the first image encoder, the encoder can divide the sample image into several image blocks and, through multi-layer attention encoding, finally output a feature vector representing the semantic content of the image, referred to as the first visual feature.
[0039] S2 projects the first visual features into the embedding space of the language model of the text generation model.
[0040] Since the feature dimension output by the first image encoder (e.g., 1024) is inconsistent with the embedding dimension of the language model in the text generation model (e.g., 4096), in order for the language model to understand image information, a learnable projection layer can be used to map the first visual features to the embedding space of the language model. This allows the visual features and language features to reside in the same semantic space and be processed uniformly by the language model. For ease of distinction, the language model of the text generation model is referred to as the first language model.
[0041] S3, in the embedding space of the language model of the text generation model, concatenates the projected first visual features with the first text features to generate a multimodal sequence.
[0042] Because text generation models require textual context to define the task, when only a sample image is input without displayed text input, the model automatically constructs an implicit text instruction as text input to guide the generation process. This implicit text instruction can include image placeholders (such as `img`) and task instructions (such as generating a scene cue word). During model processing, the implicit text instruction can be segmented to obtain token IDs. These token IDs are then input into the embedding layer of the first language model to generate the corresponding first text feature. This first text feature can be concatenated with the first visual feature to form a multimodal sequence, which will be referred to as the first multimodal sequence for ease of distinction. Furthermore, the text feature itself resides in the embedding space of the language model and does not require additional projection into the language model's embedding space.
[0043] Optionally, in addition to sample images, the sample dataset may also include initial prompts corresponding to the sample images. The initial prompts can describe the desired scene style, objects within the scene, the positions of sample objects, and the relative positions between objects. During model training, the sample images and initial prompts can be input into the text generation model. The text generation model can perform word segmentation on the initial prompts to obtain lexical numbers, which are then input into the embedding layer of the language model to generate the corresponding first text features. In this implementation, the text generation model has explicit text input, eliminating the need for implicit text instructions to generate the first text features. During the inference phase, the initial prompts can be input by the user to represent the user's expectations for the target scene.
[0044] S4. Using a language model and guided by multimodal sequences, scene prompts are generated for the sample images to obtain sample scene prompts that are semantically consistent with the sample images.
[0045] The language model takes the first multimodal sequence as input and generates scene cue words word by word through an autoregressive approach. During the generation process, the model uses semantic information from the first visual features to understand the image content and combines it with the first text features to generate a coherent, accurate scene description that is semantically consistent with the sample image; this is called the sample scene cue word.
[0046] The following will combine Figure 3 The architecture and processing flow of the above text generation model are illustrated with examples.
[0047] like Figure 3 As shown, the text generation model may include: a first image encoder, a first multimodal alignment module, and a first language model.
[0048] The first image encoder is responsible for converting the input sample image into first visual features rich in semantic information. The first image encoder may include an embedding layer, a positional encoding module, and a multi-layer encoder.
[0049] The embedding layer of the first image encoder divides the input sample image into fixed-size non-overlapping image patches, each of which is linearly projected into a multi-dimensional vector. Subsequently, the positional encoding module of the first image encoder adds learnable two-dimensional positional codes to the multiple image patch vectors to preserve the original spatial structure information. The resulting image patch label sequence is then input into a backbone network consisting of multiple stacked encoders. Each encoder layer contains a multi-head self-attention mechanism and a feedforward neural network. Through the multi-head self-attention mechanism and the feedforward neural network (FNN), local details and global semantics of the image can be gradually extracted.
[0050] The multi-head self-attention mechanism models the global dependencies between multiple image patches using multiple attention heads, achieving semantic fusion of local details (such as texture and edges) and overall structure (such as contours). The feedforward neural network introduces non-linear expressive power through two-layer linear transformation and Gaussian error linear unit activation functions, enhancing feature discriminative power. Both multi-head self-attention and the feedforward neural network incorporate residual connections and layer normalization to ensure effective information transfer and training stability. After layer-by-layer feature extraction, a dense feature sequence composed of multiple visual markers is finally output. Each marker encodes fine-grained semantic information about the corresponding image region in terms of shape, material, color, orientation, and proportion, thus transforming the original pixel matrix of the sample image into structured first visual features. This lays the foundation for subsequent multimodal alignment and understanding with the language model in a unified semantic space.
[0051] The first multimodal alignment module is the core bridge connecting the visual and first language modalities, responsible for aligning the first visual features output by the first image encoder with the text semantic space of the first language model. For example, the first multimodal alignment module may include a learnable linear projection layer (MLP). The projection layer transforms the first visual features into vectors with the same embedding dimension as the first language model through a learnable linear (or non-linear) mapping. This process is equivalent to "translating" visual information into "language" that the language model can understand, thereby enabling visual information to participate in the reasoning process of the first language model.
[0052] The first language model is responsible for outputting structured sample scene prompts that conform to a predefined format. The first language model can include a pre-defined tokenizer, an embedding layer, a positional encoding module, and a self-attention layer. The pre-defined tokenizer segments the implicit text instructions or initial prompts into a series of discrete tokens and converts them into corresponding token codes. These token codes are fed into the embedding layer of the first language model, which maps the token codes into continuous high-dimensional vectors to obtain the first text features, thus transforming the original text into a numerical semantic representation that the first language model can understand and process. Subsequently, in the embedding space of the first language model, the first visual features and the first text features are concatenated in the sequence dimension to form a multimodal sequence (also called a multimodal embedding representation). The positional encoding module of the first language model adds positional information to the concatenated sequence, ensuring that the position of each token (whether from text or image) in the sequence can be perceived.
[0053] The self-attention layer of the first language model consists of multiple stacked decoders, each containing a multi-head self-attention mechanism and a feedforward neural network. When a multimodal sequence is input into the self-attention layer, the dependencies between positions in the multimodal sequence are dynamically calculated through the self-attention mechanism. In each layer, each token only focuses on all tokens to its left (including itself), thus achieving deep fusion and cross-modal alignment of visual information and linguistic context in each layer. Based on this, in each generation step, the first language model predicts the next most likely token step by step based on the existing multimodal context, and iterates until the end token is generated, finally outputting a sample scene cue word consistent with the image semantics. For example, when the sample image includes a table, the generated sample scene cue word could be "A table is placed in a restaurant environment, with decorations on the walls, green plants along the walls, and a light-colored floor, which complements the modern and open design of the room."
[0054] Optionally, when generating sample scene prompts, the text generation model can also use the current model parameters to perform positional inference on the sample object to obtain its positional information in the sample scene graph. When the sample dataset includes initial prompts, and the initial prompts contain positional information, the text generation model can perform positional inference based on the input positional information to obtain the sample object's positional information in the sample scene graph. When the input to the text generation model does not include positional information, the text generation model can infer the sample object's positional information in the sample scene by learning the strong correlation between visual regions and spatial descriptions in the training data. Optionally, the positional information of the sample object in the sample scene graph can also be included within the sample scene prompts, as part of the sample scene prompts. During the inference phase, the initial prompts can be input by the user.
[0055] This application does not limit the specific implementation of the aforementioned image generation model. In some optional embodiments, step 12 can be implemented as follows: inputting the sample image and sample scene prompts into the image generation model, using the sample image and sample scene prompts as noise prediction conditions, and denoising the noise tensor to obtain the sample scene image.
[0056] The image generation model can adopt a denoising generation architecture based on a diffusion mechanism. The principle is as follows: the input is converted into the corresponding conditional embedding; multiple conditional embeddings drive the image generation model to add and remove noise, and finally the scene map is output through the decoder.
[0057] In this embodiment, the image generation model extracts visual features corresponding to the sample image and text features corresponding to the prompt words in the sample scene. The model maintains an initial noise tensor and iteratively predicts and removes the noise across multiple time steps. In each denoising step, the image generation model relies not only on the current state of the noisy image but also on the visual priors provided by the visual features and the semantic content expressed by the text features. Guided by these visual and text features, it predicts the noise components that should be removed. After several rounds of denoising iterations, the initial noise tensor gradually transforms into a sample scene image that conforms to the prompt semantics.
[0058] In some optional embodiments, the above embodiment of "inputting the sample image and sample scene prompts into the image generation model, using the sample image and sample scene prompts as noise prediction conditions, and denoising the noise tensor to obtain the sample scene image" can be implemented based on the following steps R1-R4: R1 uses the first text encoder of the image generation model to encode sample scene cue words into text conditional embeddings.
[0059] The first text encoder can be implemented using a pre-trained text model or other models, such as a multimodal encoding model. This embodiment does not impose any restrictions.
[0060] The text pre-trained model can be used to capture the semantic information of natural language (i.e., sample scene prompts) and map it into a high-dimensional vector representation, namely the semantic vector of the sample scene prompt, which is then used as a conditional embedding of the text. When encoding sample scene prompts from any image, the text pre-trained model can be implemented based on the following steps: preprocessing the sample scene prompts, including word segmentation, standardization, and padding or truncation operations, to generate text prompts that meet the input requirements of the text pre-trained model; loading a word segmenter and using it to encode the preprocessed sample scene prompts to generate an input tensor; extracting hidden representations of the sample scene prompts from the input tensor; and extracting the semantic vectors of the sample scene prompts from these hidden representations.
[0061] In this context, a multimodal coding model refers to a model capable of processing and fusing data from multiple different information sources (i.e., modalities). These modalities can include text, images, audio, video, etc. A multimodal coding model also possesses unimodal coding capabilities for any single modality. Therefore, unimodal coding capabilities can be used to perform vector embedding based on the input sample scene prompts to output the text conditional embeddings corresponding to the sample scene prompts.
[0062] R2 uses a second image encoder from an image generation model to encode sample images into visual conditional embeddings.
[0063] The second image encoder in the image generation model is used to convert the input image into a representation that can be understood and processed by a machine, namely a feature map or feature vector. This application does not limit the specific implementation of the second image encoder. In general, the second image encoder is a visual feature extraction module that can encode sample images into visual conditional embedding vectors aligned with text embeddings.
[0064] R3 initializes the random noise tensor in the latent space.
[0065] R4, based on the conditional diffusion network in the image generation model, uses text conditional embedding and visual conditional embedding as noise prediction conditions to denoise the random noise tensor to obtain the target latent representation.
[0066] R5 uses the decoder of the image generation model to decode the latent representation of the target in order to obtain the sample scene map.
[0067] The conditional diffusion network (CDN) is an image generation network based on a diffusion process. It possesses the following capability: guided by the input generation conditions, it continuously denoises a random noise vector, ultimately obtaining a feature vector that meets the generation conditions. In this embodiment, the text conditional embeddings corresponding to the sample scene prompts and the visual conditional embeddings corresponding to the sample images are used as the generation conditions of the CDN. These are input into the CDN, which starts with an initial noise tensor randomly generated in the latent space. Guided by the text and visual conditional embeddings, it gradually denoises through a series of steps. Specifically, the CDN predicts the noise to be removed in each step and performs corresponding denoising, thereby inversely recovering the target latent representation that meets the given conditions. This target latent representation is then decoded by a decoder, and the sample scene image is finally obtained through row upsampling and feature reconstruction. In this embodiment, the number of times the random noise vector is denoised is not limited; for example, it can be 30-50 times. Text and visual conditional embeddings are continuously injected into the noise vector during each denoising process, resulting in a feature vector with rich semantic information.
[0068] The following will combine Figure 4 The architecture and processing flow of the above image generation model are illustrated with examples.
[0069] like Figure 4 As shown, the image generation model may include: a first text encoder, a second image encoder, a conditional diffusion network, and a decoder.
[0070] On one hand, a sample scene cue phrase, "A table placed in a restaurant environment, with decorations on the walls, green plants along the walls, and a light-colored floor, complementing the room's modern and open design," can be input into the first text encoder to obtain coded text (i.e., text conditional embedding). On the other hand, a sample image can be input into the second image encoder for vector embedding to obtain the corresponding visual conditional embedding. Additionally, a completely random noise tensor is generated in the latent space. Then, the text conditional embedding and the visual conditional embedding can be input together into a conditional diffusion network. After denoising the noise tensor through the conditional diffusion network, the target latent representation corresponding to the sample scene image can be obtained. This target latent representation is then input into the decoder to obtain the sample scene image. For example, by... Figure 4 As can be seen, the input sample image only includes a table, while the processed sample scene image includes scene information, such as added chairs, wall hangings, green plants, and windows, presenting an overall modern style.
[0071] The initial noise tensor has the same shape as the sample scene image in the latent space. The shape of the noise tensor is set according to the size of the sample scene image. The image generation model first determines the spatial dimensions (such as height and width) of the latent representation of the sample scene image in the latent space, and then initializes the noise tensor with the corresponding shape.
[0072] In some implementations, the size of the sample scene image can be a default size, for example, the same as the size of the input sample image.
[0073] In other implementations, the sample dataset includes the dimensions of the sample scene graph corresponding to the sample image. The dimensions of the sample scene graph can serve as another input condition for the image generation model. In this implementation, a size encoding module can be added to the architecture of the image generation model. This size encoding module can encode the size information into size-conditional embeddings that the model can understand. Encoding methods include, but are not limited to: normalizing the height and width and concatenating them into a vector, then mapping it through a multilayer perceptron (MLP) to a size-conditional embedding aligned with the text embedding dimension, or converting the height and width into a learnable positional encoding form. The workflow of the image generation model is as follows: input the dimensions of the sample image, sample scene cue words, and sample scene graph into the image generation model; in the image generation model, the first text encoder encodes the sample scene cue words into text-conditional embeddings, the second image encoder encodes the sample image into visual-conditional embeddings, and the size encoding module encodes the dimensions of the sample scene graph into size-conditional embeddings; in the latent space, the spatial dimension of the target latent representation is calculated based on the size-conditional embeddings, and a random noise tensor is initialized on this spatial dimension as the starting point for the denoising process. Based on a conditional diffusion network, textual conditional embedding, visual conditional embedding, and size conditional embedding are used as noise prediction conditions to denoise the random noise tensor to obtain the latent representation of the target. The decoder maps the latent representation of the target back to the pixel space through row upsampling and feature reconstruction, and outputs a sample scene map.
[0074] Optionally, the aforementioned location information can also serve as another input condition for the image generation model. In this implementation, a location encoder can be added to the architecture of the image generation model to convert the raw location information into a location-conditional embedding that the model can understand. The location encoder can be implemented in ways including, but not limited to: first, inputting the raw location coordinates (e.g., x, y values) into a multilayer perceptron and generating a fixed-dimensional location-conditional embedding through nonlinear transformation; or using a sinusoidal location encoding method, mapping each coordinate dimension using high-frequency trigonometric functions and then projecting to obtain the location-conditional embedding; or discretizing the location into region indices in an image grid and using a learnable embedding table for lookup encoding. In implementation, the image generation model's workflow is as follows: Sample images, sample scene cues, and the size and location information of the sample scene image are input into the image generation model. Within the model, a first text encoder encodes the sample scene cues into text conditional embeddings, a second image encoder encodes the sample images into visual conditional embeddings, a size encoding module encodes the size of the sample scene image into size conditional embeddings, and a location encoder encodes the location information into location conditional embeddings. In the latent space, the spatial dimension of the target latent representation is calculated based on the size conditional embeddings. A random noise tensor is initialized in this spatial dimension as the starting point for the denoising process. Based on a conditional diffusion network, text conditional embeddings, visual conditional embeddings, size conditional embeddings, and location conditional embeddings are used as noise prediction conditions to denoise the random noise tensor to obtain the target latent representation. The decoder maps the target latent representation back to the pixel space through row upsampling and feature reconstruction, outputting the sample scene image.
[0075] The sample scene image can be obtained through steps 11 and 12 above. Further, the evaluation model can be used to evaluate the sample scene image from a spatial aesthetic perspective, and a reward signal can be generated based on the evaluation results to update the model parameters of the text generation model. This application does not limit the specific implementation method of the evaluation model.
[0076] In some optional embodiments, a first evaluation model can be used to evaluate the aesthetic fit of the sample scene image, obtain a first evaluation result, and generate a reward signal based on the first evaluation result. The model parameters of the text generation model are then updated based on the reward signal. The aesthetic fit reflects the overall visual experience of the sample scene image in terms of layout balance, harmonious arrangement, and the sense of light and shadow hierarchy.
[0077] In some optional embodiments, a second evaluation model can be used to evaluate the semantic alignment of the sample scene graph, obtain a second evaluation result, and generate a reward signal based on the second evaluation result. The model parameters of the text generation model are then updated according to the reward signal. Here, semantic alignment reflects the degree of matching between the semantic content expressed by the sample scene graph and the initial prompt word. If there is no initial prompt word, the second evaluation result can use a default value.
[0078] In some optional embodiments, a third evaluation model can be used to evaluate the scene richness of the sample scene graph, obtain the third evaluation result, generate a reward signal based on the third evaluation result, and update the model parameters of the text generation model according to the reward signal. Scene richness reflects the diversity of objects and the complexity of spatial hierarchy in the sample scene graph.
[0079] In some optional embodiments, a first evaluation model can be used to evaluate the aesthetic fit of the sample scene graph to obtain a first evaluation result; a second evaluation model can be used to evaluate the semantic alignment of the sample scene graph to obtain a second evaluation result; a reward signal can be generated based on the first evaluation result and the second evaluation result, and the model parameters of the text generation model can be updated according to the reward signal.
[0080] In some optional embodiments, a first evaluation model can be used to evaluate the aesthetic fit of the sample scene image to obtain a first evaluation result; a third evaluation model can be used to evaluate the scene richness of the sample scene image to obtain a third evaluation result; a reward signal can be generated based on the first evaluation result and the third evaluation result, and the model parameters of the text generation model can be updated according to the reward signal.
[0081] In some optional embodiments, a second evaluation model can be used to evaluate the semantic alignment of the sample scene graph to obtain a second evaluation result; a third evaluation model can be used to evaluate the scene richness of the sample scene graph to obtain a third evaluation result; a reward signal can be generated based on the second and third evaluation results, and the model parameters of the text generation model can be updated according to the reward signal.
[0082] In some optional embodiments, a first evaluation model can be used to evaluate the aesthetic fit of the sample scene graph to obtain a first evaluation result; a second evaluation model can be used to evaluate the semantic alignment of the sample scene graph to obtain a second evaluation result; a third evaluation model can be used to evaluate the scene richness of the sample scene graph to obtain a third evaluation result; a reward signal is generated based on the first evaluation result, the second evaluation result, and the third evaluation result, and the model parameters of the text generation model are updated according to the reward signal.
[0083] For example, when multiple evaluation results are obtained from multiple dimensions, a reward signal can be generated based on the multiple evaluation results through weighted calculation. For instance, weights w1, w2, and w3 are assigned to the first evaluation result x1, the second evaluation result x2, and the third evaluation result x3, respectively, and the reward signal = x1w1 + x2w2 + x3w3.
[0084] The embodiments of this application do not limit the specific implementation of the first evaluation model, the second evaluation model and the third evaluation model described above. Exemplary descriptions are given below.
[0085] In an exemplary embodiment, the step of "using the first evaluation model to evaluate the aesthetic fit of the sample scene image and obtaining the first evaluation result" in the foregoing embodiment can be implemented based on the following steps Q1-Q2: Q1. The third image encoder of the first evaluation model is used to divide the sample scene image into multiple image blocks, and interactive learning is performed between multiple image blocks based on the self-attention mechanism to obtain the second visual features. The second visual features represent the coordination relationship between the scene content and scene structure in the sample scene image in terms of layout balance, matching coordination and light and shadow hierarchy.
[0086] The first evaluation model is a regression model pre-trained on a large number of image-text pairs and fine-tuned on human-annotated aesthetic rating data. Through supervised fine-tuning on a large number of images with comprehensive human aesthetic ratings, the first evaluation model implicitly learns the complex combination rules of aesthetic factors such as layout, color, and lighting. It can distinguish which visual patterns correspond to high scores (aesthetically pleasing) and which correspond to low scores (unattractive), and thus evaluate the aesthetics of the sample scene images from aspects such as layout balance, matching harmony, and the sense of lighting and shadow hierarchy.
[0087] Layout balance refers to the placement of objects in a scene that conforms to the logic of actual space, conveys a sense of life, and makes the image full yet uncluttered. The placement needs to consider the positional relationships of all objects in the scene, avoiding objects appearing to float or unconventional placements (such as placing objects in reverse). In addition to the main sample objects, some auxiliary objects are needed in the scene to convey a sense of life. Inappropriate placement, lack of atmosphere, or a cluttered image will all result in poor layout balance.
[0088] Harmony and coordination refer to the harmonious matching of style and color between scene content and scene structure, conforming to spatial aesthetics. If there are objects or colors with severely conflicting styles in the scene, creating a noticeable sense of disharmony, then harmony and coordination are seriously problematic. If the styles or color combinations of different objects in the scene clash, creating a sense of disharmony, then harmony and coordination are poor. If the styles or colors of different objects in the scene are generally unified, although there may be slight differences, the overall visual flow is smooth and does not produce abruptness, then harmony and coordination are good. If the styles of different objects in the scene are harmonious and their colors echo each other, presenting a natural, balanced, and aesthetically pleasing visual effect that makes people feel comfortable and immersed, then harmony and coordination are excellent. For example, a scene structure with a Chinese landscape background wall paired with scene content featuring a luxurious European crystal chandelier creates a style conflict. Another example is a scene structure with blue walls, and scene content featuring a purple sofa and a black coffee table, creating a color conflict.
[0089] A sense of depth in lighting and shadow refers to the natural, harmonious, and unified interplay of light and shadow in a scene, conforming to aesthetic principles. The light and shadow within a scene should closely match the display lighting, avoiding overexposure or darkness caused by excessively strong or weak light. Furthermore, the overall lighting and shadow within the scene should be unified, with appropriate primary and secondary light sources and natural transitions, and the direction of light sources for each object should align with the ambient light. Poor depth in lighting and shadow is indicated by large areas of overexposure or darkness, noticeable light spots from lights, or a severe disharmony between the scene's lighting and the objects' own lighting.
[0090] In this embodiment, the third image encoder of the first evaluation model is used to convert the input image into a representation that can be understood and processed by the machine, possessing strong semantic and style understanding capabilities. The third image encoder can divide the input sample scene image into multiple image blocks after normalization, model the global spatial relationship between multiple image blocks through a multi-layer self-attention mechanism, and extract a second visual feature representing the global semantics of the image. This second visual feature is a comprehensive aesthetic representation of the entire image, implicitly encoding aesthetic dimensions such as layout rationality, matching harmony, and light and shadow hierarchy, and can characterize the harmonious relationship of the sample scene image in terms of layout balance, matching harmony, and light and shadow hierarchy.
[0091] Q2, perform regression processing on the second visual features to obtain the first evaluation result.
[0092] Regression processing predicts a continuous value from input features. By performing regression processing on the second visual features, the layout balance, matching harmony, and sense of light and shadow hierarchy can be mapped into a quantifiable aesthetic score. The higher the score, the higher the aesthetic fit of the sample scene image, and the closer it is to the subjective preferences of real humans.
[0093] The following combination Figure 5The architecture and processing flow of the first evaluation model described above are illustrated by example.
[0094] The first evaluation model may include a third image encoder, a multimodal feature enhancement module, and a regression head. The third image encoder may include an embedding layer, a positional encoding module, and a multilayer encoder.
[0095] The embedding layer in the third image encoder divides the input scene graph (e.g., 224×224) into fixed-size image patches (e.g., 16×16) and maps each patch to a multi-dimensional (e.g., 768) embedding vector through a linear projection layer. The positional encoding module in the third image encoder adds learnable two-dimensional positional encodings to multiple image patch vectors to preserve the original spatial structure information. The encoder in the third image encoder consists of multiple identical stacked layers, each containing two core sublayers: a multi-head self-attention mechanism and a feedforward neural network. The multi-head self-attention mechanism captures the dynamic relationships between image patches in the input sequence by calculating a query (Q), key (K), and value (V) matrix. By segmenting the self-attention mechanism into multiple heads, the model can learn information from different subspaces in parallel, thereby enhancing the model's expressive power. The feedforward neural network is a fully connected network consisting of two linear transformation layers and an activation function, which can perform independent nonlinear transformations on the representation of each location, further extracting and integrating features. By using multi-layer stacked encoders, global and local visual features of an image can be extracted step by step, and the final output is a serialized feature map (second visual feature) with dimensions [N, 768], where N is the number of blocks, such as a 224×224 image corresponding to 196 blocks.
[0096] The multimodal feature enhancement module optimizes visual feature representations through cross-modal skip connections. The module comprises multiple skip connection fusion blocks, each containing S asymmetric co-attention layers. These layers consist of self-attention, cross-attention, and feedforward neural networks, recursively enhancing cross-modal fusion efficiency. In pure image input scenarios, the multimodal feature enhancement module simplifies to a single-modal feature enhancer, utilizing inter-layer shortcuts to skip some self-attention layers, reducing the computational overhead of long sequences. The input to the multimodal feature enhancement module is the serialized feature map output by the third image encoder, and the output is the enhanced high-density visual features, maintaining the same dimensionality (e.g., [196, 768]).
[0097] The regression head can consist of lightweight fully connected layers, including a linear projection layer and an output layer. The linear projection layer reduces the feature dimension (e.g., 768 dimensions) of the third image encoder output to a smaller intermediate dimension (e.g., 256 dimensions). The output layer maps the features to the final scalar output (e.g., a 0-1 rating value) through linear transformations (e.g., activation functions), completing the conversion from visual features to continuous ratings. During training, the regression head performs end-to-end optimization by minimizing the mean squared error (MSE) between the predicted score and the average human-annotated rating, thereby learning to implicitly weight and fuse different aesthetic dimensions.
[0098] The processing flow of the first evaluation model is as follows: The third image encoder divides the input image into fixed-size image blocks, each image block encoding the visual content of a corresponding local region (e.g., "lower left corner of the sofa," "middle of the wall," or "rug by the window"). These image blocks are transformed into embedding vectors after linear projection, and their initial representations already contain local RGB color, texture, and brightness information. The image block label sequence is input into the encoder, which dynamically models the relationship between any two image blocks globally through a multi-head self-attention mechanism, thereby systematically learning the following aesthetic dimensions.
[0099] In terms of scene layout, the self-attention mechanism can capture spatial composition rules: such as the relative positional relationship between furniture (e.g., whether the sofa and coffee table form visual symmetry or functional alignment), whether the traffic flow is smooth (judging the traffic logic by the continuity and distribution density of image blocks in open areas), whether the visual center of gravity of the picture is balanced (whether the distribution of attention weight in the upper / lower / left / right areas is balanced), and whether the white space is appropriate (whether the spatial proportion and position of large areas of low texture or low saturation conform to the aesthetic rhythm), etc.
[0100] In terms of color matching, after embedding the RGB information of each image patch, the encoder implicitly learns the distribution of the dominant hue (the clustering tendency of high-frequency color image patches), color contrast intensity (whether complementary colors or warm and cool colors create harmonious tension rather than conflict and glaring), and the global coordination of brightness and saturation (such as whether high-saturation areas are excessively concentrated and cause visual fatigue) by comparing the color representation of different areas such as walls, furniture, and soft furnishings. It then determines whether these conform to the color matching principles.
[0101] In terms of style and material coordination, a global matching semantic is gradually constructed through multi-layer stacking: at the shallow layer, the categories and material attributes of local elements can be identified (such as fabric sofas, marble countertops, and matte walls); at the middle layer, semantic relationships between elements can be established, associating element styles, such as metal and glass corresponding to minimalist industrial style; at the deep layer, it can assess whether the overall combination is stylistically unified, material transitions are natural, colors are matched, and there is a clear distinction between primary and secondary elements. If the scene content and scene structure present consistent style semantics in the feature space (such as both pointing to "modern minimalism"), high attention weight is generated, contributing to a positive score; conversely, if semantic conflicts occur (such as a Chinese carved screen paired with Nordic minimalist furniture), a negative mode is activated, reducing the aesthetic score.
[0102] In terms of lighting and shadow representation, the encoder is sensitive to changes in brightness gradients. It can implicitly learn the distribution of brightness histograms and local contrast through pixel value statistics, and infer from the brightness differences between image blocks whether the light source direction is consistent, whether the shadow transition is natural, whether the highlight areas are realistic, and whether the overall illumination is balanced. These lighting and shadow cues together constitute the characteristics of lighting and shadow aesthetics.
[0103] After systematically learning the aesthetic dimension features, a serialized feature map is output. The multimodal feature enhancement module will serialize the feature map and output the enhanced high-density visual features. The regression head performs implicit weighted fusion on different aesthetic dimensions, mapping the features to aesthetic scores, thereby obtaining the first evaluation result corresponding to the aesthetic fit.
[0104] In addition to using the first evaluation model to assess aesthetic fit, a second evaluation model can be used to assess semantic alignment. By determining whether the scene graph generated from the samples contains elements from the initial prompts, the semantic fidelity of the text generation model can be ensured, preventing task failure in subsequent inference applications. The second evaluation model needs to understand the key semantic units in the initial prompts and locate or confirm their existence in the sample scene graph. For example, if the initial prompts mention "retro desk lamp," the second evaluation model not only needs to identify "lamp" but also determine whether it possesses the "retro" style characteristics.
[0105] In an exemplary embodiment, the step of "using a second evaluation model to evaluate the semantic alignment of the sample scene graph and obtaining a second evaluation result" in the foregoing embodiment can be implemented based on the following steps P1-P4: P1 uses a fourth image encoder to encode the sample scene image to obtain third visual features. The third visual features characterize the object category, environmental attributes, spatial relationships and scene style expressed by the sample scene image.
[0106] Among them, object category refers to the types of objects and decorative details included in the sample scene image, such as tables, chairs, and green plants. Environmental attributes refer to the color, material, and lighting atmosphere of the objects in the sample scene image. Spatial relationship refers to the relative layout and position of objects in the sample scene image, such as green plants on the left side of the wall and a sofa in the middle. Scene style refers to the overall style of the environment in the sample scene image, such as a modern living room or a retro bedroom.
[0107] P2, the initial prompt words are encoded using a second text encoder to obtain the second text features.
[0108] In the embodiments of this application, a standalone text encoder or a non-standalone text encoder can be used. For example, the embedding layer of a language model can be used as a text encoder to process text.
[0109] P3 utilizes a cross-attention mechanism to align the second text features and the third visual features in a shared semantic space, resulting in semantic alignment features. These semantic alignment features represent consistency in object category, environmental attributes, spatial relationships, and scene style.
[0110] Cross-attention is a natural manifestation of self-attention in mixed sequences. The self-attention mechanism acts on both visual and textual features simultaneously. By using the second textual feature as the query and the third visual feature as the key and value, cross-modal interaction occurs within a shared semantic space. This allows textual features to dynamically focus on the most relevant visual regions in the image, while visual features also focus according to the semantic guidance of the text. This alignment process enables the model to establish a connection between the textual description and the image content, thus effectively measuring whether the sample scene image accurately captures the key elements in the initial prompt.
[0111] Among these, object category consistency refers to whether the sample scene image contains the object category and decorative details specified by the text prompt; environmental attribute consistency refers to whether the color, material, and lighting atmosphere of objects in the sample scene image match the text description; spatial relationship consistency refers to whether the relative layout positions of objects in the sample scene image conform to the text instructions; and scene style consistency refers to whether the overall environment in the sample scene image is consistent with the text style description.
[0112] P4 performs natural language reasoning on the semantic alignment features to obtain the second evaluation result.
[0113] The second evaluation model can perform natural language inference on semantic alignment features, outputting text that includes dimensional analysis and numerical scores. Although the score is presented in linguistic form, it essentially reflects the model's internal judgment on the degree of semantic alignment between the image and text, and can be extracted into a structured score through post-processing.
[0114] The following will combine Figure 6 The architecture and processing flow of the second evaluation model are illustrated with examples.
[0115] like Figure 6 As shown, the second evaluation model can include a fourth image encoder, a second multimodal alignment module, and a second language model. The embedding layer of the fourth image encoder divides the input sample scene image into fixed-size image patches, each of which is linearly projected into a multidimensional vector. Subsequently, the positional encoding module adds learnable two-dimensional positional codes to multiple image patch vectors to preserve the original spatial structure information. The resulting image patch label sequence is input into a backbone network composed of multi-layered stacked encoders, each layer containing a multi-head self-attention mechanism and a feedforward neural network. The encoder dynamically models the relationship between any two image patches globally using the multi-head self-attention mechanism and the feedforward neural network, systematically learning cues such as local color, texture, and material, and gradually abstracting high-level semantics such as objects, layout, lighting, decoration, and style, ultimately outputting third visual features.
[0116] The second multimodal alignment module of the second evaluation model is responsible for aligning the third visual features output by the fourth image encoder with the text semantic space of the language model. For example, the second multimodal alignment module may include a learnable linear projection layer that transforms the third visual features into a vector consistent with the embedding dimension of the second language model through a learnable linear (or non-linear) mapping.
[0117] The second language model of the second evaluation model serves as the final output head, enabling reasoning and score generation based on textual and visual features. The second language model can include a pre-processor (word segmenter), an embedding layer, a positional encoding module, a self-attention layer, and an output head. The initial prompt word is segmented into a series of discrete units by the word segmenter and then converted into second textual features by the word embedding layer of the second language model. In the embedding space of the second language model, the third visual features and the second textual features are concatenated into a multimodal input sequence. The positional encoding module adds positional information to the concatenated sequence. Subsequently, this multimodal input sequence is fed into the self-attention layer, which consists of multiple stacked decoders, each containing a multi-head self-attention mechanism and a feedforward neural network.
[0118] The self-attention mechanism calculates association weights between any text unit and any visual unit, thereby dynamically establishing a "word-region" correspondence. For example, when the text mentions "mahogany dining table," the attention mechanism focuses on the area in the image that has a wood texture, a deep red tone, and a tabletop shape; when describing "left-side white space," the model focuses on the low-complexity area in the left half of the image. In the self-attention layer, the model implicitly performs consistency verification from dimensions such as object category, environmental attributes, spatial relationships, and scene style, generating the final hidden state after multimodal fusion, called semantic alignment features.
[0119] The output head of the second language model in the second evaluation model transforms consistency evaluation into a structured natural language reasoning task. By designing specific instructional prompts (such as "Please evaluate the fidelity of this image to the input description from dimensions such as object category, environmental attributes, spatial relationships, and scene style, and give a comprehensive score of 0-1"), the second evaluation model can output text containing dimensional analysis and numerical scores as the second evaluation result. Furthermore, structured numerical scores can be automatically extracted through text parsing (such as regular expression matching) for calculating reward signals.
[0120] For example, the initial prompt could be: "This is a living room space. The TV is mounted on a light beige wall with decorative lines on both sides. A TV cabinet is placed against the wall. On the right side of the screen, near the balcony, there is a vintage floor lamp. In front of the TV cabinet, there is a light-colored, simple-lined rug with a bunch of pink roses and an open book. Recessed lights are installed on the ceiling, emitting a warm glow throughout the space and enhancing the overall atmosphere. The floor is covered with light-colored wood flooring, complementing the room's modern and open design. Natural light enters from the window on the right, ensuring the room is bright and comfortable. A large potted plant sits on the balcony on the right." A sample scene image could be like this: Figure 6 As shown, after processing by the second evaluation model, the output of the second evaluation result can be "{Score: 0.85, Reason: The picture mostly matches the description, but there are slight differences in some details. Analysis: {Advantages: [The TV cabinet is placed on the wall and the TV is hung in the middle, which matches the description; the storage cabinet is light-colored and has multiple drawers, located in the center of the picture against the wall, which is consistent with the description; a potted plant is placed next to the window on the right, which increases the vitality of the space; there is a dark blue gradient decoration (carpet) on the left floor, which creates a comfortable and open space atmosphere; the overall design is neat, combining functionality and aesthetics, which meets the description requirements], Disadvantages: [The description mentions an oil painting, but the picture actually shows a gradient carpet, which is a slight difference; the picture does not clearly show the window on the right, and this detail is not obvious in the picture]}}".
[0121] In addition to the evaluation content mentioned above, a third evaluation model can also be used to evaluate scene richness. In an exemplary embodiment, the "using a third evaluation model to evaluate the scene richness of the sample scene graph and obtain the third evaluation result" in the aforementioned embodiment can be implemented based on the following steps K1-K4: K1 uses the fifth image encoder to extract multi-level features from the sample scene image to obtain multi-scale visual features. Among them, the multi-scale visual features are used to analyze the complexity of the three-dimensional structure of the wall, the diversity of spatial shapes, the number of material types, and the number of identifiable objects.
[0122] Scene richness can be divided into scene structure richness and scene content richness. Scene structure richness can be evaluated from dimensions such as the complexity of the wall's three-dimensional structure, the diversity of spatial shapes, and the variety and quantity of materials. If the wall is relatively flat and uses a single material, it indicates a simple scene structure, and the corresponding scene structure richness score will be low. If the wall has a clear three-dimensional sense, contains some design elements, and is composed of different materials, the scene structure richness is at a medium level, and the corresponding scene structure richness score will be moderate. If the wall structure is complex, has a sophisticated shape, and uses multiple materials interwoven and integrated, the scene structure is very rich, and the corresponding scene structure richness score will be high. Scene content richness can be evaluated by the number of identifiable objects. Here, identifiable objects refer to movable items in the scene used to construct the living scene and for decorative embellishment.
[0123] K2 uses a third text encoder to encode the evaluation prompts to obtain third text features.
[0124] Evaluation prompts are used to clearly define the dimensions of the visual attributes to be analyzed, specify the corresponding quantitative or qualitative evaluation criteria, and standardize the format and content requirements of the output results. Specifically, these prompts, as explicit representations of the task context, guide the third-party evaluation model to focus on specific scene elements (such as the complexity of the three-dimensional structure of a wall, the types and quantities of materials, etc.), interpret the extracted visual features according to preset scoring criteria or evaluation logic, and generate analysis results that conform to the specified format. In simple terms, it tells the model "what to analyze," "what standards to use," and "how to output the results."
[0125] The evaluation model obtains third text features after being encoded by a third text encoder. These text features not only contain the literal meaning of keywords but also implicitly encode the task intent, the logical relationships between evaluation dimensions, and the constraints of the scoring rules. These semantic vectors serve as anchor points for understanding, providing contextual guidance and judgment criteria for the subsequent interpretation of multi-scale visual features, enabling the model to focus on relevant visual content according to specific instructions.
[0126] K3 uses third-party text features as the evaluation standard to analyze multi-scale visual features to obtain scene content richness scores and scene structure richness scores.
[0127] The third evaluation model, based on the semantic standards defined by the third text features, performs targeted analysis and quantitative evaluation of multi-scale visual features. On the one hand, by matching the semantic consistency between visual regions and text descriptions, it statistically analyzes the types and quantities of identifiable objects to measure the richness of scene content. On the other hand, by combining geometric, topological, and material distribution information, it evaluates the morphological complexity, structural hierarchy, shape diversity, and material diversity of the space to form the richness of scene structure. This process can be achieved through cross-modal alignment mechanisms (such as attention, similarity calculation, or inference modules).
[0128] K4 combines the scene content richness score and the scene structure richness score to obtain the third evaluation result.
[0129] The fusion strategy can take the form of weighted average, rule mapping or learnable gating mechanism, etc., and its goal is to integrate multi-dimensional information and output an overall evaluation result that reflects both the diversity of objects and the sense of spatial hierarchy.
[0130] The following will combine Figure 7 The architecture and processing flow of the third evaluation model described above are illustrated with examples.
[0131] like Figure 7 As shown, the third evaluation model can include a fifth image encoder, a third multimodal alignment module, and a fourth language model. The embedding layer of the fifth image encoder divides the input sample scene image into fixed-size image blocks, each of which is linearly projected into a multidimensional vector. Subsequently, the position encoding module adds learnable two-dimensional positional codes to multiple image block vectors to preserve the original spatial structure information. The resulting image block label sequence is input into a backbone network composed of multi-layered stacked encoders, each layer containing a multi-head self-attention mechanism and a feedforward neural network. The encoder dynamically models the relationship between any two image blocks globally using the multi-head self-attention mechanism and feedforward neural network, systematically learning cues such as local edges and materials, and gradually abstracting multi-scale visual features: the shallow feature layer preserves high-resolution edges, textures, and local geometric details for material boundaries and microstructure recognition; the middle feature layer captures medium-scale spatial structure information such as wall contours, concavity and convexity variations, and shape transitions, providing a basis for quantifying the complexity of three-dimensional structures and the diversity of spatial shapes; the high-level feature layer encodes global semantic information, including high-level concepts such as object categories and material semantics, supporting the semantic recognition and counting of decorations.
[0132] The third multimodal alignment module of the third evaluation model is responsible for aligning the multi-scale visual features output by the fifth image encoder with the text semantic space of the language model. For example, the third multimodal alignment module may include a learnable linear projection layer that transforms the multi-scale visual features into vectors consistent with the embedding dimension of the language model through a learnable linear (or non-linear) mapping.
[0133] The third language model of the third evaluation model serves as the final output head, enabling reasoning and score generation based on textual and visual features. The third language model can include a pre-processor (word segmenter), an embedding layer, a positional encoding module, a self-attention layer, and an output head. The evaluation prompts are segmented into a series of discrete units by the word segmenter and then converted into third textual features by the word embedding layer of the third language model. In the embedding space of the third language model, multi-scale visual features and third textual features are concatenated into a multimodal input sequence. The positional encoding module adds positional information to the concatenated sequence. Subsequently, this multimodal input sequence is fed into the self-attention layer, which consists of multiple stacked decoders, each containing a multi-head self-attention mechanism and a feedforward neural network.
[0134] Self-attention mechanisms calculate association weights between any text unit and any visual unit, achieving fine-grained semantic alignment. For example, "spatial design" guides attention to compare the types and combinations of design units (flat, curved, sloping, hollow, etc.) across different areas of a wall. Another example is "material types and quantities," which triggers the model to cluster the material semantics of different wall blocks and estimate the number of types based on the differences in cross-regional material embedding. Yet another example is "three-dimensional structure," which guides attention to focus on the visual features corresponding to depth changes in a wall area, assessing complexity through its geometric distribution and structural diversity. Finally, "object quantity" guides the model to locate semantically clear objects across the entire image, counting them based on the activation intensity and category confidence of object-level visual features.
[0135] The third language model of the third evaluation model generates natural language text containing analytical logic and scoring results in an autoregressive manner. By designing specific instructional prompts, the third evaluation model first outputs a step-by-step observation and quantitative judgment of the wall's three-dimensional structure, spatial form, material types, and the quantity of objects. Then, based on the scoring criteria provided in the evaluation prompts, it outputs structured scores for scene content richness and scene structure richness. Furthermore, the scene content richness score and scene structure richness score can be fused to obtain the third evaluation result.
[0136] For example, the evaluation prompt could be: "You are a professional interior design analyst. Your task is to analyze the interior scene images I provide and provide a rigorous quantitative score based on the following two dimensions: 'Richness of Scene Structure' and 'Richness of Scene Content.' To facilitate only counting the background, I have covered the main products with gray. Please do not count the gray-covered content. Please strictly adhere to the following scoring criteria: Richness of Scene Structure (background_wall_richness): Score 0 (Easy): The walls are basically flat, mainly composed of a single material (such as solid-color latex paint, wallpaper with simple repeating textures). There is almost no three-dimensional structure or shape. Score 1 (Medium): The walls have obvious three-dimensional layering, and begin to use a combination of two or three different materials (such as latex paint + wood veneer / stone / lines). There are design elements such as geometric shapes, wainscoting, wooden grilles, or light strips. Score 2 (Very Rich):" The wall is a highly integrated and complex system. It utilizes three or more materials in a complex interweaving and fusion, exhibiting sophisticated forms such as sculptural qualities, curved surfaces, and floating elements, and systematically integrates functions such as hidden doors, custom cabinetry, smart lighting, or fireplaces. Background wall richness: Definition: "Scene content" refers to movable items used to create a living scene and for decorative embellishment. Please only count the number of items in the following categories: sofa, bed, table, chair, cabinet, side table, bedside table, green plants, decorative paintings, carpet, sculptures, vases, books, dressing table, mirror, table lamp, and robot vacuum cleaner. Note that auxiliary items on the main furniture, such as duvet covers and pillows on the bed and blankets on the sofa, are not counted. Items of the same purpose or symmetrical items, such as one bedside table on each side, are counted as only one item. A score of 0 is given if the total number of soft furnishing items is less than or equal to 2. A score of 1 is given if the total number of soft furnishing items is between 3 and 5. A score of 2 is given if the total number of soft furnishing items is greater than or equal to 6. Output requirements: After analyzing the image, please first output your thought process. Finally, your answer must and can only be a strict JSON object, in the following format: JSON {"background_wall_richness_score": [score], "soft_furnishing_richness_score": [score]}, where [score] must be 0, 1, or 2. "An integer in 2".
[0137] Correspondingly, for Figure 7The structured score for the image shown is: "Background Wall Richness Score: 1. Wall Material and Structure: Part of the background wall is a light-colored flat wall with a single material (e.g., latex paint); the right background wall is combined with a bookshelf, whose design adds a sense of depth, but the overall wall surface remains predominantly flat; 2. Design Elements: The bookshelf features an embedded design with lighting effects, increasing its depth and functionality; the wall itself lacks complex geometric shapes, wainscoting, or wooden grilles. 3. Matching Scoring Criteria: The background wall is primarily flat, but the embedded design of the right-hand bookshelf adds a degree of depth. This design meets the criteria of 'initiating a combination of two or three different materials (e.g., latex paint + wood veneer / stone / molde)'. Therefore, the background wall richness score is 1 (Medium)."
[0138] Scene Content Richness (soft_furnishing_richness_score): 1. Soft Furnishing Item Recognition: Curtains: There are curtains on the left; Green Plants: There are several potted plants on the bookshelf; Decorations: There are some decorations (such as small ornaments) on the bookshelf; Books: There are books on the bookshelf; Carpet: There is a carpet on the floor; 2. Quantity of Soft Furnishing Items: Curtains: 1 item; Green Plants: At least 2 items (plants on the bookshelf); Decorations: At least 2 items (small ornaments on the bookshelf); Books: Counted as 1 item (books on the bookshelf are considered as one category); Carpet: 1 item; 3. Total Quantity: There are at least 1 (curtain) + 2 (green plants) + 2 (decorations) + 1 (book) + 1 (carpet) = 7 soft furnishing items in total. 4. Scoring Criterion Matching: The total number of soft furnishing items is between 5 and 7. Therefore, the background soft furnishing richness score is 1.
[0139] json{ "background_wall_richness_score": 1, "soft_furnishing_richness_score": 1}".
[0140] By implementing the first, second, and third evaluation models described above, multi-dimensional evaluation of the sample scene graph can be performed, yielding corresponding first, second, and third evaluation results. Furthermore, a reward signal can be generated, and the model parameters of the text generation model can be updated based on the reward signal.
[0141] In some optional embodiments, before inputting the sample scene prompts into the image generation model, the format of the sample scene prompts can be evaluated based on the target format to obtain the format evaluation result.
[0142] In this embodiment, after the text generation model outputs sample scene prompts, it first determines whether the output sample scene prompts conform to a standard target format (e.g., JSON), which includes description and position fields. If a position is pre-defined, it also determines whether the output position matches the pre-defined position. If the format of the sample scene prompts conforms to the target format, the format evaluation passes, and subsequent evaluations can proceed. Then, based on at least one of the first, second, and third evaluation results, combined with the format evaluation result, a reward signal is generated for reinforcement learning (RL). If the format of the sample scene prompts does not conform to the target format, the format evaluation fails, the format evaluation result is 0, and reinforcement learning is temporarily suspended; instead, supervised fine-tuning (SFT) training is performed.
[0143] The purpose of supervised fine-tuning is to teach the model basic instruction following ability and the target format required for output. The purpose of reinforcement learning is to allow the text generation model, based on SFT, to autonomously explore how to generate scene prompts that can obtain higher aesthetic scores in a "creation-feedback" cycle. By interspersing supervised fine-tuning and reinforcement learning stages during training, it is possible to ensure that the text generation model smoothly transitions from imitating examples to autonomous creation.
[0144] For the supervised fine-tuning phase, approximately 200,000 "product image / user input - ideal output" data pairs can be constructed as a training dataset. For example, a single sample data pair includes: sample {"user_prompt": "", "position":null, ...} + [product image / user input - ideal output]. Figure 1 The tag is {"prompt": "A gold-colored sculpture...", "position": ...}. For example, a sample data point might be {"user_prompt": "Kitchen scene...", "position": {"left_top_x": 0.2275, "left_top_y": 0.03125, ...}. "right_bottom_x":0.77125,"right_bottom_y":0.9875}}+[item Figure 2 ], tagged with {"prompt": "A wooden shelving unit...", "position": {"left_top_x": 0.2275, "left_top_y":0.03125, "right_bottom_x": 0.77125, "right_bottom_y": 0.9875}}.
[0145] For reinforcement learning, this application does not limit the specific implementation of the reinforcement learning algorithm. In an exemplary embodiment, the text generation model generates multiple sample scene prompts for the same sample image, and the multiple sample scene prompts correspond to multiple sample scene images. The evaluation model evaluates the sample scene images from a spatial aesthetic perspective and generates a reward signal based on the evaluation results. This includes: evaluating the multiple sample scene images separately to obtain multiple evaluation results; sorting the multiple evaluation results and determining the relative weights corresponding to each of the multiple evaluation results based on the sorting results. The relative weights are used to adjust the contribution of the sample scene images in the model parameter update process; and generating a reward signal based on the multiple evaluation results and their corresponding relative weights.
[0146] In this embodiment, the text generation model generates multiple sample scene prompts for the same sample image, and then the image generation model generates a corresponding sample scene image for each sample scene prompt. Each sample scene image is evaluated in terms of target format, aesthetic fit, semantic alignment, and scene richness, yielding evaluation results. Then, the evaluation results of the sample scene images corresponding to the same sample image are ranked from highest to lowest to determine their relative merit. Based on this ranking, a weight allocation strategy (e.g., linear interpolation based on ranking position) is used to assign a relative weight to each sample scene image, so that the sample scene images ranked higher (i.e., of higher quality) receive larger weight values. During the model parameter update phase, the evaluation results of each sample scene image are multiplied by their corresponding relative weights to generate a weighted reward signal. Based on the policy gradient method, the logarithmic probability gradient of the generation action for each sample scene image is calculated using the weighted reward signal as a supervision signal, and the weighted gradients of multiple sample scene images are summed as the final gradient update direction used for backpropagation. High-quality sample scene images, with their larger relative weights, dominate the summation of their corresponding gradient terms, significantly enhancing their contribution to model parameter updates. In contrast, the gradient contribution of low-weight sample scene images is suppressed, thus guiding the model to focus more on improving its ability to generate high-quality scene images during training.
[0147] Through the above model training process, sample scene prompts are first generated based on sample images containing sample objects. Then, guided by the sample images and sample scene prompts, an image generation model is used to generate sample scene images containing sample objects. The sample scene images are then evaluated from a spatial aesthetic perspective. Based on the evaluation results of the sample scene images, the model parameters of the text generation model are updated, thereby training a text generation model. This text generation model can generate scene prompts that meet the requirements of spatial aesthetics. Furthermore, in the inference stage, guided by the scene prompts, the image generation model can generate scene images that better meet the requirements of spatial aesthetics.
[0148] For the reasoning stage, this application embodiment also provides a scene graph generation method, including the following steps: Step 21: Obtain the target image including the target object.
[0149] Step 22: Input the target image into the text generation model to generate scene prompts, so as to obtain the target scene prompts corresponding to the target image. The target scene prompts describe the target scene information that needs to be generated by the model in text form.
[0150] The text generation model was trained using the aforementioned model training method.
[0151] Step 23: Input the target image and target scene prompts into the image generation model to generate a target scene map containing target object and target scene information.
[0152] For detailed implementation methods and beneficial effects of each step in this embodiment, please refer to the foregoing embodiments; they will not be described in detail here.
[0153] In addition, this application embodiment also provides another method for generating scene graphs, including the following steps: Step 31: Display the image generation interface, which includes an image upload control.
[0154] Step 32: In response to the triggering operation of the image upload control, obtain the target image uploaded by the user. The target image includes the target object.
[0155] Step 33: Input the target image into the text generation model to generate scene prompts, so as to obtain the target scene prompts corresponding to the target image. The target scene prompts describe the target scene information that needs to be generated by the model in text form.
[0156] Step 34: Input the target image and target scene prompts into the image generation model to generate a target scene map containing target object and target scene information.
[0157] For example, such as Figure 8As shown, the client's user interface can display an image generation interface for generating scene diagrams. This image generation interface includes an image upload control 81.
[0158] The image upload control 81 allows users to upload target images. The target image is a product image uploaded by the user, containing the main product element, i.e., the target object. The target image can be an image uploaded from a local file or selected from the history. For example, when the image upload control 81 is triggered by the user, the electronic device can retrieve the user-uploaded target image. The user can trigger the "Start Generation" control; in response, the target image is input into the text generation model and image generation model for processing, resulting in a target scene image. The generated target scene image can be displayed on the image generation interface. Users can save the target scene image and then display it in marketing scenarios. For example, the target scene image can be used as an e-commerce main image, a product details page image, or a social media content image.
[0159] Optionally, the image generation interface may also include a size selection control 82, which allows the user to upload and select the size of the target scene image, such as 1:1 or 3:4. In response to a triggering operation on the size selection control 82, the size of the target scene image selected by the user can be obtained. Then, the target scene image size, the target image, and the target scene prompt are input together into the image generation model, and the target scene image is generated under the guidance of these factors.
[0160] Optionally, the image generation interface may also include a link control 83, which allows users to bind e-commerce products. If a user binds an e-commerce product, after generating the target scene image, the target scene image can be directly applied to the main image of the e-commerce product or the product details page, allowing users to more conveniently and quickly display the target scene image in marketing scenarios.
[0161] In this embodiment, during the scene image generation process, artificial intelligence (AI) technology is introduced to optimize the scene image generation process and quickly generate scene images, thereby reducing shooting costs and improving generation efficiency. On the other hand, a text generation model is obtained based on the aforementioned model training method. This model constructs scene prompts that match the characteristics of the target object and guide the image generation model to produce high-quality, aesthetically pleasing images, making the generated target scene image more in line with human aesthetics.
[0162] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0163] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution subject of steps 11 to 14 can be device E; or the execution subject of steps 11 and 12 can be device E, and the execution subject of steps 13 and 14 can be device F, etc.
[0164] In some of the processes described in the above embodiments and accompanying drawings, multiple operations are included that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or they may be executed in parallel. The sequence numbers of the operations are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0165] Figure 9 This is a schematic diagram of a model training device provided as another exemplary embodiment of this application. (See diagram below.) Figure 9 As shown, the device includes: a sample prompt word generation module 901, a sample scene diagram generation module 902, an evaluation module 903, and an update module 904.
[0166] The sample prompt word generation module 901 is used to input a sample image containing a sample object into a text generation model, and generate scene prompt words using the current model parameters to obtain the sample scene prompt words corresponding to the sample object.
[0167] The sample scene graph generation module 902 is used to input sample images and sample scene prompts into the image generation model to generate sample scene graphs containing sample objects.
[0168] Evaluation module 903 is used to evaluate the sample scene image from a spatial aesthetic perspective using an evaluation model, and generate a reward signal based on the evaluation results.
[0169] The update module 904 is used to update the model parameters of the text generation model based on the reward signal, so that the text generation model generates scene prompt words that meet the requirements of spatial aesthetics.
[0170] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0171] Figure 10 This is a schematic diagram of a scene graph generation apparatus provided as another exemplary embodiment of this application. (See diagram below.) Figure 10As shown, the device includes: an acquisition module 1001, a target prompt word generation module 1002, and a target scene graph generation module 1003.
[0172] Acquisition module 1001 is used to acquire a target image including the target object; The target prompt word generation module 1002 inputs the target image into the text generation model to generate scene prompt words, so as to obtain the target scene prompt words corresponding to the target image. The target scene prompt words describe the target scene information that needs to be generated by the model in text form.
[0173] The target scene graph generation module 1003 is used to input the target image and target scene prompts into the image generation model to generate a target scene graph containing target object and target scene information.
[0174] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0175] Figure 11 This is a schematic diagram of a computing device provided in an embodiment of this application. This computing device can be used for model training and image generation. Figure 11 As shown, in practice, the computing device includes a memory 1101 and a processor 1102.
[0176] Memory 1101 is used to store computer programs and can be configured to store various other data to support operation on the computing platform. Examples of this data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0177] Processor 1102, coupled to memory 1101, is used to execute computer programs stored in memory 1101. When the computing device is used for model training, processor 1102 is used to: input sample images containing sample objects into a text generation model; generate scene cue words using the current model parameters to obtain sample scene cue words corresponding to the sample objects; input sample images and sample scene cue words into an image generation model; generate sample scene images containing sample objects under the guidance of sample images and sample scene cue words; evaluate the sample scene images from a spatial aesthetic perspective using an evaluation model; and generate a reward signal based on the evaluation results; update the model parameters of the text generation model based on the reward signal so that the text generation model generates scene cue words that meet the spatial aesthetic requirements.
[0178] When the computing device is used for image generation, the processor 1102 is used to: acquire a target image including the target object; input the target image into a text generation model to generate scene prompts, so as to obtain the target scene prompts corresponding to the target image, wherein the target scene prompts describe the target scene information to be generated by the model in text form; input the target image and the target scene prompts into the image generation model, and generate a target scene image containing the target object and target scene information under the guidance of the target image and the target scene prompts.
[0179] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.
[0180] Furthermore, such as Figure 11 As shown, the computing device also includes other components such as a communication component 1103, a display 1104, a power supply component 1105, and an audio component 1106. Figure 11 The diagram only shows some components and does not mean that the computing device includes only these components. Figure 11 The components shown. Additionally... Figure 11 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the work node. In this embodiment, the work node can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device, or a server-side device such as a conventional server, cloud server, or server array. If the work node in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 11 The components within the dashed box; if the working node in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, it may be omitted. Figure 11 The component within the dashed box.
[0181] The aforementioned memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0182] The aforementioned communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0183] The aforementioned display includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0184] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0185] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0186] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, digital video disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium.
[0187] Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.
[0188] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary hardware. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product, or it can be embodied in the process of data migration. The computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0189] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0190] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A model training method, characterized in that, include: Input the sample image containing the sample object into the text generation model, and use the current model parameters to generate scene prompt words to obtain the sample scene prompt words corresponding to the sample object; The sample image and the sample scene prompt are input into the image generation model to generate a sample scene image containing the sample object; The sample scene image is evaluated from a spatial aesthetic perspective using an evaluation model, and a reward signal is generated based on the evaluation results. The model parameters of the text generation model are updated based on the reward signal so that the text generation model generates scene prompts that meet the requirements of spatial aesthetics.
2. The method according to claim 1, characterized in that, The step of inputting a sample image containing a sample object into a text generation model to generate scene prompts, so as to obtain sample scene prompts corresponding to the sample object, includes: The sample image containing the sample object is input into the text generation model. The multimodal sequence generated based on the sample image is used as a guiding condition to perform scene prompt word generation processing on the sample image to obtain the sample scene prompt word corresponding to the sample object.
3. The method according to claim 2, characterized in that, The sample image containing the sample object is input into the text generation model. Guided by a multimodal sequence generated based on the sample image, scene cue word generation is performed on the sample image to obtain the sample scene cue words corresponding to the sample object, including: The sample image is input into the text generation model, and the following operations are performed in the text generation model: The first image encoder of the text generation model is used to extract semantic features from the sample image to obtain the first visual features; The first visual feature is projected into the embedding space of the language model of the text generation model; In the embedding space of the language model, the projected first visual feature is concatenated with the first text feature to generate the multimodal sequence; Using the language model and guided by the multimodal sequence, scene prompt words are generated for the sample image to obtain sample scene prompt words that are semantically consistent with the sample image.
4. The method according to claim 3, characterized in that, Also includes: After the initial prompt words are segmented, they are input into the language model to convert the initial prompt words into first text features. The initial prompt words are constructed for the sample images and describe the sample scene information that the model needs to generate in text form.
5. The method according to any one of claims 1-4, characterized in that, After inputting a sample image containing the sample object into the text generation model, the method further includes: In the text generation model, the current model parameters are used to perform positional reasoning on the sample object to obtain the positional information of the sample object in the sample scene graph. The positional information is used together with the sample scene prompt words to guide the image generation model to generate a sample scene graph containing the sample object.
6. The method according to any one of claims 1-4, characterized in that, The sample image and the sample scene prompts are input into an image generation model to generate a sample scene map containing the sample objects, including: The sample image and the sample scene prompts are input into the image generation model. The sample image and the sample scene prompts are used as noise prediction conditions to denoise the noise tensor to obtain the sample scene image.
7. The method according to claim 6, characterized in that, The sample image and the sample scene cue words are input into the image generation model. Using the sample image and the sample scene cue words as noise prediction conditions, the noise tensor is denoised to obtain the sample scene image, including: The sample image and the sample scene prompt are input into the image generation model, and the following operations are performed in the image generation model: The first text encoder of the image generation model encodes the sample scene prompts into text conditional embeddings. The sample image is encoded into a visual conditional embedding using the second image encoder of the image generation model; Initialize the random noise tensor in the latent space; Based on the conditional diffusion network in the image generation model, the text conditional embedding and the visual conditional embedding are used as noise prediction conditions to denoise the random noise tensor in order to obtain the target latent representation. The latent representation of the target is decoded using the decoder of the image generation model to obtain the sample scene image.
8. The method according to any one of claims 1-4 or 7, characterized in that, The evaluation of the sample scene image from a spatial aesthetic perspective using an evaluation model includes: Using the first evaluation model, the aesthetic fit of the sample scene image is evaluated, yielding a first evaluation result. The aesthetic fit reflects the overall visual experience of the sample scene image in terms of layout balance, harmonious arrangement, and the sense of light and shadow hierarchy; and / or, Using a second evaluation model, the semantic alignment of the sample scene graph is evaluated to obtain a second evaluation result. The semantic alignment reflects the degree of matching between the semantic content expressed by the sample scene graph and the initial prompt word; and / or, Using a third evaluation model, the scene richness of the sample scene image is evaluated, and the third evaluation result is obtained. The scene richness reflects the diversity of objects and the complexity of spatial hierarchy in the sample scene image. Accordingly, reward signals are generated based on the evaluation results, including: The reward signal is generated based on at least one of the first evaluation result, the second evaluation result, and the third evaluation result.
9. The method according to claim 8, characterized in that, Using the first evaluation model, the aesthetic fit of the sample scene images is evaluated, and the first evaluation result is obtained, including: Input the sample scene graph into the first evaluation model, and perform the following operations in the first evaluation model: The sample scene image is divided into multiple image blocks using a third image encoder, and interactive learning is performed between the multiple image blocks based on a self-attention mechanism to obtain a second visual feature. The second visual feature represents the coordination relationship between the scene content and scene structure in the sample scene image in terms of layout balance, matching coordination and light and shadow hierarchy. The second visual feature is subjected to regression processing to obtain the first evaluation result.
10. The method according to claim 8, characterized in that, Using the second evaluation model, the semantic alignment of the sample scene graph is evaluated to obtain the second evaluation result, including: Input the sample scene image and initial prompt words into the second evaluation model, and perform the following operations in the second evaluation model: The sample scene image is encoded using a fourth image encoder to obtain a third visual feature, which characterizes the object category, environmental attributes, spatial relationships, and scene style expressed by the sample scene image. The initial prompt word is encoded using a second text encoder to obtain second text features; By using a cross-attention mechanism, the second text feature and the third visual feature are aligned in a shared semantic space to obtain a semantic alignment feature, which represents the consistency of object category, environmental attributes, spatial relationships and scene style. Natural language reasoning is performed on the semantic alignment features to obtain the second evaluation result.
11. The method according to claim 8, characterized in that, Using a third evaluation model, the scene richness of the sample scene graph is evaluated, and the third evaluation results are obtained, including: Input the sample scene image and evaluation prompts into the third evaluation model, and perform the following operations in the third evaluation model: The sample scene image is subjected to multi-level feature extraction using the fifth image encoder to obtain multi-scale visual features, wherein the multi-scale visual features are used to analyze the three-dimensional structural complexity of the wall, the diversity of spatial shapes, the number of material types and the number of identifiable objects. The evaluation prompts are encoded using a third text encoder to obtain third text features; Using the third text feature as the evaluation standard, the multi-scale visual features are analyzed to obtain scene content richness score and scene structure richness score; The scene content richness score and the scene structure richness score are fused to obtain the third evaluation result.
12. The method according to claim 8, characterized in that, Before inputting the sample scene prompts into the image generation model, the method further includes: The format of the prompt words in the sample scene is evaluated based on the target format as the criterion, and the format evaluation results are obtained. Accordingly, a reward signal is generated based on at least one of the first evaluation result, the second evaluation result, and the third evaluation result, including: The reward signal is generated based on at least one of the first evaluation result, the second evaluation result, and the third evaluation result, combined with the formatted evaluation result.
13. The method according to any one of claims 1-4, 7, and 9-12, characterized in that, The text generation model generates multiple sample scene prompts for the same sample image, and the multiple sample scene prompts correspond to multiple sample scene images. The evaluation model evaluates the sample scene images from a spatial aesthetic perspective and generates a reward signal based on the evaluation results, including: The multiple sample scene images were evaluated separately, resulting in multiple evaluation results; The multiple evaluation results are sorted, and the relative weights corresponding to each of the multiple evaluation results are determined based on the sorting results. The relative weights are used to adjust the contribution of the sample scene graph in the model parameter update process. The reward signal is generated based on the multiple evaluation results and their respective relative weights.
14. A method for generating a scene graph, characterized in that, include: Obtain the target image including the target object; The target image is input into a text generation model to generate scene prompts, so as to obtain the target scene prompts corresponding to the target image. The target scene prompts describe the target scene information that needs to be generated by the model in text form. The target image and the target scene prompt are input into the image generation model to generate a target scene image containing the target object and the target scene information.
15. A method for generating a scene graph, characterized in that, include: The image generation interface is displayed, and the image generation interface includes an image upload control; In response to a trigger operation on the image upload control, the target image uploaded by the user is obtained, the target image including a target object; The target image is input into a text generation model to generate scene prompts, so as to obtain the target scene prompts corresponding to the target image. The target scene prompts describe the target scene information that needs to be generated by the model in text form. The target image and the target scene prompt are input into the image generation model to generate a target scene image containing the target object and the target scene information.
16. A computing device, characterized in that, include: A memory and a processor; wherein the memory stores executable code, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 1-15.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores executable code that, when executed by a processor of a computing device, causes the processor to perform the method as described in any one of claims 1-15.
18. A computer program product, characterized in that, include: A computer program / instruction that, when executed by a processor, causes the processor to perform the steps of the method according to any one of claims 1-15.