Method, system and device for generating decoration effect picture based on multi-modal artificial intelligence
Patent Information
- Application Number
- CN202610946484.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
由于装修效果图需要在基础空间图片所呈现的原始空间基础上生成,如果用户补充的装修需求信息没有与基础空间图片中的空间内容建立明确对应关系,图像生成过程就容易将不同需求混合处理,使部分需求被弱化、遗漏或者作用到不合适的图像区域,进而出现原始空间结构被明显改变、用户选择的装修内容体现不稳定、家具饰品或素材与空间位置不匹配、整体画面效果与用户预期不一致等情况
[0015]通过先获取基础空间图像和多模态装修需求信息,并分别进行预处理;再对预处理后的基础空间图像进行室内场景分割,得到基础空间图像中的空间固定区域和装修可变区域;随后对预处理后的多模态装修需求信息进行分析,确定对应空间固定区域的固定区域需求、对应装修可变区域的可变区域需求以及同时对应空间固定区域和装修可变区域的整体协同需求;再对固定区域需求、可变区域需求和整体协同需求进行融合处理,生成装修生成约束条件;最后基于基础空间图像和装修生成约束条件调用预设的图像生成模型生成候选装修效果图,并对候选装修效果图进行一致性筛选得到最终装修效果图。由此可见,本申请并不是将基础空间图像和多模态装修需求信息简单合并后直接输入图像生成模型,而是先将基础空间图像划分为空间固定区域和装修可变区域,再将多模态装修需求信息对应分析为固定区域需求、可变区域需求和整体协同需求,使不同来源、不同作用范围的装修需求能够与基础空间图像中的不同区域建立对应关系;在此基础上,通过融合处理生成装修生成约束条件,使图像生成模型在基础空间图像上生成候选装修效果图时受到区域和需求的共同约束,并通过一致性筛选从候选结果中确定最终装修效果图。因此,该方案能够使用户通过多模态方式表达的装修需求更准确地作用于基础空间图像中的相应区域,减少不同需求被混合处理、需求作用区域不准确、原始空间内容被不合理改变以及生成结果与用户实际需求不一致的情况,从而提高装修效果图生成结果对用户多模态装修需求的反映准确性和稳定性。
Smart Images

Figure CN122821030A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent generation technology, and in particular to a method, system and device for generating interior decoration renderings based on multimodal artificial intelligence. Background Technology
[0002] As the demand for high-quality living spaces increases, users often want to preview the final look of their space before decorating, furnishing, and remodeling. This helps them determine the decorating style, furniture configuration, lighting effects, and decorative elements. Traditional interior design renderings are typically generated by designers based on floor plans, interior photos, or design sketches through modeling, layout, and rendering. While this method can visually showcase the finished effect, it relies heavily on the designer's professional experience and modeling / rendering tools, resulting in a long production cycle and high modification costs. With the development of image generation technology, AI-based interior design rendering generation is increasingly being applied to interior design scenarios. Users can upload basic space images and input their decoration requirements to have the system automatically generate corresponding renderings, thus lowering the barrier to entry for previewing interior design effects.
[0003] Current AI-powered interior design rendering generation methods typically involve users uploading interior photos, design sketches, or other basic spatial images, along with text descriptions of the desired style or renovation effect. Upon receiving the basic spatial images and text descriptions, the system treats the images as the target image and the text descriptions as generation prompts. Then, an image generation model stylizes or modifies the content of the target image to produce the corresponding interior design rendering. In this approach, the image generation model can generate preview images with a certain visual appeal based on the interior scene content provided by the basic spatial images and the renovation intentions conveyed by the text descriptions.
[0004] However, in actual interior design rendering scenarios, users' needs for the desired effect cannot usually be fully expressed through a single text description. Besides basic spatial images and text descriptions, users may further supplement their design requirements, such as by selecting space type, design style, lighting type, furniture and accessories type, or choosing specific materials from a resource library. Existing generation methods, when processing these multiple types of design requirements, typically still rely primarily on basic spatial images and text descriptions as input, or simply merge the user-supplemented design requirements into a single generation prompt. This fails to establish a stable association between design requirements from different sources and the spatial content in the basic spatial image. Since interior design renderings are generated based on the original space presented in the basic spatial image, if the user-supplemented design requirements do not establish a clear correspondence with the spatial content in the basic spatial image, the image generation process can easily mix different requirements, weakening, omitting, or applying some requirements to inappropriate image areas. This can lead to situations such as significant alterations to the original spatial structure, unstable representation of user-selected design content, mismatch between furniture, accessories, or materials and their spatial placement, and an overall image effect inconsistent with the user's expectations. Therefore, how to effectively link multimodal decoration needs information with basic spatial images so that the generated decoration renderings can accurately reflect the user's actual decoration needs has become a problem that needs to be solved. Summary of the Invention
[0005] This application provides a method, system, and device for generating interior design renderings based on multimodal artificial intelligence. It effectively correlates multimodal interior design requirements with basic spatial images, enabling the generated renderings to accurately reflect the user's actual interior design needs. This application provides the following technical solutions: Firstly, this application provides a method for generating interior design renderings based on multimodal artificial intelligence, the method comprising: Acquire basic spatial images and multimodal decoration requirement information, and preprocess the basic spatial images and multimodal decoration requirement information respectively; The preprocessed base space image is segmented into an indoor scene to obtain a fixed spatial region and a variable decoration region in the base space image. The preprocessed multimodal decoration demand information is analyzed to determine the fixed area demand corresponding to the fixed spatial area, the variable area demand corresponding to the variable decoration area, and the overall coordination demand corresponding to both the fixed spatial area and the variable decoration area. The fixed area requirements, the variable area requirements, and the overall collaborative requirements are integrated to generate decoration generation constraints. Based on the basic spatial image and the decoration generation constraints, a preset image generation model is invoked to generate candidate decoration effect images, and the candidate decoration effect images are subjected to consistency screening to obtain the final decoration effect image.
[0006] In one specific implementation scheme, the acquisition of basic spatial images and multimodal decoration demand information, and the preprocessing of the basic spatial images and the multimodal decoration demand information respectively, includes: After receiving the base spatial image, the file data of the base spatial image is read first. It is determined whether the file data can be parsed into image data normally. If it can be parsed into image data normally, the image width, image height, image orientation information and color channel information of the base spatial image are obtained. After the file data reading is completed, the image orientation is corrected, the image size is unified and the color format is converted into the base spatial image to obtain the preprocessed base spatial image. The multimodal decoration requirement information refers to the requirement data input or selected by the user around the process of generating decoration renderings; after receiving the multimodal decoration requirement information, the multimodal decoration requirement information is divided into selection-type requirement information, text-type requirement information, and image-type requirement information according to the data presentation format; For selection-based requirements, selection requirement records are generated according to preset fields; for text-based requirements, the text content entered by the user is processed to obtain text requirement records; for image-based requirements, the source images are read, their size is standardized, and their color format is converted to obtain source image records; selection requirement records, text requirement records, and source image records together constitute the preprocessed multimodal decoration requirement information.
[0007] In one specific implementation scheme, the step of performing indoor scene segmentation on the preprocessed base space image to obtain fixed spatial regions and variable decoration regions in the base space image includes: The preprocessed base space image is converted to the Lab color space, and the SLIC superpixel segmentation algorithm is used to perform superpixel segmentation on the converted base space image to obtain a superpixel set. A saliency analysis is performed on the base spatial image to obtain a saliency map corresponding to the base spatial image. The saliency map is used to represent the visual prominence of each pixel in the base spatial image relative to the whole image. Based on the superpixel set and the saliency map, the region saliency corresponding to each superpixel unit is calculated. The region saliency is calculated by the average value of the values of each pixel in the superpixel unit in the saliency map. If the regional saliency of a superpixel unit is greater than or equal to a preset saliency threshold, then the superpixel unit is recorded as an object superpixel; if the regional saliency of a superpixel unit is less than the preset saliency threshold, then the superpixel unit is recorded as a background superpixel. The adjacent background superpixels are merged to obtain the fixed spatial region, and the adjacent object superpixels are merged to obtain the variable decoration region. The fixed spatial region represents a continuous image region in the base spatial image that serves as the interior background, and the variable decoration region represents a local object region in the base spatial image that is presented independently relative to the interior background.
[0008] In a specific feasible implementation, the analysis of the preprocessed multimodal decoration demand information to determine the fixed area demand corresponding to the fixed spatial area, the variable area demand corresponding to the variable decoration area, and the overall coordinated demand corresponding to both the fixed spatial area and the variable decoration area includes: The selection requirement record and the text requirement record are merged into a text form requirement, and dependency parsing and semantic role labeling are used to semantically decompose the text form requirement to obtain at least one text requirement unit. Based on the mask corresponding to the fixed spatial area, generate a description text for the fixed spatial area; and based on the mask corresponding to the variable decoration area, generate a description text for the variable decoration area. Calculate the first one respectively Semantic similarity between each text requirement unit and the text describing the fixed spatial region and the Semantic similarity between each text requirement unit and the text describing the variable decoration area and based on and Calculate the first Regional bias value of each text demand unit and effective correlation value ,in: ; ; like If the text value is less than the preset effective text threshold, then discard the first one. Each text requirement unit; if If the value is greater than or equal to the preset text validity threshold, then according to... The comparison result with the text region bias threshold will be the first Each text requirement unit is determined as a fixed area requirement, a variable area requirement, or an overall collaborative requirement; The source images are recorded as image requirement units, and the fixed region image prototype vector, variable region image prototype vector, and overall collaborative image prototype vector are determined based on a pre-constructed image calibration sample set. The first The source image corresponding to the first image requirement unit is converted into an image semantic vector, and the first image is calculated according to the following formula. The image demand unit and the first Image semantic similarity between image prototype vectors: ; in, This indicates the image category corresponding to the requirement of a fixed area. This indicates the image category corresponding to the variable region requirement. This indicates the image category corresponding to the overall collaborative requirements. Indicates the first The image semantic vector corresponding to each image requirement unit Indicates the first Image prototype vector; According to the The maximum image similarity among the three types of semantic similarity corresponding to each image demand unit, and the difference between it and the second largest image similarity, will be used to determine the first... Each image demand unit is determined as a fixed region demand, a variable region demand, an overall collaborative demand, or an unoriented image demand, and the unoriented image demand is discarded.
[0009] In a specific feasible implementation, the process of integrating the fixed area requirements, the variable area requirements, and the overall collaborative requirements to generate decoration generation constraints includes: A fixed region mask is generated based on the fixed spatial region, and a content editing mask is generated based on the variable decoration region. Both the fixed region mask and the content editing mask are in the same pixel coordinate system as the base spatial image. A fixed-area prompt text is generated based on the text requirement units in the fixed-area requirement; a variable-area prompt text is generated based on the text requirement units in the variable-area requirement; an overall prompt text is generated based on the text requirement units in the overall collaboration requirement; and the text requirement units in the overall collaboration requirement are associated with the fixed-area prompt text and the variable-area prompt text. Material reference conditions are generated based on image requirement units in the fixed area requirement, the variable area requirement, and the overall collaborative requirement. Specifically, image requirement units belonging to the fixed area requirement are associated with the fixed area mask, image requirement units belonging to the variable area requirement are associated with the content editing mask, and image requirement units belonging to the overall collaborative requirement are associated with the whole image mask. The fixed area mask, the content editing mask, the fixed area prompt text, the variable area prompt text, the overall prompt text, and the material reference conditions are combined to form the decoration generation constraints.
[0010] In one specific feasible implementation, the preset image generation model is obtained in the following manner: Acquire training samples, which include sample base spatial images, sample multimodal decoration demand information, and sample decoration effect images; The sample base space image is segmented into an indoor scene to obtain a fixed sample space region and a variable sample decoration region; the sample multimodal decoration requirement information is analyzed to obtain the sample fixed region requirement, the sample variable region requirement, and the sample overall coordination requirement; and the sample fixed region requirement, the sample variable region requirement, and the sample overall coordination requirement are merged into the sample decoration generation constraints. The sample base space image and the sample decoration generation constraints are used as model inputs, and the sample decoration effect image is used as the target output. The initial image editing and generation model is trained or fine-tuned to obtain the preset image generation model. The constraints for generating the sample decoration include a fixed area mask, a content editing mask, a fixed area prompt text, a variable area prompt text, an overall prompt text, and reference conditions for sample materials.
[0011] In one specific implementation scheme, the step of performing consistency screening on the candidate decoration renderings to obtain the final decoration rendering includes: For the Based on the candidate decoration renderings, the structural retention value of the fixed area is calculated using the fixed area mask. The constraint response value is calculated based on the set of constraint terms in the decoration generation constraint conditions. And based on the fixed region structure retention value and the constraint response value Calculate the overall consistency value ; Wherein, the fixed region structure retention value Based on the aforementioned basic spatial image and the first The normalized gradient difference of each candidate decoration rendering image within the effective pixel position of the fixed area mask is determined; The constraint response value The minimum value among the semantic matching values between the candidate image content and the corresponding constraint content within the mask corresponding to each constraint term is determined. The overall consistency value Determined according to the following formula: ; The candidate design rendering with the highest overall consistency score will be selected as the final design rendering.
[0012] Secondly, this application provides a decoration rendering generation system based on multimodal artificial intelligence, which adopts the following technical solution: A system for generating interior design renderings based on multimodal artificial intelligence, comprising: The data acquisition module is used to acquire basic spatial images and multimodal decoration demand information, and to preprocess the basic spatial images and multimodal decoration demand information respectively; The scene segmentation module is used to segment the preprocessed base space image into indoor scenes to obtain fixed spatial areas and variable decoration areas in the base space image. The demand analysis module is used to analyze the preprocessed multimodal decoration demand information to determine the fixed area demand corresponding to the fixed space area, the variable area demand corresponding to the variable decoration area, and the overall collaborative demand corresponding to both the fixed space area and the variable decoration area. The constraint generation module is used to integrate the fixed area requirements, the variable area requirements, and the overall collaborative requirements to generate decoration generation constraints. The image generation module is used to generate candidate decoration effect images by calling a preset image generation model based on the basic spatial image and the decoration generation constraints, and to perform consistency screening on the candidate decoration effect images to obtain the final decoration effect image.
[0013] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement a method for generating interior decoration renderings based on multimodal artificial intelligence as described in the first aspect.
[0014] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a method for generating interior decoration renderings based on multimodal artificial intelligence as described in the first aspect.
[0015] First, basic spatial images and multimodal decoration requirements information are acquired and preprocessed separately. Then, the preprocessed basic spatial images are segmented into indoor scenes to obtain fixed spatial areas and variable decoration areas. Subsequently, the preprocessed multimodal decoration requirements information is analyzed to determine the fixed area requirements corresponding to the fixed spatial areas, the variable area requirements corresponding to the variable decoration areas, and the overall collaborative requirements corresponding to both fixed spatial areas and variable decoration areas. Then, the fixed area requirements, variable area requirements, and overall collaborative requirements are fused to generate decoration generation constraints. Finally, based on the basic spatial images and decoration generation constraints, a preset image generation model is called to generate candidate decoration effect images, and consistency screening is performed on the candidate decoration effect images to obtain the final decoration effect image. Therefore, this application does not simply merge the basic spatial image and multimodal decoration demand information and directly input them into the image generation model. Instead, it first divides the basic spatial image into fixed spatial regions and variable decoration regions, and then analyzes the multimodal decoration demand information into fixed region demands, variable region demands, and overall collaborative demands. This allows decoration demands from different sources and with different scopes of application to establish a correspondence with different regions in the basic spatial image. Based on this, decoration generation constraints are generated through fusion processing, so that the image generation model is subject to the joint constraints of regions and demands when generating candidate decoration renderings on the basic spatial image. The final decoration rendering is determined from the candidate results through consistency screening. Therefore, this scheme enables users' decoration demands expressed in a multimodal manner to more accurately act on the corresponding regions in the basic spatial image, reducing the situation where different demands are mixed, the demand application areas are inaccurate, the original spatial content is unreasonably altered, and the generated results are inconsistent with the user's actual needs. This improves the accuracy and stability of the decoration renderings in reflecting the user's multimodal decoration demands.
[0016] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the method for generating interior decoration renderings based on multimodal artificial intelligence in this application embodiment.
[0018] Figure 2 This is a schematic diagram of the data input interface for generating decoration renderings in an embodiment of this application.
[0019] Figure 3 This is a schematic diagram of the overall process of the decoration rendering generation method based on multimodal artificial intelligence in the embodiments of this application.
[0020] Figure 4This is a structural block diagram of the decoration rendering generation system based on multimodal artificial intelligence in the embodiments of this application.
[0021] Figure 5 This is a block diagram of an electronic device generated from a decoration rendering based on multimodal artificial intelligence, as described in this application embodiment. Detailed Implementation
[0022] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.
[0023] Optionally, this application uses the method for generating decoration renderings based on multimodal artificial intelligence provided in various embodiments as an example for application in electronic devices. The electronic device is a terminal or a server. The terminal can be a computer, tablet computer, etc. This embodiment does not limit the type of electronic device.
[0024] Reference Figure 1 This is a flowchart illustrating a method for generating interior design renderings based on multimodal artificial intelligence, provided in one embodiment of this application. The method includes at least the following steps: Step S101: Obtain basic spatial images and multimodal decoration requirement information, and preprocess the basic spatial images and multimodal decoration requirement information respectively.
[0025] In step S101, the basic spatial image and multimodal decoration requirement information are acquired, and preprocessed separately. This step converts the user-provided raw image data and decoration requirement data into a data format that can be uniformly read, recorded, and retrieved. The basic spatial image represents the original interior space from which decoration renderings are generated, while the multimodal decoration requirement information represents the user's requirements for the decoration renderings in the form of images, text, or options. Since the basic spatial image and multimodal decoration requirement information have different data formats, without preprocessing, issues such as inconsistent image sizes, incorrect image orientation, inconsistent color channels, inconsistent option fields, invalid characters in text content, and missing source identifiers for source images can easily arise, affecting the stability of the processing. Therefore, this step preprocesses the basic spatial image and multimodal decoration requirement information separately to obtain input data with a uniform format.
[0026] Specifically, the base space image is an indoor space image that the user takes, uploads, or selects from their local photo album. Upon receiving the base space image, the system first reads its file data and determines whether it can be correctly parsed into image data. If it cannot be parsed, a prompt to re-enter is generated. If it can be correctly parsed into image data, the system obtains the image width, image height, image orientation, and color channel information of the base space image.
[0027] After reading the file data, the base spatial image undergoes image orientation correction, image size unification, and color format conversion. Image orientation correction involves rotating or flipping the base spatial image based on the orientation information it carries, ensuring the indoor space is displayed in the user's normal viewing orientation when captured or uploaded. If no orientation information is present in the base spatial image, its current display orientation is maintained. Image size unification involves adjusting the base spatial image to a preset input size. During adjustment, the image is first scaled proportionally to its original aspect ratio, ensuring the scaled image falls within the preset input size range. Then, preset background pixels are filled into the boundary areas not covered by the image, giving the adjusted base spatial image a fixed width and height. Color format conversion involves converting the base spatial image to RGB color format. When the base spatial image contains an alpha channel, the area corresponding to the alpha channel is first composited with a preset background color before conversion to RGB color format. After these processes, a preprocessed base spatial image is obtained.
[0028] Multimodal renovation requirement information comprises user input or selection data related to the process of generating renovation renderings. Upon receiving this information, it is categorized into selection-based, text-based, and image-based requirements according to their format. Selection-based requirements are those obtained through user interface controls; text-based requirements are those entered via text input boxes; and image-based requirements are those uploaded, photographed, or selected from a resource library. This categorization ensures that different formats of renovation requirement information are recorded according to corresponding rules, preventing the mixing of options, text, and images that could lead to unclear requirement origins.
[0029] For selection-based requirement information, selection requirement records are generated according to preset fields. Each selection requirement record includes a selection category and selection content. The selection category represents the requirement dimension corresponding to that selection type, and the selection content represents the actual option selected by the user. For example, when a user selects "living room" in the space type control, the selection category record is "space type," and the selection content record is "living room." When a user selects "modern minimalist" in the decoration style control, the selection category record is "decoration style," and the selection content record is "modern minimalist." In this way, the requirements entered by the user through interface options are converted into selection requirement records with unified fields.
[0030] For text-based requirements, the user-inputted text is processed to create a text requirement record. Text processing includes removing leading and trailing whitespace, merging consecutive spaces, removing duplicate punctuation, and deleting default placeholder text in the input box. After processing, the text content and input time are recorded to form a text requirement record. If the user does not input any text, the text requirement record is empty, which does not affect the acquisition and processing of other types of requirements.
[0031] For image-based information requirements, the source images are read, resized, and converted in color format to obtain a source image record. Image reading refers to reading source images uploaded by the user, captured by the user, or selected from the source library, and determining whether the source image can be correctly parsed into image data; if it cannot be parsed into image data, a prompt for re-entry of the source image is generated. The processing rules for resizing and color format conversion are consistent with the processing rules for basic spatial images. After processing, the source image, source identifier, and processed source image data are recorded to form a source image record. The source image is used to distinguish whether the source image comes from user upload, user capture, or selection from the source library, and the source identifier is used to distinguish different source images.
[0032] After the above processing, the selected requirement records, text requirement records, and material image records together constitute the preprocessed multimodal decoration requirement information. Thus, the basic spatial images are processed into image data with uniform size, orientation, and color format, and the multimodal decoration requirement information is processed into data records with clear sources and uniform fields, thereby ensuring that the image content and decoration requirement content input by the user can be accurately read and retrieved.
[0033] In one specific embodiment Figure 2 This is a schematic diagram of the input interface for generating decoration renderings in an embodiment of this application. (Refer to...) Figure 2The input interface includes a basic space image input area, multiple decoration requirement input areas, and an AI-generated button. The basic space image input area is used to receive basic space images taken or uploaded by the user; the multiple decoration requirement input area is used to receive multimodal decoration requirement information provided by the user through interface selection, text input, material upload, or material library selection; the AI-generated button is used to trigger data acquisition operation after the user completes data input. Figure 2 The space type, style selection, lighting type, addition of furniture or decorations, material image information, and text requirement information are only one implementation of the input interface and do not constitute a limitation on the specific types of multimodal decoration requirement information.
[0034] Step S102: Perform indoor scene segmentation on the preprocessed basic spatial image to obtain the fixed spatial area and the variable decoration area in the basic spatial image.
[0035] In step S102, the preprocessed base space image obtained in step S101 is segmented into an interior scene to obtain fixed spatial regions and variable decoration regions in the base space image. This step converts the base space image from a whole image to a region image and distinguishes the interior background region from the local object region in the base space image. The fixed spatial region represents the continuous image region in the base space image that serves as the interior background, while the variable decoration region represents the local object region in the base space image that is presented independently relative to the interior background. After this step, the background content and local object content in the base space image are recorded separately, avoiding the mistaken treatment of local objects as interior background when generating decoration renderings.
[0036] Specifically, the preprocessed base spatial image is converted to the Lab color space, and the existing SLIC superpixel segmentation algorithm is used to perform superpixel segmentation on the converted base spatial image, resulting in a set of superpixels. The SLIC superpixel segmentation algorithm generates superpixel units based on the color similarity and spatial proximity between pixels, with each superpixel unit having a continuous pixel range. The reason for choosing the SLIC superpixel segmentation algorithm is that the base spatial image usually includes a large area of indoor background and multiple local objects. Ordinary threshold segmentation is easily affected by changes in lighting and darkness, edge detection can only obtain the edges and cannot directly obtain the regions, and object detection tends to output object boxes and is difficult to represent background areas. The SLIC superpixel segmentation algorithm can first convert the image into superpixel units with boundaries that closely fit the image content and have a clear area range, thus making it suitable as an initial processing method for indoor scene region segmentation. Converting the base spatial image to the Lab color space is to separate the brightness and color information, so that superpixel segmentation can utilize both brightness and color differences.
[0037] After obtaining the superpixel set, saliency analysis is performed on the base spatial image to obtain a saliency map corresponding to the base spatial image. The saliency map is used to represent the visual prominence of each pixel in the base spatial image relative to the entire image. Specifically, based on the color distribution of the base spatial image in the Lab color space, the L-channel, a-channel, and b-channel values of each pixel in the Lab color space are obtained, and the three-dimensional Euclidean distance between the color value of each pixel and the color values of other pixels in the base spatial image is calculated. The weighted sum of multiple three-dimensional Euclidean distances corresponding to the same pixel is used as the color difference value of that pixel. Subsequently, the color difference values corresponding to each pixel are smoothed to obtain the saliency map. Therefore, the higher the value in the saliency map, the more prominent the region where the corresponding pixel is located relative to the overall color distribution of the base spatial image; the lower the value in the saliency map, the closer the region where the corresponding pixel is located is to the overall background distribution of the base spatial image.
[0038] Based on the superpixel set and the saliency map, the region saliency corresponding to each superpixel unit is calculated. Region saliency is calculated as the average value of each pixel within that superpixel unit in the saliency map. If the region saliency of a superpixel unit is greater than or equal to a preset saliency threshold, the superpixel unit is recorded as an object superpixel; if the region saliency of a superpixel unit is less than the preset saliency threshold, the superpixel unit is recorded as a background superpixel. The preset saliency threshold is recorded in the parameter configuration file and remains consistent throughout the same batch of image processing. Through this process, the superpixel set is divided into background superpixels and object superpixels.
[0039] The preset saliency threshold is determined in advance based on indoor image calibration samples. Specifically, multiple indoor images that are the same as or similar to the scene generated from the current decoration rendering are selected as calibration sample images. The calibration sample images undergo the same color space conversion, SLIC superpixel segmentation, and saliency analysis as the base spatial images to obtain the regional saliency of each superpixel unit. Then, based on the labeled spatial background area and local object area in the calibration sample images, the regional saliency distribution of the two types of superpixel units is statistically analyzed. The saliency value that can better distinguish between the spatial background area and the local object area is determined as the preset saliency threshold. The determined preset saliency threshold is recorded in the parameter configuration file and remains consistent throughout the same batch of image processing. This avoids arbitrary fluctuations in the threshold due to local illumination, local color differences, or local texture changes in a single image, improving the stability of the results for dividing fixed spatial areas and variable decoration areas.
[0040] Adjacent background superpixels are merged to obtain spatially fixed regions. Specifically, superpixel units that are adjacent in position and all belong to background superpixels are merged into the same background region, and the merged background region is determined as the spatially fixed region. The spatially fixed region originates from background superpixels with low saliency, and therefore is used to represent visually relatively stable indoor background content in the base spatial image. Adjacent object superpixels are merged to obtain decoratively variable regions. Specifically, superpixel units that are adjacent in position and all belong to object superpixels are merged into the same local object region, and the merged local object region is determined as the decoratively variable region. The decoratively variable region originates from object superpixels with high saliency, and therefore is used to represent local object content that is presented independently relative to the indoor background in the base spatial image. Since the division of spatially fixed regions and decoratively variable regions is based on regional saliency, rather than whether the image region is connected to the image boundary, large objects near the image edge can still be classified as decoratively variable regions based on their visual prominence relative to the indoor background.
[0041] The system records the region outlines and positions corresponding to both fixed spatial regions and variable decorative regions. The region outline is formed by the outer boundary pixels of the corresponding region, and the region position is determined by the minimum x-coordinate, minimum y-coordinate, maximum x-coordinate, and maximum y-coordinate of the region outline. For fixed spatial regions, the region outline and position represent the image range of the interior background within the base spatial image; for variable decorative regions, the region outline and position represent the image range of local objects within the base spatial image. After this processing, the base spatial image is divided into fixed spatial regions and variable decorative regions. The former represents the interior background, and the latter represents local objects presented independently relative to the interior background, thus giving the background content and local object content in the base spatial image clear image ranges.
[0042] Step S103: Analyze the preprocessed multimodal decoration demand information to determine the fixed area demand for the corresponding fixed space area, the variable area demand for the corresponding variable decoration area, and the overall coordination demand for both the fixed space area and the variable decoration area.
[0043] In step S103, the preprocessed multimodal decoration requirement information obtained in step S101 is analyzed to determine the fixed area requirements corresponding to the fixed spatial area, the variable area requirements corresponding to the variable decoration area, and the overall collaborative requirements corresponding to both the fixed spatial area and the variable decoration area. The preprocessed multimodal decoration requirement information includes text-based requirements obtained by merging selection requirement records and text requirement records, and image-based requirements formed from material image records. Since text-based requirements express the user's decoration intentions in a verbal way, and image-based requirements express the decoration references provided by the user in a visual way, this step uses different analysis methods for text-based and image-based requirements, and the analysis results are uniformly categorized into fixed area requirements, variable area requirements, and overall collaborative requirements.
[0044] For textual requirements, dependency parsing and semantic role labeling are first used to semantically decompose the textual requirements, resulting in at least one textual requirement unit. Specifically, dependency parsing is used to determine the modification and dominance relationships between words in the text, while semantic role labeling is used to determine the correspondence between action intent, the object being acted upon, and the effect description. Based on this, content with a complete correspondence between the object being acted upon, action intent, and effect description is divided into a textual requirement unit. If a text contains multiple independent decoration intents, it is split into multiple textual requirement units; if multiple objects being acted upon correspond to the same effect description, they are retained as a single textual requirement unit. For textual requirements formed from selection requirement records, the selection category and selection content are combined into a single textual requirement unit; for textual requirements formed from textual requirement records, they are split according to the results of dependency parsing and semantic role labeling. The advantage of this approach is that the splitting is based on the relationships within the language structure, rather than simply relying on separators or keywords, enabling a more accurate distinction between independent and collaborative decoration intents within a user description.
[0045] After obtaining the text requirement unit, based on the fixed spatial region and the variable decoration region obtained in step S102, descriptive text for the fixed spatial region and descriptive text for the variable decoration region are generated respectively. Specifically, based on the mask corresponding to the fixed spatial region, the area ratio, region continuity, saliency mean, dominant color distribution, and boundary stability of the fixed spatial region are statistically analyzed, and descriptive text for the fixed spatial region is generated according to a preset description template; based on the mask corresponding to the variable decoration region, the area ratio, region dispersion, saliency mean, dominant color distribution, and boundary independence of the variable decoration region are statistically analyzed, and descriptive text for the variable decoration region is generated according to a preset description template. The preset description template includes a region source field, a region visual field, and a region generation function field. The region source field records the segmentation result of the region from step S102, the region visual field records the visual features obtained from the mask statistics, and the region generation function field records the adjustment range of the region during the generation of the decoration rendering. In this way, the region description text is generated from the actual segmentation result of the basic spatial image, rather than using a pre-fixed classification vocabulary.
[0046] The first Each text requirement unit, the spatial fixed area description text, and the decoration variable area description text are converted into semantic vectors, and the semantic vectors are calculated for each. Semantic similarity between individual text requirement units and texts describing fixed spatial areas and variable decoration areas: ; ; in, Indicates the first The semantic similarity between a textual demand unit and a text describing a fixed spatial region. Indicates the first Semantic similarity between individual text requirement units and text describing variable decoration areas; Indicates the first The semantic vector corresponding to each text requirement unit A semantic vector representing a fixed spatial region describing the text. This represents the semantic vector corresponding to the text describing the variable areas of the decoration. This represents cosine similarity. Through the above calculation, semantic similarity is mapped to a value range of 0 to 1.
[0047] Furthermore, according to the first The semantic similarity between a text demand unit and two types of region description text is used to calculate the region bias value of the text demand unit: ; in, Indicates the first The region bias value of each text requirement unit This is a preset minimum positive number to avoid a denominator of 0. Since the numerator of this formula is the first... The semantic similarity between a textual demand unit and the text describing the variable area of decoration, minus the semantic similarity between it and the text describing the fixed area of space, is therefore, when Greater than When the region bias value is negative, it indicates that the text requirement unit is more biased towards a fixed spatial region; when Greater than When the region bias value is positive, it means that the text requirement unit is more biased towards the variable decoration area; when the two are close, the region bias value is close to 0, which means that the text requirement unit is associated with both types of areas at the same time and is suitable as an overall collaborative requirement.
[0048] To avoid misjudging textual requirement units that lack effective association with both types of regions as overall collaborative requirements, further calculations are performed. Valid association values for each text requirement unit: ; in, Indicates the first The valid association value of each text requirement unit. If If the text is less than the preset valid text threshold, then the first... Each text requirement unit is identified as an undirected text requirement, and this undirected text requirement is discarded; if If the value is greater than or equal to the preset effective text threshold, then the region bias value is applied to the first... The text requirement units are categorized. Specifically, when At that time, the first Each text requirement unit is determined as a fixed area requirement; when At that time, the first Each text requirement unit is determined as a variable region requirement; when At that time, the first Each textual requirement unit is identified as an overall collaborative requirement. Among them, This is the threshold for text region bias.
[0049] The reason for discarding undirected text requirements is that the semantic similarity between these text requirement units and both the spatial fixed region description text and the decoration variable region description text does not meet the effective association requirements, making it impossible to stably determine their regional scope in the base spatial image. If such text requirements were retained and included in the generation of decoration generation constraints, it would introduce requirement content lacking regional boundaries, making it difficult to define the scope of the generated constraints, thereby increasing the risk of misallocation of requirements, regional constraint conflicts, and unstable generation results. By discarding undirected text requirements, all text requirements participating in the generation of decoration generation constraints can have clear regional correspondences, thus improving the stability of the correspondence between multimodal decoration requirement information and the regions of the base spatial image.
[0050] Text region bias threshold The preset effective threshold for text is determined through a text labeling sample set. Specifically, a text labeling sample set is pre-constructed, including text samples labeled as fixed-area requirements, variable-area requirements, overall collaborative requirements, and undirected requirements; the region bias value and effective association value corresponding to each text sample are calculated respectively. For text samples labeled as overall collaborative requirements, the high quantile value of their absolute region bias value is calculated; for text samples labeled as fixed-area requirements and variable-area requirements, the low quantile value of their absolute region bias value is calculated; the median value between the high quantile value and the low quantile value is determined as the text region bias threshold. Among them, the regional bias value of the overall collaborative demand sample is usually close to 0, while the absolute value of the regional bias value of the fixed regional demand sample and the variable regional demand sample is usually larger. Therefore, the text regional bias threshold is used to distinguish between dual-region collaborative orientation and single-region bias orientation. The preset effective text threshold is determined based on the distribution of effective association values of undirected demand samples and directed demand samples. Specifically, the median value between the high quantile of the effective association value of undirected demand samples and the low quantile of the effective association value of directed demand samples is determined as the preset effective text threshold.
[0051] For image-based requirements, each source image is treated as a single image requirement unit. Image-based requirements do not undergo textual semantic segmentation because each source image already constitutes a complete visual reference unit. Furthermore, image-based requirements do not use cross-modal similarity between image semantic vectors and region description text semantic vectors for classification, to avoid unstable attribution results due to modal differences between image visual content and textual descriptions. Therefore, for image-based requirements, intramodal similarity between image semantic vectors is used for processing.
[0052] Specifically, an image calibration sample set is pre-constructed, including image samples labeled as fixed-area requirements, variable-area requirements, and overall collaborative requirements. Image semantic vectors are extracted from each type of image sample, and the semantic vectors within the same category are averaged and normalized to obtain prototype vectors for fixed-area images, variable-area images, and overall collaborative images. The image calibration sample set originates from pre-organized decoration material samples or reference images with completed type labeling from historical tasks; therefore, the classification of image form requirements is based on the visual feature distribution of the calibrated image samples, rather than textual descriptions.
[0053] The first Each image requirement unit's corresponding source image is converted into an image semantic vector. And calculate the first according to the following formula The image demand unit and the first Image semantic similarity between image prototype vectors:
[0054] in, Indicates the first The image demand unit and the first Image semantic similarity between image prototype vectors; This indicates the image category corresponding to the requirement of a fixed area. This indicates the image category corresponding to the variable region requirement. This indicates the image category corresponding to the overall collaborative requirements; Indicates the first The image semantic vector corresponding to each image requirement unit; Indicates the first Image prototype vector; This represents cosine similarity. Using the above formula, the source image is compared with the three types of labeled image prototypes in the same modal semantics to obtain its visual similarity relative to the three types of requirements.
[0055] After obtaining the semantic similarity of the three types of images, the maximum value is selected as the maximum image similarity, and the corresponding requirement type is determined. If the maximum image similarity is less than a preset effective image threshold, it indicates that the source image lacks effective visual association with the prototype images of all three requirement types, and the image is then... The first image demand unit is identified as an undirected image demand and discarded. If the maximum image similarity is greater than or equal to a preset effective image threshold, then it is determined whether the difference between the maximum image similarity and the second largest image similarity is greater than or equal to a preset similarity interval threshold; if so, the first image demand unit is discarded. The first image requirement unit is determined as the requirement type corresponding to the highest image similarity; otherwise, it indicates that the material image does not show a clear single-region bias in visual semantics, and the first image requirement unit is determined as the requirement type corresponding to the highest image similarity. Each image requirement unit is identified as an overall collaborative requirement.
[0056] The effective image threshold and similarity interval threshold are determined using an image calibration sample set. Specifically, for both oriented and unoriented image samples in the image calibration sample set, the maximum image similarity is calculated, and the median value between the highest quantile of the maximum image similarity for the unoriented image samples and the lowest quantile value for the maximum image similarity for the oriented image samples is determined as the preset effective image threshold. For image samples correctly labeled as single-region requirements in the image calibration sample set, the similarity interval between their maximum and second-highest similarity values is calculated; for image samples representing overall collaborative requirements, their similarity interval is calculated; and the median value between their distribution boundaries is determined as the preset similarity interval threshold. In this way, the preset effective image threshold is used to exclude material images that lack effective visual association with the decoration requirement type, and the preset similarity interval threshold is used to distinguish between single-region image requirements and overall collaborative image requirements.
[0057] Textual and image-based requirements are not processed in the same way. Textual requirements require semantic segmentation first because a piece of text may contain multiple decoration intentions simultaneously; only by segmenting can the mixing of different intentions be avoided. Image-based requirements, on the other hand, typically use a single source image as a visual reference unit and are not suitable for segmentation according to textual logic. Textual requirements are best judged by their relative bias to the text describing the underlying spatial image region, as the text expresses the user's linguistic requirements for the region. Image-based requirements are best judged by their semantic similarity to the image prototype formed by the image calibration sample, as the image expresses visual reference content. Therefore, textual and image requirements employ different analysis logics, but both ultimately output fixed region requirements, variable region requirements, and overall collaborative requirements, enabling the unified construction of decoration generation constraints.
[0058] After the above processing, text and image requirement units identified as fixed-area requirements are categorized into fixed-area requirements; text and image requirement units identified as variable-area requirements are categorized into variable-area requirements; text and image requirement units identified as overall collaborative requirements are categorized into overall collaborative requirements; and undirected text and image requirements are discarded. Therefore, multimodal decoration requirement information is not simply concatenated into a unified prompt text, but its scope is determined separately based on the semantic region bias of text requirements and the visual semantic attribution of image requirements. This reduces the interference of requirements without clear regional correspondence on the image generation process and improves the stability of the correspondence between decoration generation constraints and basic spatial image regions.
[0059] Step S104: Integrate the fixed area requirements, variable area requirements, and overall collaborative requirements to generate decoration generation constraints.
[0060] In step S104, the fixed-area requirements, variable-area requirements, and overall collaborative requirements obtained in step S103 are fused to generate decoration generation constraints. This step converts the multimodal decoration requirements with defined scopes into data conditions that can be input into the preset image generation model along with the base spatial image. Since unoriented text requirements and unoriented image requirements have been discarded in step S103, this step only processes fixed-area requirements, variable-area requirements, and overall collaborative requirements with clear regional correspondences, avoiding requirements without clear regional scopes from entering the decoration generation constraints.
[0061] The constraints for decoration generation include region masking conditions, text prompting conditions, and material reference conditions. Region masking conditions limit the scope of different requirements during decoration generation; text prompting conditions limit the content to be generated; and material reference conditions limit the source images that need to be referenced during the generation process and their scope. Through the combination of these conditions, fixed region requirements, variable region requirements, and overall collaborative requirements no longer participate in image generation as scattered requirement records, but are instead organized into corresponding input conditions for regions, text, and materials.
[0062] Specifically, a fixed-area mask is generated based on the fixed spatial area obtained in step S102, and a content editing mask is generated based on the variable decoration area obtained in step S102. Both the fixed-area mask and the content editing mask have the same width and height as the base spatial image and are located in the same pixel coordinate system as the base spatial image. The fixed-area mask is used to limit the scope of content related to the spatial background in terms of fixed area requirements and overall coordination requirements; the content editing mask is used to limit the scope of content related to local objects in terms of variable area requirements and overall coordination requirements. The fixed-area mask does not mean that the corresponding area will not undergo any visual changes during the generation process, but rather that when the area participates in the generation as a fixed spatial area, its scope is limited by the fixed-area mask; the content editing mask means that when the variable decoration area participates in the generation, its scope is limited by the content editing mask.
[0063] In one embodiment, the fixed region mask and the content editing mask are generated as follows: ; ; in, The coordinates in the fixed region mask are: The pixel values; The coordinates in the content editing mask are The pixel values; This represents the fixed spatial region obtained in step S102; This represents the variable decoration area obtained in step S102; a value of 1 indicates that the corresponding pixel position belongs to the valid area of the corresponding mask; a value of 0 indicates that the corresponding pixel position does not belong to the valid area of the corresponding mask. Through the above mask generation method, the fixed spatial area and the variable decoration area are converted into image input conditions with the same size as the basic spatial image.
[0064] Furthermore, fixed-area prompt text is generated based on text requirement units in fixed-area requirements, variable-area prompt text is generated based on text requirement units in variable-area requirements, and overall prompt text is generated based on text requirement units in overall collaboration requirements. Fixed-area prompt text describes the generation requirements for the area corresponding to the fixed-area mask, variable-area prompt text describes the generation requirements for the area corresponding to the content editing mask, and overall prompt text describes the overall generation requirements for the entire candidate decoration rendering. For text requirement units in overall collaboration requirements, they are written into the overall prompt text and synchronously associated with the fixed-area and variable-area prompt texts, ensuring that both fixed spatial areas and variable decoration areas are constrained by the overall visual collaboration requirements during generation.
[0065] When generating each prompt text, the source record identifier and requirement type identifier of each text requirement unit are retained. The source record identifier records whether the text requirement unit originates from a selection requirement record or a text requirement record, and the requirement type identifier records whether the text requirement unit belongs to a fixed area requirement, a variable area requirement, or an overall collaborative requirement. If the same source record has been edited multiple times to form multiple text contents, the text content finally confirmed by the user shall prevail. If the same prompt text contains multiple text requirement units, they are written in the order of user confirmation. By retaining the source record identifier and requirement type identifier, overall collaborative requirements and single area requirements can be distinguished, preventing overall collaborative requirements from being mistakenly considered to apply only to a single area after being written into the area prompt text.
[0066] Furthermore, reference conditions for source materials are generated based on image requirement units within the categories of fixed-area requirements, variable-area requirements, and overall collaborative requirements. For image requirement units identified as having fixed-area requirements, their source image data is used as fixed-area reference material, and their source reference mask is set as a fixed-area mask. For image requirement units identified as having variable-area requirements, their source image data is used as variable-area reference material, and their source reference mask is set as a content editing mask. For image requirement units identified as having overall collaborative requirements, their source image data is used as overall reference material, and their source reference mask is set as a whole-image mask, where all pixel positions of the base spatial image in the whole-image mask are recorded as valid values. In this way, source images are not generated indiscriminately as independent reference images, but are bound to the requirement type and corresponding scope determined in step S103, enabling the preset image generation model to reference the corresponding source images within a specified area.
[0067] Finally, the fixed area mask, content editing mask, fixed area prompt text, variable area prompt text, overall prompt text, and material reference conditions are combined into decoration generation constraints. These constraints, along with the base spatial image, serve as input to the preset image generation model. This enables the model to generate candidate decoration renderings based on the base spatial image, according to the area defined by the region mask, the content defined by the prompt text, and the reference content defined by the material reference conditions.
[0068] Through the above processing, fixed area requirements, variable area requirements, and overall collaborative requirements are integrated into input conditions that correspond to each other, such as area masks, text prompts, and material references. This allows multimodal decoration requirements to participate in image generation with a clear area of influence, reducing confusion between the areas of influence of different requirements and improving the stability of the correspondence between candidate decoration renderings and user decoration requirements.
[0069] Step S105: Based on the basic spatial image and decoration generation constraints, call the preset image generation model to generate candidate decoration effect images, and perform consistency screening on the candidate decoration effect images to obtain the final decoration effect image.
[0070] In step S105, based on the basic spatial image obtained in step S101 and the decoration generation constraints obtained in step S104, a preset image generation model is called to generate multiple candidate decoration effect images. These candidate images are then subjected to consistency screening to obtain the final decoration effect image. The preset image generation model in this step is not a general image generation model that only supports basic images and single text prompts. Instead, it is obtained through conditional adaptation training or fine-tuning using decoration image samples, enabling it to read fixed-area masks, content editing masks, fixed-area prompt text, variable-area prompt text, overall prompt text, and material reference conditions.
[0071] Specifically, the preset image generation model is obtained through training or fine-tuning of training samples. Each training sample includes a sample base space image, sample multimodal decoration requirement information, and sample decoration effect image. For the sample base space image, it is first converted to the same input size and color format as the base space image, and then Lab color space conversion, SLIC superpixel segmentation, and saliency analysis are performed. According to the preset saliency threshold, the superpixel units in the sample base space image are divided into sample space fixed regions and sample decoration variable regions. For the sample multimodal decoration requirement information, the selection requirement records and text requirement records are merged into sample text form requirements, and semantic segmentation and region bias judgment are performed on the sample text form requirements to obtain the sample fixed region requirements, sample variable region requirements, or sample overall collaborative requirements corresponding to the sample text requirement units. At the same time, the sample material image records are used as sample image requirement units, and the sample fixed region requirements, sample variable region requirements, or sample overall collaborative requirements corresponding to the sample image requirement units are determined according to the image semantic similarity between the sample image requirement units and the image prototype vectors.
[0072] After obtaining the sample fixed-region requirements, sample variable-region requirements, and sample overall coordination requirements, a sample fixed-region mask is generated based on the sample spatial fixed-region, and a sample content editing mask is generated based on the sample decoration variable-region. Sample fixed-region prompt text is generated based on the text requirement units in the sample fixed-region requirements, sample variable-region prompt text is generated based on the text requirement units in the sample variable-region requirements, and sample overall prompt text is generated based on the text requirement units in the sample overall coordination requirements. Sample material reference conditions are also generated based on the image requirement units in the sample fixed-region requirements, sample variable-region requirements, and sample overall coordination requirements. Thus, each training sample establishes a correspondence between the sample basic spatial image, sample decoration generation constraints, and sample decoration effect image. The sample decoration generation constraints include the sample fixed-region mask, sample content editing mask, sample fixed-region prompt text, sample variable-region prompt text, sample overall prompt text, and sample material reference conditions.
[0073] During training or fine-tuning, the initial image editing and generation model is trained or fine-tuned by using the sample base spatial image and sample decoration generation constraints as model inputs, and the sample decoration effect image as the target output. During training, the initial image editing and generation model learns the following correspondences: the region defined by the sample fixed region mask responds to the sample fixed region prompt text while maintaining the main spatial structure; the region defined by the sample content editing mask responds to the sample variable region prompt text; the overall sample prompt text affects the overall visual effect of the sample decoration effect image; and the material image data in the sample material reference conditions are referenced within the range defined by the corresponding material reference mask. Through the above training or fine-tuning, a preset image generation model capable of generating candidate decoration effect images based on the base spatial image and decoration generation constraints is obtained.
[0074] When generating candidate interior design renderings, the base spatial image and the design generation constraints are input into a preset image generation model. Specifically, a fixed-area mask is associated with fixed-area prompt text, a content editing mask is associated with variable-area prompt text, the overall prompt text serves as the input for the entire image, and the material reference conditions serve as a reference input with a material reference mask. The preset image generation model generates multiple candidate interior design renderings based on the above inputs. These multiple candidate renderings can be obtained by setting different random seeds, different sampling times, or different generation batches.
[0075] Because the image generation process is random, different candidate decoration renderings may satisfy the decoration generation constraints to varying degrees. Therefore, after obtaining multiple candidate decoration renderings, a consistency screening is performed on them. The consistency screening does not simply compare the overall differences between the candidate decoration renderings and the base spatial image, but rather determines whether the candidate decoration renderings simultaneously meet the structural preservation requirements of the fixed spatial area and all the requirements in the decoration generation constraints.
[0076] For the For each candidate interior design rendering, first calculate the structural integrity value for a fixed area: ; in, Indicates the first The fixed area structure retention value of each candidate decoration rendering; Represents the basic spatial image. Indicates the first One candidate interior design rendering; This represents the set of valid pixel positions in a fixed region mask, where valid pixel positions are the pixel positions in the fixed region mask that have a value of 1. Represents the base space image in coordinates The normalized gradient value at that point, Indicates the first The candidate interior design renderings are located at the coordinates The normalized gradient value at that point; Indicates the number of valid pixel positions in a fixed region mask; This is a preset minimum positive number. The larger the value of the fixed region structure preservation, the stronger the... Within the corresponding area of the fixed-area mask, each candidate interior design rendering maintains the structural continuity of the base spatial image. Normalized gradient values are used for calculation because the fixed area can undergo color or style changes according to decoration needs, but its spatial structure should not be disrupted.
[0077] Then, calculate the first... The constraint response values of each candidate decoration rendering are determined. The fixed area prompt text, variable area prompt text, overall prompt text, and material reference conditions from step S104 are used as constraint terms, and a corresponding action mask is determined for each constraint term. Specifically, the fixed area prompt text corresponds to the fixed area mask, the variable area prompt text corresponds to the content editing mask, the overall prompt text corresponds to the whole image mask, and the material reference conditions correspond to the material reference mask. All pixel positions in the base spatial image within the whole image mask are considered valid pixel positions.
[0078] For any constraint term From the first Extract the image content within the mask corresponding to the constraint from each candidate decoration rendering, and calculate the semantic matching value between the image content and the constraint content. If the constraint... If the prompt text is a fixed-area prompt text, a variable-area prompt text, or a global prompt text, then the prompt text is input into a preset text semantic encoder to obtain the conditional semantic vector corresponding to the constraint term. If constraint terms If the material reference condition is used, the material image data in the material reference condition is input into a preset image semantic encoder to obtain the conditional semantic vector corresponding to the constraint term. In other words, conditional semantic vectors Used to represent constraint terms The content features; when the constraint is text, the conditional semantic vector comes from the text content; when the constraint is a source image, the conditional semantic vector comes from the source image data. For the first... The candidate decoration renderings are masked. A defined image region is input into a preset image semantic encoder to obtain a region image semantic vector. .
[0079] No. The constraint response value of each candidate decoration rendering is calculated according to the following formula: ; in, Indicates the first The constraint response values of each candidate interior design rendering; This represents the set of constraint terms in the decoration generation constraint conditions. The set of constraint terms includes fixed area prompt text, variable area prompt text, overall prompt text, and material reference conditions with existing material image data. Represents any constraint term in the set of constraint terms; Represents constraint terms Corresponding function mask; Indicates from the first The candidate decoration renderings are masked. The semantic vector of the region image obtained by extracting the defined region; Represents constraint terms The corresponding conditional semantic vector; This represents cosine similarity. By taking the minimum semantic matching value among all constraint terms as the constraint response value, we can avoid situations where candidate decoration renderings only meet some requirements while omitting others.
[0080] Based on the fixed region structure preservation value and constraint response value, calculate the first... Overall consistency value of candidate interior design renderings: ; in, Indicates the first The overall consistency value of candidate interior design renderings. This formula is used to simultaneously measure the response to both the preservation of the basic spatial structure and the constraints of the interior design; when a candidate interior design rendering responds to some of the interior design requirements but disrupts the structure of a fixed area, the overall consistency is affected. This reduces and lowers the overall consistency value; when candidate interior design renderings maintain a fixed area structure but do not consistently meet the conditions for fixed area prompt text, variable area prompt text, overall prompt text, or material reference. This reduces and lowers the overall consistency score. Therefore, a better performance in one area will not mask a significant deficiency in another.
[0081] Finally, the candidate decoration rendering with the highest overall consistency value is determined as the final decoration rendering. If multiple candidate decoration renderings have the same highest overall consistency value, the candidate decoration rendering with the higher fixed area structure preservation value is selected first. If multiple candidate decoration renderings have the same overall consistency value and fixed area structure preservation value, the candidate decoration rendering with the smallest generation sequence number is selected as the final decoration rendering, according to the generation order of the candidate decoration renderings. Through the above processing, the preset image generation model can generate candidate decoration renderings based on decoration generation constraints. The consistency screening process can evaluate the candidate decoration renderings based on fixed area structure preservation and constraint response, and complete the screening through a deterministic generation order when the evaluation results are the same, thereby improving the correspondence stability between the final decoration rendering and the basic spatial image and multimodal decoration requirement information.
[0082] In summary, combining Figure 3First, basic spatial images and multimodal decoration requirement information are acquired. These are then preprocessed to ensure that user-uploaded or selected images, text, options, and source images are formatted uniformly. Next, the preprocessed basic spatial images are segmented into fixed spatial areas and variable decoration areas, distinguishing and recording background content and local object content. The preprocessed multimodal decoration requirement information is analyzed, treating selection and text requirement records as text-based requirements and source image records as image-based requirements, determining their corresponding fixed-area requirements, variable-area requirements, and overall collaborative requirements. Further, these requirements are fused to generate decoration generation constraints, including area masking conditions, text prompting conditions, and source image reference conditions. Based on the basic spatial images and decoration generation constraints, a pre-defined image generation model is used to generate multiple candidate decoration renderings. Consistency screening is performed on the candidate renderings based on the preservation of the fixed-area structure and the response of each constraint in the decoration generation constraints, resulting in the final decoration rendering.
[0083] Through the above scheme, the basic spatial image is no longer directly used in generation as a single whole image and unified prompt text. Instead, it is first segmented into fixed spatial regions and variable decoration regions, allowing the image generation model to distinguish the scope of influence of different image regions based on region masks. Multimodal decoration requirement information is no longer simply pieced together into a single generated prompt, but is analyzed into fixed region requirements corresponding to the fixed spatial region, variable region requirements corresponding to the variable decoration region, and overall collaborative requirements affecting both types of regions simultaneously. Therefore, the user's requirements expressed through text, options, and source images can be transformed into fixed region prompt text, variable region prompt text, overall prompt text, and source reference conditions, respectively. These are then bound to the corresponding region masks and input into the image generation model, enabling the decoration requirements to function within the corresponding regions of the basic spatial image. Furthermore, the decoration generation constraints are used to constrain the generation of candidate decoration renderings during the generation stage, and participate in the consistency evaluation of candidate decoration renderings as a set of constraints during the screening stage. The final decoration renderings are screened from two aspects: the preservation of fixed area structure and the response of constraint items. This reduces the situation where the original spatial structure is significantly damaged, the demand content is not stably responded to, the scope of material reference is inaccurate, and the overall picture is inconsistent with the user's needs. This allows the final decoration renderings to more stably reflect the actual decoration needs expressed by the user through various input methods.
[0084] Figure 4 This is a structural block diagram of a decoration rendering generation system based on multimodal artificial intelligence provided in one embodiment of this application. The system includes at least the following modules: The data acquisition module is used to acquire basic spatial images and multimodal decoration demand information, and to preprocess the basic spatial images and multimodal decoration demand information respectively; The scene segmentation module is used to segment the preprocessed base space image into indoor scenes, obtaining the fixed spatial region and the variable decoration region in the base space image. The requirements analysis module is used to analyze the pre-processed multimodal decoration requirements information to determine the fixed area requirements of the corresponding fixed spatial area, the variable area requirements of the corresponding variable decoration area, and the overall collaborative requirements of both the fixed spatial area and the variable decoration area. The constraint generation module is used to integrate fixed area requirements, variable area requirements, and overall collaborative requirements to generate decoration generation constraints. The image generation module is used to generate candidate decoration effect images based on the basic spatial image and decoration generation constraints by calling the preset image generation model, and to perform consistency screening on the candidate decoration effect images to obtain the final decoration effect image.
[0085] For relevant details, please refer to the above method implementation examples.
[0086] Figure 5 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 501 and a memory 502.
[0087] Processor 501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. Processor 501 may be implemented using at least one hardware form of DSP, FPGA, PLA, or a combination of CPU, GPU, NPU, AI accelerator, or the above processors. The CPU can be used to perform general computing tasks such as data scheduling, interface interaction, file reading, parameter configuration, and flow control; the GPU can be used to perform parallel computing tasks such as basic spatial image preprocessing, indoor scene segmentation, image calculation in the candidate decoration rendering generation process, and image display; the NPU or AI accelerator can be used to perform semantic segmentation for text-based requirements, region bias judgment, image semantic similarity calculation for image-based requirements, image generation model inference, and semantic matching calculation in consistency screening. In some embodiments, processor 501 may also include a main processor and a coprocessor. The main processor is used to execute the main processing flow in the decoration rendering generation method based on multimodal artificial intelligence, and the coprocessor is used to perform image preprocessing, material image processing, model inference acceleration, or data caching processing in a low-power state.
[0088] The memory 502 may include one or more computer-readable storage media, which may be non-transitory. The memory 502 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash memory devices, or solid-state storage devices. In some embodiments, the memory 502 may be used to store a base spatial image, multimodal decoration requirement information, preprocessed image data, fixed spatial areas, variable decoration areas, text requirement units, image requirement units, fixed area requirements, variable area requirements, overall coordination requirements, decoration generation constraints, candidate decoration renderings, and the final decoration rendering. The memory 502 may also be used to store a preset image generation model, parameter configuration file, image calibration sample set, text calibration sample set, material library data, and program instructions for executing the method of this application. In some embodiments, the non-transitory computer-readable storage medium in the memory 502 stores at least one program instruction, which, when executed by the processor 501, implements the decoration rendering generation method based on multimodal artificial intelligence provided in the embodiments of this application.
[0089] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the multimodal artificial intelligence-based method for generating interior decoration renderings described in the above method embodiments.
[0090] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the multimodal artificial intelligence-based decoration rendering generation method of the above method embodiments.
[0091] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0092] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for generating interior design renderings based on multimodal artificial intelligence, characterized in that, The method includes: Acquire basic spatial images and multimodal decoration requirement information, and preprocess the basic spatial images and multimodal decoration requirement information respectively; The preprocessed base space image is segmented into an indoor scene to obtain a fixed spatial region and a variable decoration region in the base space image. The preprocessed multimodal decoration demand information is analyzed to determine the fixed area demand corresponding to the fixed spatial area, the variable area demand corresponding to the variable decoration area, and the overall coordination demand corresponding to both the fixed spatial area and the variable decoration area. The fixed area requirements, the variable area requirements, and the overall collaborative requirements are integrated to generate decoration generation constraints. Based on the basic spatial image and the decoration generation constraints, a preset image generation model is invoked to generate candidate decoration effect images, and the candidate decoration effect images are subjected to consistency screening to obtain the final decoration effect image.
2. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 1, characterized in that, The process of acquiring basic spatial images and multimodal decoration requirement information, and preprocessing the basic spatial images and multimodal decoration requirement information respectively, includes: After receiving the base spatial image, the file data of the base spatial image is read first. It is determined whether the file data can be parsed into image data normally. If it can be parsed into image data normally, the image width, image height, image orientation information and color channel information of the base spatial image are obtained. After the file data reading is completed, the image orientation is corrected, the image size is unified and the color format is converted into the base spatial image to obtain the preprocessed base spatial image. The multimodal decoration requirement information refers to the requirement data input or selected by the user around the process of generating decoration renderings; after receiving the multimodal decoration requirement information, the multimodal decoration requirement information is divided into selection-type requirement information, text-type requirement information, and image-type requirement information according to the data presentation format; For selection-based requirements, selection requirement records are generated according to preset fields; for text-based requirements, the text content entered by the user is processed to obtain text requirement records; for image-based requirements, the source images are read, their size is standardized, and their color format is converted to obtain source image records; selection requirement records, text requirement records, and source image records together constitute the preprocessed multimodal decoration requirement information.
3. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 1, characterized in that, The step of performing indoor scene segmentation on the preprocessed base space image to obtain fixed spatial regions and variable decoration regions in the base space image includes: The preprocessed base space image is converted to the Lab color space, and the SLIC superpixel segmentation algorithm is used to perform superpixel segmentation on the converted base space image to obtain a superpixel set. A saliency analysis is performed on the base spatial image to obtain a saliency map corresponding to the base spatial image. The saliency map is used to represent the visual prominence of each pixel in the base spatial image relative to the whole image. Based on the superpixel set and the saliency map, the regional saliency corresponding to each superpixel unit is calculated. The regional saliency is calculated by the average value of the values of each pixel in the superpixel unit in the saliency map. If the regional saliency of a superpixel unit is greater than or equal to a preset saliency threshold, then the superpixel unit is recorded as an object superpixel; if the regional saliency of a superpixel unit is less than the preset saliency threshold, then the superpixel unit is recorded as a background superpixel. The adjacent background superpixels are merged to obtain the fixed spatial region, and the adjacent object superpixels are merged to obtain the variable decoration region. The fixed spatial region represents a continuous image region in the base spatial image that serves as the interior background, and the variable decoration region represents a local object region in the base spatial image that is presented independently relative to the interior background.
4. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 2, characterized in that, The analysis of the preprocessed multimodal decoration demand information to determine the fixed area demand corresponding to the fixed spatial area, the variable area demand corresponding to the variable decoration area, and the overall collaborative demand corresponding to both the fixed spatial area and the variable decoration area includes: The selection requirement record and the text requirement record are merged into a text form requirement, and dependency parsing and semantic role labeling are used to semantically decompose the text form requirement to obtain at least one text requirement unit. Based on the mask corresponding to the fixed spatial area, generate a description text for the fixed spatial area; and based on the mask corresponding to the variable decoration area, generate a description text for the variable decoration area. Calculate the first one respectively Semantic similarity between each text requirement unit and the text describing the fixed spatial region and the Semantic similarity between each text requirement unit and the text describing the variable decoration area and based on and Calculate the first Regional bias value of each text demand unit and effective correlation value ,in: ; ; like If the text value is less than the preset effective text threshold, then discard the first one. Each text requirement unit; if If the value is greater than or equal to the preset text validity threshold, then according to... The comparison result with the text region bias threshold will be the first Each text requirement unit is determined as a fixed area requirement, a variable area requirement, or an overall collaborative requirement; The source images are recorded as image requirement units, and the fixed region image prototype vector, variable region image prototype vector, and overall collaborative image prototype vector are determined based on a pre-constructed image calibration sample set. The first The source image corresponding to the first image requirement unit is converted into an image semantic vector, and the first image is calculated according to the following formula. The image demand unit and the first Image semantic similarity between image prototype vectors: in, This indicates the image category corresponding to the requirement of a fixed area. This indicates the image category corresponding to the variable region requirement. This indicates the image category corresponding to the overall collaborative requirements. Indicates the first The image semantic vector corresponding to each image requirement unit Indicates the first Image prototype vector; According to the The maximum image similarity among the three types of semantic similarity corresponding to each image demand unit, and the difference between it and the second largest image similarity, will be used to determine the first... Each image demand unit is determined as a fixed region demand, a variable region demand, an overall collaborative demand, or an unoriented image demand, and the unoriented image demand is discarded.
5. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 4, characterized in that, The process of integrating the fixed area requirements, the variable area requirements, and the overall collaborative requirements to generate decoration generation constraints includes: A fixed region mask is generated based on the fixed spatial region, and a content editing mask is generated based on the variable decoration region. Both the fixed region mask and the content editing mask are in the same pixel coordinate system as the base spatial image. A fixed-area prompt text is generated based on the text requirement units in the fixed-area requirement; a variable-area prompt text is generated based on the text requirement units in the variable-area requirement; an overall prompt text is generated based on the text requirement units in the overall collaboration requirement; and the text requirement units in the overall collaboration requirement are associated with the fixed-area prompt text and the variable-area prompt text. Material reference conditions are generated based on image requirement units in the fixed area requirement, the variable area requirement, and the overall collaborative requirement. Specifically, image requirement units belonging to the fixed area requirement are associated with the fixed area mask, image requirement units belonging to the variable area requirement are associated with the content editing mask, and image requirement units belonging to the overall collaborative requirement are associated with the whole image mask. The fixed area mask, the content editing mask, the fixed area prompt text, the variable area prompt text, the overall prompt text, and the material reference conditions are combined to form the decoration generation constraints.
6. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 5, characterized in that, The preset image generation model is obtained in the following way: Acquire training samples, which include sample base spatial images, sample multimodal decoration demand information, and sample decoration effect images; The sample base space image is segmented into an indoor scene to obtain a fixed sample space region and a variable sample decoration region; the sample multimodal decoration requirement information is analyzed to obtain the sample fixed region requirement, the sample variable region requirement, and the sample overall coordination requirement; and the sample fixed region requirement, the sample variable region requirement, and the sample overall coordination requirement are merged into the sample decoration generation constraints. The sample base space image and the sample decoration generation constraints are used as model inputs, and the sample decoration effect image is used as the target output. The initial image editing and generation model is trained or fine-tuned to obtain the preset image generation model. The constraints for generating the sample decoration include a fixed area mask, a content editing mask, a fixed area prompt text, a variable area prompt text, an overall prompt text, and reference conditions for sample materials.
7. The method for generating interior design renderings based on multimodal artificial intelligence according to claim 6, characterized in that, The process of obtaining the final decoration rendering by performing consistency screening on the candidate decoration renderings includes: For the Based on the candidate decoration renderings, the structural retention value of the fixed area is calculated using the fixed area mask. The constraint response value is calculated based on the set of constraint terms in the decoration generation constraint conditions. And based on the fixed region structure retention value and the constraint response value Calculate the overall consistency value ; Wherein, the fixed region structure retention value Based on the aforementioned basic spatial image and the first The normalized gradient difference of each candidate decoration rendering image within the effective pixel position of the fixed area mask is determined; The constraint response value The minimum value among the semantic matching values between the candidate image content and the corresponding constraint content within the mask corresponding to each constraint term is determined. The overall consistency value Determined according to the following formula: ; The candidate design rendering with the highest overall consistency score will be selected as the final design rendering.
8. A system for generating interior design renderings based on multimodal artificial intelligence, characterized in that, include: The data acquisition module is used to acquire basic spatial images and multimodal decoration demand information, and to preprocess the basic spatial images and multimodal decoration demand information respectively; The scene segmentation module is used to segment the preprocessed base space image into indoor scenes to obtain fixed spatial areas and variable decoration areas in the base space image. The demand analysis module is used to analyze the preprocessed multimodal decoration demand information to determine the fixed area demand corresponding to the fixed space area, the variable area demand corresponding to the variable decoration area, and the overall collaborative demand corresponding to both the fixed space area and the variable decoration area. The constraint generation module is used to integrate the fixed area requirements, the variable area requirements, and the overall collaborative requirements to generate decoration generation constraints. The image generation module is used to generate candidate decoration effect images by calling a preset image generation model based on the basic spatial image and the decoration generation constraints, and to perform consistency screening on the candidate decoration effect images to obtain the final decoration effect image.
9. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement a method for generating interior decoration renderings based on multimodal artificial intelligence as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a program, which, when executed by a processor, is used to implement a method for generating interior decoration renderings based on multimodal artificial intelligence as described in any one of claims 1 to 7.