Effect picture generation method, electronic equipment, readable storage medium and program product
By obtaining furniture images and room type text, the rendering of furniture in a specific room is automatically generated, which solves the problem of low efficiency in manual production of renderings in the prior art, and achieves efficient and accurate rendering generation.
Patent Information
- Application Number
- CN202411999091.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, designers need to manually create renderings of furniture in the room, which is inefficient, and consumers need to go through repeated communication, modification and waiting, which is a high time cost.
By obtaining furniture images and text that characterizes room types, determine the category text corresponding to the furniture images, and generate target renderings based on this information to achieve the effect of placing furniture in a specific room type.
The rendering generation process is automated, saving designers time and labor costs, and improving the efficiency and accuracy of rendering generation.
Smart Images

Figure CN119941891A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of computers, and more particularly to a method for generating an effect diagram, an electronic device, a readable storage medium, and a computer program product. Background Art
[0002] In the modern interior design and decoration industry, the production of renderings plays a vital role. In related technologies, designers manually produce renderings of furniture placed in the room, which is inefficient. If consumers want to get a satisfactory rendering, they often need to go through repeated communication, modification and waiting, which is time-consuming and costly. Summary of the invention
[0003] In order to solve at least one of the above technical problems, the present disclosure provides a method for generating an effect diagram, an electronic device, a readable storage medium and a computer program product.
[0004] According to one aspect of the present disclosure, a method for generating an effect diagram is provided, comprising: Obtaining a furniture image and a first text representing a room type; Determining a second text corresponding to the furniture image, wherein the second text can represent a category of furniture in the furniture image; A target rendering is generated based at least on the furniture image, the first text and the second text corresponding to the furniture image, wherein the target rendering can at least represent the placement effect of the furniture in the furniture image in a room of the same room type as the first text.
[0005] According to at least one embodiment of the present disclosure, a method for generating a rendering, generating a target rendering based on at least the furniture image, the first text, and the second text corresponding to the furniture image, includes: Extracting image features of the furniture image to obtain image features of the furniture image; Performing text feature extraction on the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image; fusing the image feature of the furniture image and the corresponding furniture category feature to obtain a first fused feature; and The target effect graph is generated based on at least the first fusion feature and the text feature.
[0006] According to the rendering generation method of at least one embodiment of the present disclosure, the number of the furniture images is at least one, and the image features of the furniture images and the corresponding furniture category features are fused together to obtain a first fused feature, including: Performing linear mapping on image features of each furniture image in at least one furniture image to obtain mapped image features of each furniture image; Fusing the mapped image features of each furniture image with the corresponding furniture category features to obtain a second fused feature corresponding to each furniture image; Fusing the second fusion features corresponding to the furniture images together to obtain a third fusion feature; and The third fused feature is resampled to obtain the first fused feature, wherein the first fused feature and the text feature have the same spatial dimension.
[0007] According to at least one embodiment of the present disclosure, the method for generating an effect diagram generates the target effect diagram at least according to the first fusion feature and the text feature, including: The first fusion feature and the text feature are input into a first image generation network of a pre-trained first model, and image generation processing is performed on the first fusion feature and the text feature through the first image generation network to obtain the target effect image.
[0008] According to the rendering generation method of at least one embodiment of the present disclosure, the first model also includes a first image feature extraction network, a first text feature extraction network, a first linear mapping network and a first resampling network; Wherein, the first image feature extraction network is used to extract image features from the furniture image to obtain image features of the furniture image; The first text feature extraction network is used to extract text features from the first text and the second text corresponding to the furniture image, respectively, to obtain text features of the first text and furniture category features corresponding to the furniture image; The first linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain mapped image features of the furniture image; The first resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature.
[0009] According to the method for generating a rendering of at least one embodiment of the present disclosure, the first model is trained by the following steps: Acquire a first sample data set, the first sample data set comprising a plurality of first data pairs, each of the first data pairs comprising a sample rendering and a first sample text corresponding to the sample rendering representing a room type in the sample rendering, at least one sample furniture image and a second sample text representing a furniture category corresponding to the sample furniture image, the sample rendering comprising at least one sample furniture, and the sample furniture image and the second sample text corresponding one to one; Freezing the first image feature extraction network and the first text feature extraction network; and The first linear mapping network, the first resampling network and the first image generation network are trained based on the first sample data set to obtain the trained first model.
[0010] According to at least one embodiment of the present disclosure, a method for generating an effect diagram, obtaining a first sample data set, includes: Acquire a plurality of sample renderings and a first sample text representing the room types in the sample renderings; Performing target detection on each of the plurality of sample renderings respectively to obtain a position of at least one sample furniture in each of the sample renderings and a second sample text representing a category of at least one sample furniture; Based on the position, segmenting at least one sample furniture image from each sample rendering; and The first data pairs are generated based on the sample renderings and the first sample texts corresponding to the sample renderings, at least one sample furniture image and the second sample text to obtain the first sample data set.
[0011] According to at least one embodiment of the present disclosure, the method for generating a rendering further includes: Get the target room image; Determining room depth information corresponding to the target room image; Performing feature extraction on the room depth information to obtain room depth features; Generating the target effect graph at least according to the first fusion feature and the text feature, further comprising: The target rendering is generated according to the first fusion feature, the text feature and the room depth feature, wherein the target rendering can represent the placement effect of the furniture in the furniture image in the target room whose room type is the same as that of the first text.
[0012] According to at least one embodiment of the present disclosure, the method for generating an effect map generates the target effect map according to the first fusion feature, the text feature and the room depth feature, including: The first fusion feature, the text feature and the room depth feature are input into a second image generation network of a pre-trained second model, and the first fusion feature, the text feature and the room depth feature are processed by the second image generation network to generate an image, so as to obtain the target effect image.
[0013] According to the effect diagram generation method of at least one embodiment of the present disclosure, the second model also includes a second image feature extraction network, a second text feature extraction network, a second linear mapping network, a second resampling network and a control network; Wherein, the second image feature extraction network is used to extract image features from the furniture image to obtain image features of the furniture image; The second text feature extraction network is used to extract text features from the first text and the second text corresponding to the furniture image, respectively, to obtain text features of the first text and furniture category features corresponding to the furniture image; The second linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain mapped image features of the furniture image; The second resampling network is used to resample the second fused feature or the third fused feature to obtain the first fused feature; The control network is used to extract features from the room depth information to obtain the room depth features.
[0014] According to the method for generating a rendering of at least one embodiment of the present disclosure, the second model is trained by the following steps: Acquire a second sample data set, the second sample data set includes a plurality of second data pairs, each of the second data pairs includes a sample rendering and room depth information corresponding to the sample rendering, a first sample text representing the room type in the sample rendering, at least one sample furniture image and a second sample text representing the furniture category corresponding to the sample furniture image, the sample rendering includes at least one sample furniture, and the sample furniture image corresponds to the second sample text one-to-one; Freezing the second image feature extraction network and the second text feature extraction network; and The second linear mapping network, the second resampling network, the control network and the second image generation network are trained based on the second sample data set to obtain the trained second model.
[0015] According to at least one embodiment of the effect diagram generation method of the present disclosure, obtaining a second sample data set includes: Acquire a plurality of sample renderings and a first sample text representing the room types in the sample renderings; Performing target detection on each of the plurality of sample renderings respectively to obtain a position of at least one sample furniture in each of the sample renderings and a second sample text representing a category of at least one sample furniture; Based on the position, segmenting at least one sample furniture image from each sample rendering; Extracting room depth information from each sample rendering respectively to obtain room depth information corresponding to each sample rendering; and The second data pairs are generated based on the sample renderings and the room depth information corresponding to the sample renderings, the first sample text, at least one sample furniture image and the second sample text to obtain the second sample data set.
[0016] According to another aspect of the present disclosure, an electronic device is provided, comprising: a memory storing a computer program; and a processor executing the computer program stored in the memory, so that the processor executes the effect diagram generation method of any embodiment of the present disclosure.
[0017] According to another aspect of the present disclosure, a readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement the effect diagram generation method of any embodiment of the present disclosure.
[0018] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the method for generating a rendering according to any embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings illustrate exemplary embodiments of the present disclosure and together with the description serve to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are incorporated in and constitute a part of this specification.
[0020] Figure 1 It is a flowchart of a method for generating a rendering according to an embodiment of the present disclosure.
[0021] Figure 2 It is a schematic diagram of a process of generating a target effect diagram according to an embodiment of the present disclosure.
[0022] Figure 3 It is a schematic diagram of a process of obtaining a first fusion feature according to an embodiment of the present disclosure.
[0023] Figure 4 It is a schematic structural diagram of a first model according to an embodiment of the present disclosure.
[0024] Figure 5 It is a schematic diagram of the process of training the first model according to one embodiment of the present disclosure.
[0025] Figure 6 It is a schematic diagram of an application example of the first model according to one embodiment of the present disclosure.
[0026] Figure 7 It is a schematic diagram of a process of obtaining a first sample data set according to an embodiment of the present disclosure.
[0027] Figure 8 The diagram is a visualization example of generating a first data pair according to one embodiment of the present disclosure.
[0028] Fig. 9 It is a schematic structural diagram of a second model according to an embodiment of the present disclosure.
[0029] Fig.10 It is a schematic diagram of the process of training the second model according to one embodiment of the present disclosure.
[0030] Fig.11 It is a schematic diagram of an application example of the second model according to an embodiment of the present disclosure.
[0031] Fig.12 It is a schematic diagram of a process of obtaining a second sample data set according to an embodiment of the present disclosure.
[0032] Fig.13 It is a schematic block diagram of the structure of an effect diagram generating device according to an embodiment of the present invention.
[0033] Fig.14 It is a schematic block diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0034] The present disclosure is further described in detail below in conjunction with the accompanying drawings and examples. It is understood that the specific examples described herein are only used to explain the relevant content, rather than to limit the present disclosure. It should also be noted that, for ease of description, only the parts related to the present disclosure are shown in the accompanying drawings.
[0035] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in the present disclosure can be combined with each other. The technical solution of the present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0036] The furniture usually placed in the living room includes sofas, coffee tables, benches, etc., and the furniture usually placed in the bedroom includes beds, wardrobes, etc. In the related art, if consumers want to know the placement effect of sofas, coffee tables, benches, etc. in the living room before decoration, the designer needs to manually draw the placement effect diagram of sofas, coffee tables, benches, etc. in the living room; if consumers want to know the placement effect of beds, wardrobes, etc. in the bedroom before decoration, the designer needs to manually draw the placement effect diagram of beds, wardrobes, etc. in the bedroom. However, manually drawing the effect diagram consumes a lot of time, which leads to low efficiency.
[0037] To this end, the present disclosure proposes a method for generating an effect diagram.
[0038] The rendering generation method disclosed in the present invention can be used for electronic devices to display the renderings of furniture placement in a room to users. In the present disclosure, electronic devices include but are not limited to mobile phones, tablet computers, laptops, personal computers, wearable devices, ATMs, etc.
[0039] For the convenience of description and to make the technical solutions of the specific implementation methods of the present disclosure easier to understand, before describing the effect diagram generation method implemented by the present disclosure, the technical terms involved in the specific implementation methods of the present disclosure are explained as follows: Encoder: A component used to convert input data (such as images, text, etc.) into a compact feature representation or latent space vector.
[0040] Stable Diffusion: is a generative model based on deep learning. It gradually generates high-quality images by iteratively adding subtle structures to noisy data. It is widely used in text-to-image synthesis tasks.
[0041] Denoising U-Net: is a neural network architecture and a key component in the Stable Diffusion model.
[0042] Controlnet: is a neural network that controls a pre-trained image diffusion model (such as StableDiffusion). It allows the input of a conditioning image and then uses the conditioning image to manipulate image generation.
[0043] Cross Attention: Mainly used for the interaction between two different sequences or feature maps, it enhances the expressiveness of the model by capturing the dependencies between different sequences.
[0044] Concat: It refers to the operation of connecting two or more tensors (such as vectors and matrices) along a specified dimension to form a new larger tensor.
[0045] Resampling: It is mainly used to change the spatial dimension of input data, such as the height and width of an image. It usually includes upsampling or downsampling.
[0046] Linear mapping: Performing a linear transformation from input data to output data through matrix multiplication can change the feature dimension of the input data.
[0047] Figure 1FIG. 1 is a schematic diagram showing the overall process of a method M100 for generating a rendering according to an embodiment of the present disclosure. Figure 1 The method shown includes steps S110 to S140. The method can be executed by electronic devices such as mobile phones and tablet computers.
[0048] Specifically, Figure 1 The methods shown include: S110, acquiring a furniture image and a first text representing a room type; The furniture image can be understood as a two-dimensional picture of a piece of furniture. The first text can characterize the room type, and the first text includes but is not limited to bedroom, living room, bathroom, kitchen, etc. Exemplarily, the user can be prompted to upload and / or select the furniture image and the first text through the display interface of the electronic device, so as to facilitate the subsequent generation of a room image (target rendering) based on the furniture image and the first text uploaded and / or selected by the user, in which the furniture corresponding to the furniture image is placed and the room type is the same as the first text. The number of furniture images input by the user each time can be one or more, and the first text can be one. A furniture image can include at least one piece of furniture.
[0049] S120, determining a second text corresponding to the furniture image, wherein the second text can represent the category of furniture in the furniture image; The second text includes but is not limited to sofas, coffee tables, benches, beds, wardrobes, etc. Exemplarily, the second text can be actively input by the user; the second text corresponding to the furniture image selected by the user can also be directly determined through the mapping relationship between the furniture image and the second text; the second text can also be obtained by automatically identifying the category of furniture in the furniture image uploaded by the user through a target detection algorithm (such as DINO). It is understandable that the placement of furniture of different categories in the room is usually different, and the placement of furniture of different categories may be related. For example, sofas and coffee tables are usually placed together. Therefore, identifying the category of furniture in the furniture image is conducive to better determining the placement of furniture in the room, thereby facilitating the accurate generation of the target rendering.
[0050] S130. Generate a target rendering based at least on the furniture image, the first text, and the second text corresponding to the furniture image, wherein the target rendering can at least represent the placement effect of the furniture in the furniture image in a room of the same room type as the first text.
[0051] The target rendering at least includes the room and the furniture in the furniture image.
[0052] As a possible implementation, a target rendering may be generated based on a furniture image, a first text, and a second text corresponding to the furniture image. For example, the furniture image, the first text, and the second text corresponding to the furniture image may be input into a pre-trained first model, and the first model may be used to automatically generate a target rendering that can characterize the placement effect of the furniture in the furniture image in a room of the same room type as the first text.
[0053] As another possible implementation, a target rendering may be generated based on the furniture image, the first text, the second text corresponding to the furniture image, and the target room image. For example, the furniture image, the first text, the second text corresponding to the furniture image, and the room depth information corresponding to the target room image may be input into a pre-trained second model, and the second model may be used to automatically generate a target rendering that can characterize the placement effect of the furniture in the furniture image in the target room of the same room type as the first text.
[0054] The rendering generation method of the disclosed embodiment automatically determines the second text corresponding to the furniture image after acquiring the furniture image and the first text representing the room type, and automatically generates a target rendering based on at least the furniture image, the first text and the second text corresponding to the furniture image. Since the target rendering is automatically generated without manual drawing, it is possible to save drawing time, reduce labor costs, and improve the efficiency of rendering generation.
[0055] Regarding step S130, in some embodiments of the present disclosure, it may include the following steps: Figure 2 Steps S131 to S134 are shown.
[0056] S131, extracting image features of the furniture image to obtain image features of the furniture image.
[0057] The image features of furniture images can fully characterize the characteristic information of furniture in furniture images, including but not limited to color, texture details, material, etc. By extracting the image features of furniture images, it is helpful to understand the furniture in the furniture images more accurately in the future, thereby improving the accuracy of the generated target renderings.
[0058] S132, performing text feature extraction on the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image.
[0059] Specifically, the text features of the first text are extracted to obtain the text features of the first text; the text features of the second text corresponding to the furniture image are extracted to obtain the text features of the second text, that is, the furniture category features corresponding to the furniture image. By extracting the text features of the first text and the second text respectively, it is helpful to understand the first text and the second text more accurately in the future, thereby improving the accuracy of the generated target rendering.
[0060] S133: Fuse the image features of the furniture image and the corresponding furniture category features to obtain a first fused feature.
[0061] The image features of the furniture image and the corresponding furniture category features are both features of the furniture in the furniture image. Therefore, the first fused features obtained by fusing the two together can better characterize the furniture in the furniture image, which is conducive to a more accurate understanding of the furniture in the furniture image in the subsequent process and improves the accuracy of the generated target rendering.
[0062] S134. Generate a target effect graph based on at least the first fusion feature and the text feature.
[0063] Exemplarily, the first fusion feature and the text feature may be input into a pre-trained neural network, and the target effect graph may be automatically generated through the neural network.
[0064] The rendering generation method of the above-mentioned embodiment can more accurately capture the visual characteristics of furniture and the semantic information of room type by extracting features from furniture images, the first text and the second text and fusing the extracted features, thereby helping to enhance the neural network's ability to understand input data and generate more realistic target renderings.
[0065] Regarding step S133, in some embodiments of the present disclosure, it may include the following steps: Figure 3 Steps S1331 to S1334 are shown.
[0066] S1331. Perform linear mapping on image features of each furniture image in at least one furniture image to obtain mapped image features of each furniture image.
[0067] Linear mapping can change the feature dimension of image features, so that the mapped image features are more in line with expectations and facilitate the execution of subsequent fusion steps.
[0068] S1332: Fuse the mapped image features of each furniture image with the corresponding furniture category features to obtain second fused features corresponding to each furniture image.
[0069] Exemplarily, the mapped image feature of a furniture image and the furniture category feature corresponding to the furniture image are concatenated together to obtain a second fused feature corresponding to the furniture image.
[0070] S1333. Fuse the second fusion features corresponding to the furniture images together to obtain a third fusion feature.
[0071] For example, when there are multiple furniture images, the second fusion features corresponding to each furniture image are concatenated together to obtain the third fusion feature. Since the second fusion features corresponding to each furniture image are pre-fused with the furniture category features corresponding to each furniture image, in the subsequent rendering generation process, the model can clearly know which category the image features of each furniture image belong to, avoiding the mixing of multiple furniture attributes in the final rendering generation process, and at the same time strengthening the image features.
[0072] It is worth noting that, when there is only one furniture image, since there is only one second fusion feature, the second fusion feature corresponding to the furniture image can be directly used as the third fusion feature.
[0073] S1334. Resample the third fused feature to obtain a first fused feature, wherein the first fused feature and the text feature have the same spatial dimension.
[0074] Since the third fused feature includes both features from the image and features from the text, and the text features of the first text only include features from the text, the spatial dimensions of the two are different. In order to facilitate subsequent use, the third fused feature needs to be aligned with the text features of the first text. The spatial dimension of the third fused feature can be changed by resampling, so as to obtain the first fused feature with the same spatial dimension as the text feature, which is convenient for subsequent use.
[0075] The rendering generation method of the above-mentioned implementation adopts linear mapping and resampling technology to enable the features from different sources to be effectively fused in the same dimension, thereby improving the consistency of feature representation and the stability and efficiency in the subsequent rendering generation process. At the same time, when there are multiple furniture images, the relevant features of multiple furniture images can be fused together as a whole, which is conducive to the complete response to each furniture image in the subsequent rendering generation process, avoiding the omission of furniture in the generated target rendering, and improving the accuracy and completeness of the target rendering.
[0076] Regarding step S134, as a possible implementation method, it can be specifically: inputting the first fusion feature and the text feature into the first image generation network of the pre-trained first model, performing image generation processing on the first fusion feature and the text feature through the first image generation network, and obtaining the target effect diagram.
[0077] The first image generation network of the pre-trained first model may contain a large amount of prior knowledge, which plays an important role in improving the realism and detail performance of the target rendering. Exemplarily, the first image generation network may be a Denoising U-Net in a StableDiffusion model, and the Denoising U-Net may include a first cross-attention module and a second cross-attention module, and the first fusion feature and the automatically generated noise feature are input into the first cross-attention module, and the text feature and the automatically generated noise feature are input into the second cross-attention module, and then under the action of the first cross-attention module and the second cross-attention module, a target rendering with better quality is generated.
[0078] The effect diagram generation method of the above-mentioned implementation uses a fully trained first model to generate a target effect diagram, which can improve the quality and accuracy of the generated result.
[0079] Please combine Figure 4 In some embodiments of the present disclosure, the first model further includes a first image feature extraction network (image encoder), a first text feature extraction network (text encoder), a first linear mapping network (Linear) and a first resampling network (Resample); wherein the first image feature extraction network is used to extract image features of the furniture image to obtain image features of the furniture image, that is, the first image feature extraction network can implement the above step S131; the first text feature extraction network is used to extract text features of the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image, that is, the first text feature extraction network can implement the above step S132; the first linear mapping network is used to linearly map the image features of the furniture image to obtain mapped image features of the furniture image, that is, the first linear mapping network can implement the above step S1331 or step S1331'; the first resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature, that is, the first resampling network can implement the above step S1333 or step S1334'.
[0080] In the effect diagram generation method of the above implementation mode, the first model is composed of multiple networks, and the modular design facilitates the optimization and adjustment of the functions of each part, and also provides convenience for the expansion and improvement of the first model. For example, the first model can be used as a plug-in in combination with other models (such as controlnet).
[0081] Please combine Figure 5 In some embodiments of the present disclosure, the first model can be trained by following steps S210 to S230.
[0082] S210. Obtain a first sample data set, where the first sample data set includes a plurality of first data pairs, each of which includes a sample rendering and a first sample text corresponding to the sample rendering that represents the room type in the sample rendering, at least one sample furniture image and a second sample text that represents the furniture category corresponding to the sample furniture image, the sample rendering includes at least one sample furniture, and the sample furniture image corresponds to the second sample text one-to-one.
[0083] Exemplarily, some sample renderings in the first sample data set include one sample furniture, and some sample renderings include multiple sample furniture. In this way, the first model trained with the first sample data set can generate a target rendering including one furniture image when receiving one furniture image, and can generate a target rendering including multiple furniture images when receiving multiple furniture images. This is more flexible and can meet different rendering generation requirements.
[0084] S220: Freeze the first image feature extraction network and the first text feature extraction network.
[0085] The first image feature extraction network and the first text feature extraction network may be pre-trained networks, and thus the two networks may be frozen during the training process, thereby increasing the training speed and saving training time.
[0086] S230: Training a first linear mapping network, a first resampling network, and a first image generation network based on a first sample data set to obtain a trained first model.
[0087] Exemplarily, a first sample text representing the room type in the sample rendering, at least one sample furniture image, and a second sample text representing the furniture category corresponding to the sample furniture image can be input into a preset first model to obtain the rendering generated by the first model, and the loss value between the rendering and the sample rendering is calculated. According to the loss value, the network parameters of the first linear mapping network, the first resampling network, and the first image generation network of the first model are adjusted to obtain a trained first model.
[0088] In one example, the size of the trained first model is 190M, and it has the ability to automatically generate a target rendering when receiving one or more furniture images, a first text representing a room type, and a second text corresponding to the furniture image. Figure 6 As shown in the figure, the sofa image, the "sofa" text, the coffee table image, the "coffee table" text and the "livingroom" text are input into the trained first model. The first model can automatically generate the renderings of the corresponding sofas and coffee tables in the living room, and on the basis of ensuring the beauty of the generated renderings, the color, texture, material and other details of the furniture image are better transferred.
[0089] The rendering generation method of the above embodiment freezes the first image feature extraction network and the first text feature extraction network of the first model, and trains other networks of the first model, thereby saving the training time of the first model and facilitating the rapid acquisition of the trained first model.
[0090] Regarding step S210, in some embodiments of the present disclosure, it may include: Figure 7 Steps S211 to S214 are shown.
[0091] S211. Obtain multiple sample renderings and first sample text representing the room types in the sample renderings.
[0092] S212, performing target detection on each of the multiple sample renderings, to obtain a position of at least one sample furniture in each sample rendering and a second sample text representing a category of at least one sample furniture.
[0093] The target detection algorithm (such as DINO) may be used to perform target detection on each of the multiple sample renderings to obtain the position of at least one sample furniture in each sample rendering and a second sample text representing the category of at least one sample furniture.
[0094] S213: Segment at least one sample furniture image from each sample rendering based on the position.
[0095] After obtaining the position of the sample furniture, a rectangular image containing the sample furniture can be cut out from the sample rendering. The rectangular image includes both the sample furniture and the background. In order to avoid the first model directly learning to directly paste the cut rectangular image into the original sample rendering, the segmentation model can be used to segment the cut rectangular image to remove its background and obtain the sample furniture image, such as Figure 8 shown.
[0096] S214, generating a first data pair based on each sample rendering and the first sample text corresponding to each sample rendering, at least one sample furniture image and the second sample text, to obtain a first sample data set.
[0097] In addition, in order to enable the first model to have the ability to adaptively adjust the angle and placement of the sample furniture images, and to solve the problem of the difficulty in obtaining a large number of sample renderings resulting in a small number of first data pairs, the sample furniture images obtained by segmentation can be subjected to image enhancement processing to obtain processed sample furniture images, and first data pairs can be generated based on each sample rendering and the first sample text corresponding to each sample rendering, at least one processed sample furniture image and the second sample text, respectively, thereby increasing the number of first data pairs in the first sample data set and facilitating the trained first model to have the ability to adaptively adjust the angle and placement of the sample furniture images. Image enhancement techniques include but are not limited to flipping, rotation, etc.
[0098] The effect diagram generation method of the above-mentioned implementation mode can automatically obtain the first sample data set, which can ensure the diversity and representativeness of the training data, thereby improving the learning ability and generalization performance of the first model.
[0099] In some embodiments of the present disclosure, the method may further include: obtaining a target room image; determining room depth information corresponding to the target room image; and extracting features of the room depth information to obtain room depth features. Accordingly, as another possible implementation, step S134 may specifically be: generating a target rendering according to the first fusion feature, the text feature, and the room depth feature, wherein the target rendering can characterize the placement effect of the furniture in the furniture image in the target room whose room type is the same as the first text.
[0100] The target room image may be an image of the target room in the user's own house uploaded by the user, or an image of the target room in other houses uploaded by the user, or an image of the target room selected by the user and provided by an electronic device, without limitation. Room depth information includes but is not limited to a depth map, and the room depth information corresponding to the target room image may be determined by a Leres depth estimation algorithm or a Zoe depth estimation algorithm. By extracting the room depth features of the room depth information, it is helpful to understand the depth of the target room more accurately in the future, thereby improving the realism of the generated target rendering. In this embodiment, the target rendering includes at least the furniture in the target room and the furniture image.
[0101] Exemplarily, the furniture image, room depth information, the first text and the second text corresponding to the furniture image can be input into a pre-trained second model, and the target rendering can be automatically generated by the second model. That is, the room depth information is feature extracted by the second model to obtain the room depth feature, and the target rendering is generated based on the first fusion feature, the text feature and the room depth feature.
[0102] The rendering generation method of the above embodiment realizes the generation of the target rendering according to the furniture image, the first text, the second text corresponding to the furniture image and the target room image, which can meet the user's need to view the placement effect of the furniture image in the target room. At the same time, since there is no need to manually draw the target rendering, but the target rendering is automatically generated, it can save drawing time, reduce labor costs, and improve the rendering generation efficiency. In addition, adding room depth information processing to the target room image can make the generated target rendering better adapt to the three-dimensional spatial structure of the target room, thereby enhancing the authenticity of the target rendering.
[0103] In some embodiments of the present disclosure, a target rendering is generated based on the first fusion feature, the text feature, and the room depth feature. Specifically, the first fusion feature, the text feature, and the room depth feature are input into a second image generation network of a pre-trained second model, and the first fusion feature, the text feature, and the room depth feature are processed by the second image generation network to generate an image to obtain the target rendering.
[0104] The second image generation network of the pre-trained second model may contain a large amount of prior knowledge, which plays an important role in improving the realism and detail performance of the target rendering. Exemplarily, the second image generation network may be a Denoising U-Net in a StableDiffusion model, and the Denoising U-Net may include a first cross-attention module and a second cross-attention module, and the first fusion feature and the automatically generated noise feature are input into the first cross-attention module, and the text feature and the automatically generated noise feature are input into the second cross-attention module, and then under the action of the first cross-attention module and the second cross-attention module, a target rendering with better quality is generated.
[0105] In the rendering generation method of the above embodiment, the second image generation network can further combine the room depth features corresponding to the target room image on the basis of the first image generation network, so that the generated target rendering is closer to the actual situation of the target room, meeting the user's need to view the effect of furniture image placement in the target room.
[0106] Please combine Fig. 9In some embodiments of the present disclosure, the second model further includes a second image feature extraction network (image encoder), a second text feature extraction network (text encoder), a second linear mapping network (Linear), a second resampling network (Resample) and a control network (Controlnet); wherein the second image feature extraction network is used to extract image features of the furniture image to obtain image features of the furniture image, that is, the second image feature extraction network can implement the above step S131; the second text feature extraction network is used to extract text features of the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image, that is, the second text feature extraction network can implement the above step S132; the second linear mapping network is used to linearly map the image features of the furniture image to obtain mapped image features of the furniture image, that is, the second linear mapping network can implement the above step S1331 or step S1331'; the second resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature, that is, the second resampling network can implement the above step S1333 or step S1334'; the control network is used to extract features of the room depth information to obtain room depth features.
[0107] In the rendering generation method of the above embodiment, the second model adds a control network on the basis of the first model, so that the second model can understand the depth situation of the target room, and generate a target rendering that conforms to the depth situation of the target room, thereby improving the quality of the target rendering and meeting the user's needs to view the placement effect of the furniture image in the target room.
[0108] Please combine Fig.10 In some embodiments of the present disclosure, the second model can be trained by following steps S310 to S330.
[0109] S310, obtaining a second sample data set, the second sample data set including a plurality of second data pairs, each second data pair including a sample rendering and room depth information corresponding to the sample rendering, a first sample text representing the room type in the sample rendering, at least one sample furniture image and a second sample text representing the furniture category corresponding to the sample furniture image, the sample rendering includes at least one sample furniture, and the sample furniture image corresponds to the second sample text one-to-one.
[0110] Exemplarily, some of the sample renderings in the second sample data set include one sample furniture, and some of the sample renderings include multiple sample furniture. In this way, the second model trained with the second sample data set can generate a target rendering including one furniture image when receiving one furniture image, and can generate a target rendering including multiple furniture images when receiving multiple furniture images. This is more flexible and can meet different rendering generation requirements.
[0111] S320: Freeze the second image feature extraction network and the second text feature extraction network.
[0112] The second image feature extraction network and the second text feature extraction network may be pre-trained networks, so the two networks may be frozen during the training process, thereby increasing the training speed and saving training time.
[0113] S330: Train the second linear mapping network, the second resampling network, the control network, and the second image generation network based on the second sample data set to obtain a trained second model.
[0114] Exemplarily, the room depth information corresponding to the sample rendering, the first sample text characterizing the room type in the sample rendering, at least one sample furniture image, and the second sample text characterizing the furniture category corresponding to the sample furniture image can be input into a preset second model to obtain the rendering generated by the second model, and the loss value between the rendering and the sample rendering is calculated, and the network parameters of the second linear mapping network, the second resampling network, the control network and the second image generation network of the second model are adjusted according to the loss value to obtain a trained second model.
[0115] like Fig.11 As shown in the figure, the sofa image, the "sofa" text, the "living room" text and the room depth information corresponding to the target room image (obtained by pre-estimating the depth of the target room image) are input into the trained second model. The second model can automatically generate a rendering of the corresponding sofa placed in a living room with a depth that is basically consistent with the target room. On the basis of ensuring the beauty of the generated rendering, the color, texture, material and other details of the furniture image are better transferred.
[0116] The rendering generation method of the above embodiment freezes the second image feature extraction network and the second text feature extraction network of the second model, and trains the other networks of the second model, thereby saving the training time of the second model and facilitating the rapid acquisition of the trained second model. In addition, the second model can understand the depth of the target room, and based on the generation of the target rendering that meets the depth of the target room, the quality of the target rendering is improved, and the user's need to view the placement effect of the furniture image in the target room can be met.
[0117] Regarding step S310, in some embodiments of the present disclosure, it may include: Fig.12 Steps S311 to S315 are shown.
[0118] S311, obtaining a plurality of sample renderings and a first sample text representing the room type in the sample renderings.
[0119] S312, performing target detection on each of the multiple sample renderings, to obtain the position of at least one sample furniture in each sample rendering and a second sample text representing the category of at least one sample furniture.
[0120] S313: Segment at least one sample furniture image from each sample rendering based on the position.
[0121] The relevant contents of step S311 to step S313 can refer to the description of the above-mentioned step S211 to step S213, and for the sake of brevity, they are not repeated here.
[0122] S314: extract room depth information from each sample rendering to obtain room depth information corresponding to each sample rendering.
[0123] The room depth information of each sample rendering can be extracted by using the Leres depth estimation algorithm or the Zoe depth estimation algorithm to obtain the room depth information corresponding to each sample rendering.
[0124] S315, generating second data pairs based on each sample rendering and the room depth information corresponding to each sample rendering, the first sample text, at least one sample furniture image and the second sample text, to obtain a second sample data set.
[0125] In addition, in order to enable the second model to have the ability to adaptively adjust the angle and placement of the sample furniture images, and to solve the problem of the difficulty in obtaining a large number of sample renderings resulting in a small number of second data pairs, the sample furniture images obtained by segmentation can be subjected to image enhancement processing to obtain processed sample furniture images, and second data pairs can be generated based on each sample rendering and the room depth information corresponding to each sample rendering, the first sample text, at least one sample furniture image, and the second sample text, respectively, thereby increasing the number of second data pairs in the second sample data set and facilitating the trained second model to have the ability to adaptively adjust the angle and placement of the sample furniture images. Image enhancement techniques include but are not limited to flipping, rotation, etc.
[0126] The rendering generation method of the above implementation can automatically obtain the second sample data set, which can ensure the diversity and representativeness of the training data, thereby improving the learning ability and generalization performance of the second model. In addition, the second sample data set adds room depth information, which can provide the second model with richer training materials, helping it to better understand and simulate the room environment in the real world, thereby generating more realistic target renderings.
[0127] In one example, a method for generating a rendering includes the following steps: obtaining a furniture image and a first text representing a room type; determining a second text corresponding to the furniture image; inputting the furniture image, the first text, and the second text corresponding to the furniture image into a pre-trained first model, and performing raw image processing on the furniture image, the first text, and the second text corresponding to the furniture image through the first model to obtain a target rendering, which can at least represent the placement effect of the furniture in the furniture image in a room with the same room type as the first text.
[0128] In another example, a rendering generation method includes the following steps: obtaining a furniture image, a target room image, and a first text representing a room type; determining a second text corresponding to the furniture image; determining room depth information corresponding to the target room image; inputting the furniture image, the first text, the second text corresponding to the furniture image, and the room depth information corresponding to the target room image into a pre-trained second model, and performing raw image processing on the furniture image, the first text, the second text corresponding to the furniture image, and the room depth information corresponding to the target room image through the second model to obtain a target rendering, which can represent the placement effect of furniture in the furniture image in a target room with the same room type as the first text.
[0129] Based on any of the above implementations, the present disclosure also provides a rendering generation device.
[0130] Fig.13 It is a schematic block diagram of the structure of an effect diagram generating device according to an embodiment of the present invention.
[0131] like Fig.13 As shown, the effect diagram generating device includes: The acquisition module 110 is used to acquire a furniture image and a first text representing a room type.
[0132] The first determination module 120 is used to determine a second text corresponding to the furniture image, wherein the second text can represent the category of the furniture in the furniture image.
[0133] The generation module 130 is used to generate a target rendering based on at least the furniture image, the first text and the second text corresponding to the furniture image, wherein the target rendering can at least represent the placement effect of the furniture in the furniture image in a room with the same room type as the first text.
[0134] The above modules can be implemented by computer program modules.
[0135] In some embodiments of the present disclosure, the generation module 130 is used to: extract image features of a furniture image to obtain image features of the furniture image; extract text features of a first text and a second text corresponding to the furniture image to obtain text features of the first text and furniture category features corresponding to the furniture image; fuse the image features of the furniture image and the corresponding furniture category features to obtain a first fused feature; and generate a target rendering based on at least the first fused feature and the text feature.
[0136] In some embodiments of the present disclosure, the number of furniture images is at least one, and the generation module 130 is used to: linearly map the image features of each furniture image in the at least one furniture image to obtain the mapped image features of each furniture image; fuse the mapped image features of each furniture image with the corresponding furniture category features to obtain the second fused features corresponding to each furniture image; fuse the second fused features corresponding to each furniture image together to obtain a third fused feature; and resample the third fused feature to obtain the first fused feature, wherein the first fused feature has the same spatial dimension as the text feature.
[0137] In some embodiments of the present disclosure, the generation module 130 is used to: input the first fusion feature and the text feature into a first image generation network of a pre-trained first model, perform image generation processing on the first fusion feature and the text feature through the first image generation network, and obtain a target effect image.
[0138] In some embodiments of the present disclosure, the first model also includes a first image feature extraction network, a first text feature extraction network, a first linear mapping network and a first resampling network; wherein the first image feature extraction network is used to perform image feature extraction on the furniture image to obtain image features of the furniture image; the first text feature extraction network is used to perform text feature extraction on the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image; the first linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain mapped image features of the furniture image; the first resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature.
[0139] In some embodiments of the present disclosure, the first model is obtained by a first model training device, and the first model training device is used to: obtain a first sample data set, the first sample data set includes multiple first data pairs, each first data pair includes a sample rendering and a first sample text corresponding to the sample rendering that represents the room type in the sample rendering, at least one sample furniture image and a second sample text that represents the furniture category corresponding to the sample furniture image, the sample rendering includes at least one sample furniture, and the sample furniture image corresponds to the second sample text one-to-one; freeze the first image feature extraction network and the first text feature extraction network; and train the first linear mapping network, the first resampling network and the first image generation network based on the first sample data set to obtain a trained first model.
[0140] In some embodiments of the present disclosure, the first model training device is used to: obtain multiple sample renderings and a first sample text representing the room type in the sample renderings; perform target detection on each of the multiple sample renderings to obtain the position of at least one sample furniture in each sample rendering and a second sample text representing the category of at least one sample furniture; based on the position, segment at least one sample furniture image from each sample rendering; and generate a first data pair based on each sample rendering and the first sample text corresponding to each sample rendering, at least one sample furniture image, and the second sample text to obtain a first sample data set.
[0141] In some embodiments of the present disclosure, an acquisition module 110 is used to acquire a target room image; the rendering generation device includes a second determination module and a feature extraction module, the second determination module is used to determine the room depth information corresponding to the target room image; the feature extraction module is used to extract features from the room depth information to obtain room depth features; the generation module 130 is used to generate a target rendering based on the first fusion feature, the text feature and the room depth feature, wherein the target rendering can characterize the placement effect of the furniture in the furniture image in the target room whose room type is the same as the first text.
[0142] In some embodiments of the present disclosure, the generation module 130 is used to: input the first fusion feature, text feature and room depth feature into a second image generation network of a pre-trained second model, perform image generation processing on the first fusion feature, text feature and room depth feature through the second image generation network, and obtain a target rendering.
[0143] In some embodiments of the present disclosure, the second model also includes a second image feature extraction network, a second text feature extraction network, a second linear mapping network, a second resampling network and a control network; wherein the second image feature extraction network is used to perform image feature extraction on the furniture image to obtain the image features of the furniture image; the second text feature extraction network is used to perform text feature extraction on the first text and the second text corresponding to the furniture image respectively to obtain the text features of the first text and the furniture category features corresponding to the furniture image; the second linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain the mapped image features of the furniture image; the second resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature; the control network is used to perform feature extraction on the room depth information to obtain the room depth feature.
[0144] In some embodiments of the present disclosure, the second model is obtained by a second model training device, and the second model training device is used to: obtain a second sample data set, the second sample data set includes multiple second data pairs, each second data pair includes a sample rendering and room depth information corresponding to the sample rendering, a first sample text representing the room type in the sample rendering, at least one sample furniture image and a second sample text representing the furniture category corresponding to the sample furniture image, the sample rendering includes at least one sample furniture, and the sample furniture image corresponds to the second sample text one-to-one; freeze the second image feature extraction network and the second text feature extraction network; and train the second linear mapping network, the second resampling network, the control network, and the second image generation network based on the second sample data set to obtain a trained second model.
[0145] In some embodiments of the present disclosure, the second model training device is used to: obtain multiple sample renderings and a first sample text representing the room type in the sample renderings; perform target detection on each of the multiple sample renderings to obtain the position of at least one sample furniture in each sample rendering and a second sample text representing the category of at least one sample furniture; based on the position, segment at least one sample furniture image from each sample rendering; perform room depth information extraction on each sample rendering to obtain room depth information corresponding to each sample rendering; and generate a second data pair based on each sample rendering and the room depth information corresponding to each sample rendering, the first sample text, at least one sample furniture image, and the second sample text to obtain a second sample data set.
[0146] The implementation process of the functions and effects of each module in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, which will not be repeated here.
[0147] The execution subject of the effect diagram generation method in the specific implementation manner of the present disclosure may be a mobile phone, a tablet computer, a server or other electronic device.
[0148] Therefore, based on any of the above embodiments, the present disclosure also provides an electronic device, which can execute the effect diagram generation method of any of the embodiments described above in the present disclosure, and the effect diagram generation device of any of the embodiments described above can be configured on the electronic device.
[0149] Fig.14 It is a schematic block diagram of the structure of an electronic device 1000 equipped with an effect graph generating device according to an embodiment of the present disclosure.
[0150] The hardware structure of the electronic device 1000 can be implemented using a bus architecture. The bus architecture can include any number of interconnected buses and bridges, depending on the specific application and overall design constraints of the hardware. The bus 1100 connects various circuits including one or more processors 1200, memory 1300 and / or hardware modules together. The bus 1100 can also connect various other circuits 1400 such as peripherals, voltage regulators, power management circuits, external antennas, etc.
[0151] The bus 1100 may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Component (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the figure only uses one connecting line, but does not mean that there is only one bus or one type of bus.
[0152] The present disclosure also provides a readable storage medium, in which a computer program is stored, and the computer program is used to implement the above method when executed by a processor. "Readable storage medium" can be any device that can contain, store, communicate, propagate or transmit a program for use in an instruction execution system, device or equipment or in combination with these instruction execution systems, devices or equipment. More specific examples of readable storage media include the following: an electrical connection portion (electronic device) with one or more wirings, a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable read-only memory (CDROM), etc.
[0153] The present disclosure also provides a computer program product. The method of the present disclosure can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed, the process or function of the present disclosure is executed in whole or in part.
[0154] The computer program or instructions may be stored in a readable storage medium or transmitted from one readable storage medium to another readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The readable storage medium may be any available medium that can be accessed or a data storage device such as a server, data center, etc. that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid state drive. The computer readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0155] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0156] The present disclosure is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0157] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0159] In the description of this specification, the description with reference to the terms "one embodiment / method", "some embodiments / methods", "example", "specific example", or "some examples" etc. means that the specific features, structures, or characteristics described in conjunction with the embodiment / method or example are included in at least one embodiment / method or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment / method or example. Moreover, the specific features, structures, or characteristics described may be combined in any one or more embodiments / methods or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments / methods or examples described in this specification and features of different embodiments / methods or examples, unless they are contradictory.
[0160] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0161] Those skilled in the art should understand that the above embodiments are only for the purpose of clearly illustrating the present disclosure, and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications may be made based on the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A method for generating an effect diagram, characterized in that: include: Obtaining a furniture image and a first text representing a room type; Determining a second text corresponding to the furniture image, wherein the second text can represent a category of furniture in the furniture image; A target rendering is generated based at least on the furniture image, the first text and the second text corresponding to the furniture image, wherein the target rendering can at least represent the placement effect of the furniture in the furniture image in a room of the same room type as the first text.
2. The method for generating an effect diagram according to claim 1, characterized in that: Generating a target rendering according to at least the furniture image, the first text, and a second text corresponding to the furniture image includes: Extracting image features of the furniture image to obtain image features of the furniture image; Performing text feature extraction on the first text and the second text corresponding to the furniture image respectively to obtain text features of the first text and furniture category features corresponding to the furniture image; fusing the image feature of the furniture image and the corresponding furniture category feature to obtain a first fused feature; and The target effect graph is generated based on at least the first fusion feature and the text feature.
3. The method for generating an effect diagram according to claim 2, characterized in that: The number of the furniture images is at least one, and the image features of the furniture images and the corresponding furniture category features are fused together to obtain a first fused feature, including: Performing linear mapping on image features of each furniture image in at least one furniture image to obtain mapped image features of each furniture image; Fusing the mapped image features of each furniture image with the corresponding furniture category features to obtain a second fused feature corresponding to each furniture image; Fusing the second fusion features corresponding to the furniture images together to obtain a third fusion feature; and The third fused feature is resampled to obtain the first fused feature, wherein the first fused feature and the text feature have the same spatial dimension.
4. The method for generating an effect diagram according to claim 2, characterized in that: Generating the target effect graph at least according to the first fusion feature and the text feature, including: The first fusion feature and the text feature are input into a first image generation network of a pre-trained first model, and image generation processing is performed on the first fusion feature and the text feature through the first image generation network to obtain the target effect image.
5. The method for generating an effect diagram according to claim 4, characterized in that: The first model also includes a first image feature extraction network, a first text feature extraction network, a first linear mapping network and a first resampling network; Wherein, the first image feature extraction network is used to extract image features from the furniture image to obtain image features of the furniture image; The first text feature extraction network is used to extract text features from the first text and the second text corresponding to the furniture image, respectively, to obtain text features of the first text and furniture category features corresponding to the furniture image; The first linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain mapped image features of the furniture image; The first resampling network is used to resample the second fusion feature or the third fusion feature to obtain the first fusion feature.
6. The method for generating an effect diagram according to claim 2, characterized in that: The method further comprises: Get the target room image; Determining room depth information corresponding to the target room image; Performing feature extraction on the room depth information to obtain room depth features; Generating the target effect graph at least according to the first fusion feature and the text feature, further comprising: The target rendering is generated according to the first fusion feature, the text feature and the room depth feature, wherein the target rendering can represent the placement effect of the furniture in the furniture image in the target room whose room type is the same as that of the first text.
7. The method for generating an effect diagram according to claim 6, characterized in that: Generating the target rendering according to the first fusion feature, the text feature, and the room depth feature includes: The first fusion feature, the text feature and the room depth feature are input into a second image generation network of a pre-trained second model, and the first fusion feature, the text feature and the room depth feature are processed by the second image generation network to generate an image, so as to obtain the target effect image.
8. The method for generating an effect diagram according to claim 7, characterized in that: The second model also includes a second image feature extraction network, a second text feature extraction network, a second linear mapping network, a second resampling network and a control network; Wherein, the second image feature extraction network is used to extract image features from the furniture image to obtain image features of the furniture image; The second text feature extraction network is used to extract text features from the first text and the second text corresponding to the furniture image, respectively, to obtain text features of the first text and furniture category features corresponding to the furniture image; The second linear mapping network is used to perform linear mapping on the image features of the furniture image to obtain mapped image features of the furniture image; The second resampling network is used to resample the second fused feature or the third fused feature to obtain the first fused feature; The control network is used to extract features from the room depth information to obtain the room depth features.
9. An electronic device, characterized in that: include: a memory storing a computer program; as well as A processor, wherein the processor executes the computer program stored in the memory, so that the processor executes the effect diagram generation method according to any one of claims 1 to 8.
10. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by a processor, it is used to implement the effect diagram generation method according to any one of claims 1 to 8.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the effect diagram generating method according to any one of claims 1 to 8 is implemented.