A method for generating a probabilistic scattering three-dimensional scene based on semantic layout constraints
Patent Information
- Application Number
- CN202611033988.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2026-04-01
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]现有技术中,基于二维平面图生成三维室内场景的方法,主要分为两类:一类是传统手工建模方法,通过设计师手动将二维平面元素转化为三维模型,该方法效率低下、耗时费力,且建模效果高度依赖设计师的专业能力,难以满足大规模、快速建模的需求;另一类是自动建模方法,通过深度学习模型或几何规则,将二维平面图自动转化为三维场景,该方法虽能提升建模效率,但其更多侧重于渲染图像的视觉真实感、多视角一致性或局部纹理细节,对室内结构与对象尺度的可控性与工程稳定性仍存在不足,其仅依赖二维平面图的像素比例或简单的几何规则进行三维尺度分配,导致在已知同一平面图布局条件下,不同对象,如:床、沙发等在生成的三维结果中往往出现相对大小漂移、占地范围偏移或与墙体、门窗等发生不合理穿插;亦或是同一对象在不同生成实例中可能出现尺寸不一以及与整体空间尺度不匹配的情况
[0062]根据上述描述可知,设置优化机制,计算所渲染出的三维场景的相对比例尺一致性损失和语义一致性损失,且相对比例尺一致性损失是通过计算出每一个像素的像素期望深度,进而计算出每一个实例的实例期望深度,结合像素尺度比例得到实际尺度,并与无量纲相对比例尺约束结合计算相对比例尺一致性损失,保证所得到的相对比例尺一致性损失的准确度,通过反向传播算法将超过第一损失阈值的总损失值反向传播至生成模型与深度预测网络,实现对生成模型与深度预测网络的优化,提升后续三维场景生成的准确度。
Smart Images

Figure CN122821045A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of 3D reconstruction technology, and in particular to a method for generating 3D scenes based on semantic layout constraints using probabilistic scattering. Background Technology
[0002] With the development of applications such as digital twins, indoor navigation, virtual simulation, home decoration design and asset management, the reconstruction and generation of three-dimensional indoor scenes based on two-dimensional floor plans has become an important research and engineering direction.
[0003] In existing technologies, methods for generating 3D interior scenes based on 2D floor plans mainly fall into two categories: one is the traditional manual modeling method, where designers manually convert 2D elements into 3D models. This method is inefficient, time-consuming, and labor-intensive, and the modeling effect is highly dependent on the designer's professional skills, making it difficult to meet the needs of large-scale, rapid modeling. The other is the automatic modeling method, which automatically converts 2D floor plans into 3D scenes through deep learning models or geometric rules. While this method can improve modeling efficiency, it focuses more on the visual realism, multi-view consistency, or local texture details of the rendered image. It still has shortcomings in the controllability and engineering stability of the scale of interior structures and objects. It relies solely on the pixel ratio of the 2D floor plan or simple geometric rules for 3D scale allocation, resulting in different objects, such as beds and sofas, often exhibiting relative size drift, offset footprint, or unreasonable intersection with walls, doors, and windows in the generated 3D results, even with a known floor plan layout. Alternatively, the same object may have different sizes and mismatches with the overall spatial scale in different generated instances. This leads to the final generated 3D scene failing to meet actual needs. Summary of the Invention
[0004] The technical problem to be solved by this invention is: to provide a probabilistic scattering 3D scene generation method based on semantic layout constraints, thereby improving the scale stability, spatial occupancy rationality, and practical usability of 3D scene generation.
[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention provides a method for generating a probabilistic scattering 3D scene based on semantic layout constraints, comprising:
[0007] A two-dimensional vector plane image is obtained, and the two-dimensional vector plane image is normalized and rasterized to obtain a multi-channel semantic layout image. At the same time, an instance mask for each instance is generated based on the instance connected regions in the multi-channel semantic layout image. The pixel scale ratio of the corresponding instance is calculated based on the instance mask. The instance category, instance mask and pixel scale ratio of all instances are integrated to generate an instance-level layout set.
[0008] For each instance in the instance-level layout set, construct a corresponding dimensionless relative scale constraint to obtain a relative scale constraint set.
[0009] The multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set are input into a pre-trained generative model to output a two-dimensional latent space feature map. The two-dimensional latent space feature map is then input into a depth prediction network for depth prediction to output the predicted depth of each pixel in the two-dimensional latent space feature map. A depth prior for each pixel is constructed based on the instance-level layout set and the relative scale constraint set. The depth prior for each pixel is then weighted and fused with the corresponding predicted depth to obtain a depth probability distribution.
[0010] All pixel coordinates in the multi-channel semantic layout map are transformed into normalized gaze coordinates, and the two-dimensional latent space feature map is scattered into the three-dimensional space along the gaze coordinates in a weighted projection manner according to the depth probability distribution to construct three-dimensional features, and a three-dimensional scene is rendered based on the three-dimensional features.
[0011] The beneficial effects of this application are as follows: By normalizing and rasterizing the two-dimensional vector planar image, the resulting multi-channel semantic layout image retains the original semantic layout information. It not only constructs instance-level pixel scale ratios but also instance-level dimensionless relative scale constraints, ensuring that the semantic layout information and scale constraint information of the two-dimensional vector planar image can be accurately transmitted to the generation process of the three-dimensional scene. This achieves precise control over the scale of indoor structures and objects. The two-dimensional latent space feature map obtained based on the multi-channel semantic layout image, instance-level layout set, and relative scale constraint set is input into the depth prediction network for depth prediction of each pixel, thus providing a basis for pixel-level depth prediction. It provides reliable constraint guidance, effectively reducing depth prediction errors. It constructs a depth prior for each pixel based on the instance-level layout set and the relative scale constraint set. The depth prior is used to constrain the relative scale relationship between different instances, which can reduce object scale drift during subsequent 3D scene generation. The depth probability distribution obtained by weighted fusion of depth prior and depth prediction is used as a 2D latent space feature map and scattered into the 3D space depth basis in a weighted projection manner, improving the accuracy and spatial hierarchy of the obtained 3D features. This improves the rendered 3D scene to have clear scale stability and reasonable spatial occupancy, which can meet practical needs and improve practical usability.
[0012] Optionally, the normalization and rasterization of the two-dimensional vector plane map to obtain a multi-channel semantic layout map includes:
[0013] All vector coordinates in the two-dimensional vector plane diagram are uniformly mapped to the same plane coordinate system and normalized to obtain a normalized two-dimensional vector plane diagram;
[0014] Obtain the target output resolution, and then rasterize the normalized two-dimensional vector plane image according to the target output resolution to obtain a rasterized two-dimensional vector plane image. Finally, split the rasterized two-dimensional vector plane image into channels according to instance categories to obtain a multi-channel semantic layout image.
[0015] As described above, by mapping all vector coordinates in the two-dimensional vector plane graph to the same plane coordinate system for normalization, the interference caused by the difference in coordinate systems is eliminated, ensuring that the position and relative relationship of each instance in the two-dimensional vector plane graph are accurate and controllable. The normalized two-dimensional vector plane graph is rasterized by the target output resolution to ensure the specific uniformity and standardization of the rasterized two-dimensional vector plane graph. At the same time, the channels are split according to the instance category, so that each instance category has an independent channel, which can accurately distinguish the semantic information of different instances and avoid semantic confusion.
[0016] Optionally, calculating the pixel scale ratio of the corresponding instance based on the instance mask includes:
[0017] Obtain the height in pixels of the multi-channel semantic layout map and the bounding rectangle of each instance mask. Input the pixel height of each bounding rectangle and the height in pixels into a first formula for calculation to obtain the pixel scale ratio of the corresponding instance. The first formula is:
[0018] ;
[0019] ;
[0020] in, H represents the pixel scale ratio of the i-th instance, and H represents the height in pixels. This represents the pixel height of instance i. This represents the maximum pixel coordinate within the instance mask of instance i. This represents the instance mask of instance i. The x-coordinate of the pixel coordinate. The ordinate represents the pixel coordinate. Represents the minimum pixel coordinates within the instance mask of instance i.
[0021] As described above, by combining the pixel height corresponding to the outer rectangle of the instance mask with the height of the multi-channel semantic layout graph, the pixel scale ratio of the instance is calculated. Instead of simply using the pixel width or height of the instance as a scale reference, this method can truly reflect the relative size of each instance in the two-dimensional vector plane graph, avoiding the scale disorder of objects in the three-dimensional scene due to scale errors.
[0022] Optionally, constructing a corresponding dimensionless relative scale constraint for each instance in the instance-level layout set includes:
[0023] Obtain the bounding rectangle of each instance mask, and input the length of the longer side of each bounding rectangle into the second formula for calculation to construct the dimensionless relative scale constraint corresponding to each instance. The second formula is:
[0024] ;
[0025] ;
[0026] in, This represents the dimensionless relative scale constraint for the i-th instance. This represents the preset tolerance hyperparameter. Indicates the dimensionless relative scale of instance i. This represents the length of the longer side of instance i. This indicates the overall reference length corresponding to the multi-channel semantic layout diagram.
[0027] As described above, according to the second formula, when performing dimensionless relative scale constraint calculation, the length of the long side of the bounding rectangle of the instance mask is combined with the overall reference length of the multi-channel semantic layout map, and a preset tolerance hyperparameter is set to set a reasonable scale range for each instance, thereby achieving precise constraint on the instance scale, effectively avoiding problems such as instance scale drift, inconsistent dimensions, and unreasonable interleaving, and improving the engineering stability of 3D scene generation.
[0028] Optionally, inputting the multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set into a pre-trained generative model to output a two-dimensional latent space feature map includes:
[0029] The multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set are input into a pre-trained generative model. At the same time, a random noise vector is introduced, so that the generative model performs multi-scale convolutional encoding on the multi-channel semantic layout map to generate multi-scale layout features. Simultaneously, based on the instance-level layout set and the relative scale constraint set, each instance is encoded to generate instance features. The random noise vector is mapped to a low-resolution latent variable feature map, wherein the multi-scale layout features include multi-scale layout spatial features and multi-scale layout semantic features.
[0030] The multi-scale layout features are concatenated and convolved with the low-resolution latent variable feature map. At the same time, the instance features are cross-attentioned with the low-resolution latent variable feature map to achieve iterative updates of the low-resolution latent variable feature map until the termination condition is met, so as to output the iteratively updated low-resolution latent variable feature map, which is a two-dimensional latent space feature map.
[0031] As described above, not only are multi-scale layout features, including multi-scale layout spatial features and multi-scale layout semantic features, extracted, but also instance features are extracted. The multi-scale layout features and instance features are fused together to iteratively update the low-resolution latent variable feature map, ensuring that the output two-dimensional latent space feature map not only contains global layout rules but also covers the personalized features of each instance. Furthermore, the initial low-resolution latent variable feature map is obtained based on random noise vector mapping, ensuring that the output two-dimensional latent space feature map has diversity.
[0032] Optionally, the step of inputting the two-dimensional latent space feature map into a depth prediction network for depth prediction, outputting the predicted depth of each pixel in the two-dimensional latent space feature map, constructing a depth prior for each pixel based on the instance-level layout set and the relative scale constraint set, and weightedly fusing the depth prior of each pixel with the corresponding predicted depth to obtain a depth probability distribution includes:
[0033] Obtain a preset number of depth layers and a preset relative depth range in three-dimensional space. Input the preset number of depth layers and the preset relative depth range into a logarithmic interval formula to generate a discrete sequence of relative depths. The discrete sequence of relative depths is a discretized depth sequence expressed in logarithmic intervals. The logarithmic interval formula is as follows:
[0034] ;
[0035] ;
[0036] in, This represents the k-th depth layer in the discretization. This represents the minimum value within the preset relative depth range. This represents the maximum value within a preset relative depth range. Indicates the preset depth number of layers;
[0037] The depth prediction network uses a multi-scale convolutional structure to enhance the features of the two-dimensional latent space feature map to extract pixel-level spatial features and pixel-level semantic features. Based on the pixel-level spatial features and the pixel-level semantic features, the network outputs the predicted depth of each pixel in each depth layer of the relative depth discrete sequence through a self-attention mechanism.
[0038] The depth prior of each pixel is constructed based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set. The depth prior of each pixel is then weighted and fused with the corresponding predicted depth to obtain the depth probability distribution.
[0039] As described above, using the logarithmic interval formula to generate a relative depth discrete sequence can accurately capture depth details at different depth layers, avoiding the loss of depth details or depth redundancy. When performing depth prediction, it combines not only pixel-level spatial features but also pixel-level semantic features, improving the accuracy of the predicted depth of each pixel at each depth layer in the relative depth discrete sequence. It also weights and fuses the predicted depth with pixel-level depth priors to achieve probabilistic quantification of depth information.
[0040] Optionally, the step of constructing a depth prior for each pixel based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set, and then weighting and fusing the depth prior of each pixel with the corresponding predicted depth to obtain the depth probability distribution includes:
[0041] The pixel scale ratio of each instance and the corresponding dimensionless relative scale constraint are input into the center formula and variance formula, respectively, to calculate the depth prior center and depth prior variance for each instance. The center formula is as follows:
[0042] ;
[0043] in, Denotes the depth prior center of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers;
[0044] The variance formula is:
[0045] ;
[0046] in, Let represent the depth prior variance of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers, This represents the variance adjustment coefficient;
[0047] Based on the depth prior center and depth prior variance of each instance, and combined with the instance mask of each instance in the instance-level layout set, the depth prior center and depth prior variance of each instance are discretized to the relative depth discrete sequence through Gaussian distribution, and the depth prior of each pixel on each depth layer in the relative depth discrete sequence is constructed.
[0048] The depth prior at each depth layer is weighted and fused with the corresponding predicted depth input using a weighted fusion formula to obtain the depth probability distribution of each pixel at each depth layer. The weighted fusion formula is as follows:
[0049] ;
[0050] in, Let represent the depth probability distribution of pixel (u,v) at the k-th depth layer, where u represents the x-coordinate of the pixel. Represents the ordinate of a pixel. Representing depth layer The corresponding index, Indicates the first A normalized function at each depth layer Indicates the first Predicted depth at each depth layer Indicates the first Depth priors at each depth layer This represents the prior weight coefficient.
[0051] As described above, by combining the pixel scale ratio of each instance with the dimensionless relative scale constraint, the depth prior center and depth prior variance of each instance are first calculated. Then, based on the instance mask, the depth prior center and depth prior variance of each instance are discretized into a relative depth discrete sequence through Gaussian distribution, so that each pixel has a corresponding depth prior at each depth layer, achieving accurate adaptation between the depth prior and the depth layer, ensuring the rationality and reliability of the fusion, and improving the accuracy of the depth probability distribution.
[0052] Optionally, the step of converting all pixel coordinates in the multi-channel semantic layout map into normalized gaze coordinates, and scattering the two-dimensional latent space feature map along the gaze coordinates in a weighted projection manner according to the depth probability distribution into three-dimensional space to construct three-dimensional features includes:
[0053] Determine whether there are calibrated camera parameters. If they exist, obtain the principal point parameters and equivalent focal length parameters from the camera parameters. If they do not exist, obtain the width and height pixel counts of the multi-channel semantic layout map, calculate the principal point parameters based on the width and height pixel counts, and obtain a preset reference focal length as the equivalent focal length parameters.
[0054] The principal point parameter and the equivalent focal length parameter are used to transform all pixel coordinates in the multi-channel semantic layout map into normalized line coordinates.
[0055] Based on the line-of-sight coordinates and the depth layer of the depth probability distribution, a projection sampling point corresponding to each two-dimensional pixel in the two-dimensional latent space feature map is constructed in three-dimensional space. Each two-dimensional pixel in the two-dimensional latent space feature map is weighted according to the depth probability distribution to obtain a weighted feature value for each two-dimensional pixel. The weighted feature value of each two-dimensional pixel is used as the three-dimensional feature value of the corresponding projection sampling point. The three-dimensional feature value of each projection sampling point is accumulated into the three-dimensional voxel grid corresponding to the three-dimensional space through an accumulation formula to construct the three-dimensional feature.
[0056] As described above, by flexibly acquiring principal point parameters and equivalent focal length parameters, pixel coordinates are transformed into normalized view coordinates, eliminating the influence of image resolution and size. This ensures accurate correspondence between 2D pixels and their 3D spatial positions, avoiding 3D scene structure chaos caused by coordinate transformation deviations. Based on view coordinates and depth probability distribution, projection sampling points are constructed for each 2D pixel in 3D space. These projection sampling points not only conform to the 2D spatial position but also match the depth characteristics of the depth probability distribution, achieving accurate 2D-to-3D projection and ensuring that 3D features can accurately convey 2D semantic layout and depth information.
[0057] Optionally, it also includes:
[0058] The three-dimensional scene is input into a pre-trained semantic prediction network to obtain the scene semantic prediction layout. At the same time, the original semantic layout of the multi-channel semantic layout map is obtained. The original semantic layout and the scene semantic prediction layout are input into a first loss function for calculation to obtain the semantic consistency loss.
[0059] The expected pixel depth of each pixel in the two-dimensional latent space feature map is calculated based on the depth probability distribution. The median of the expected pixel depths of all pixels within the same instance mask is calculated and used as the expected instance depth of the instance corresponding to the instance mask. The expected instance depth of each instance is combined with the corresponding pixel scale ratio to obtain the actual scale of each instance. The actual scales of all instances are input into the second loss function along with the corresponding dimensionless relative scale constraint to obtain the relative scale consistency loss.
[0060] The relative scale consistency loss and the semantic consistency loss are input into the total loss function to calculate the total loss value;
[0061] If the total loss value exceeds the first loss threshold, the total loss value is backpropagated to the generative model and the deep prediction network through the backpropagation algorithm to achieve iterative optimization of the generative model and the deep prediction network, resulting in an iteratively optimized generative model and an iteratively optimized deep prediction network.
[0062] As described above, an optimization mechanism is set up to calculate the relative scale consistency loss and semantic consistency loss of the rendered 3D scene. The relative scale consistency loss is calculated by calculating the expected depth of each pixel, then the expected depth of each instance, and combining the pixel scale ratio to obtain the actual scale. This is then combined with the dimensionless relative scale constraint to calculate the relative scale consistency loss, ensuring the accuracy of the obtained relative scale consistency loss. The total loss value exceeding the first loss threshold is backpropagated to the generative model and depth prediction network through the backpropagation algorithm, thereby optimizing the generative model and depth prediction network and improving the accuracy of subsequent 3D scene generation. Attached Figure Description
[0063] Figure 1 This is a flowchart of a probabilistic scattering 3D scene generation method based on semantic layout constraints provided in this embodiment;
[0064] Figure 2 This is a schematic diagram of the overall process of a probabilistic scattering 3D scene generation method based on semantic layout constraints provided in this embodiment. Detailed Implementation
[0065] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the present invention can be understood more clearly and thoroughly, and that the scope of the present invention can be fully conveyed to those skilled in the art.
[0066] Example 1
[0067] Please refer to Figures 1 to 2 This invention provides a method for generating a 3D scene based on probabilistic scattering with semantic layout constraints, comprising the following steps:
[0068] S1. Obtain a two-dimensional vector plane image, perform normalization and rasterization processing on the two-dimensional vector plane image to obtain a multi-channel semantic layout image, and generate an instance mask for each instance based on the instance connected regions in the multi-channel semantic layout image. Calculate the pixel scale ratio of the corresponding instance based on the instance mask, and integrate the instance category, instance mask, and pixel scale ratio of all instances to generate an instance-level layout set.
[0069] In this embodiment, as Figure 2 As shown, the acquired 2D vector planar image is normalized and rasterized to obtain a multi-channel semantic layout map. This 2D vector planar image includes vector coordinates, instance categories, line segments, shapes, and outlines for each instance. Instance categories include: walls, doors, windows, beds, rooms, etc. Simultaneously, an instance mask is generated for each instance based on the connected regions of the instances in the multi-channel semantic layout map. The instance mask represents the pixel position of the instance in the multi-channel semantic layout map, and its core function is to mark the association between instances and pixels. Regarding instance masks, semantic segmentation can also be performed on the multi-channel semantic layout map, and the instance mask for each instance can be generated based on the semantic segmentation results. The pixel scale ratio of the corresponding instance is calculated based on the instance mask. The instance categories, instance masks, and pixel scale ratios of all instances are integrated to generate an instance-level layout set. That is, a corresponding instance layout information is generated based on the instance category, instance mask, and pixel scale ratio of each instance, and the instance layout information of all instances is integrated to obtain the instance-level layout set.
[0070] At this point, the normalization and rasterization processing of the two-dimensional vector plane map in step S1 to obtain a multi-channel semantic layout map includes:
[0071] S11. Map all vector coordinates in the two-dimensional vector plane diagram to the same plane coordinate system for normalization processing to obtain a normalized two-dimensional vector plane diagram.
[0072] S12. Obtain the target output resolution, and perform rasterization processing on the normalized two-dimensional vector plane image according to the target output resolution to obtain a rasterized two-dimensional vector plane image. Then, perform channel splitting on the rasterized two-dimensional vector plane image according to the instance category to obtain a multi-channel semantic layout image.
[0073] In this embodiment, as Figure 2 As shown, all vector coordinates in the 2D vector plane image are uniformly mapped to the same plane coordinate system for normalization. The normalized 2D vector plane image is then rasterized according to the target output resolution to obtain a rasterized 2D vector plane image. For example, if the target output resolution is W1×W2, then the width of the rasterized 2D vector plane image is W1, and the height is W2. In fact, the height W2 is the height (in pixels H) of the multi-channel semantic layout image in step S13. The rasterized 2D vector plane image is then split into channels according to instance categories, that is, each instance category corresponds to an independent channel to obtain a multi-channel semantic layout image.
[0074] At this point, the step S1 of calculating the pixel scale ratio of the corresponding instance based on the instance mask includes:
[0075] S13. Obtain the height pixel count of the multi-channel semantic layout map and the bounding rectangle of each instance mask. Input the pixel height of each bounding rectangle and the height pixel count into the first formula for calculation to obtain the pixel scale ratio of the corresponding instance. The first formula is:
[0076] ;
[0077] ;
[0078] in, H represents the pixel scale ratio of the i-th instance, and H represents the height in pixels. This represents the pixel height of instance i. This represents the maximum pixel coordinate within the instance mask of instance i. This represents the instance mask of instance i. The x-coordinate of the pixel coordinate. The ordinate represents the pixel coordinate. Represents the minimum pixel coordinates within the instance mask of instance i.
[0079] In this embodiment, as Figure 2 As shown, the pixel height of each bounding rectangle and the number of pixels of the height of the multi-channel semantic layout map are calculated to obtain the pixel scale ratio of the corresponding instance.
[0080] S2. Construct a corresponding dimensionless relative scale constraint for each instance in the instance-level layout set to obtain a relative scale constraint set.
[0081] At this point, the step S2 of constructing a corresponding dimensionless relative scale constraint for each instance in the instance-level layout set includes:
[0082] S21. Obtain the bounding rectangle of each instance mask, and input the length of the longer side of each bounding rectangle into the second formula for calculation to construct the dimensionless relative scale constraint corresponding to each instance. The second formula is:
[0083] ;
[0084] ;
[0085] in, This represents the dimensionless relative scale constraint for the i-th instance. This represents the preset tolerance hyperparameter. Indicates the dimensionless relative scale of instance i. This represents the length of the longer side of instance i. This indicates the overall reference length corresponding to the multi-channel semantic layout diagram.
[0086] In this embodiment, as Figure 2 As shown, the ratio of the length of the longer side of the bounding rectangle of each instance mask to the overall reference length corresponding to the multi-channel semantic layout graph is used as the dimensionless relative scale of the corresponding instance. Furthermore, a preset tolerance hyperparameter is introduced into the calculated dimensionless relative scale of each instance to construct the dimensionless relative scale constraint for each instance. The dimensionless relative scale of the instance can also be obtained through the instance category prior template.
[0087] S3. Input the multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set into the pre-trained generative model to output a two-dimensional latent space feature map. Input the two-dimensional latent space feature map into the depth prediction network to perform depth prediction and output the predicted depth of each pixel in the two-dimensional latent space feature map. Construct a depth prior for each pixel based on the instance-level layout set and the relative scale constraint set. Weight and fuse the depth prior of each pixel with the corresponding predicted depth to obtain the depth probability distribution.
[0088] At this point, step S3, which involves inputting the multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set into the pre-trained generative model to output a two-dimensional latent space feature map, includes:
[0089] S31. Input the multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set into the pre-trained generative model, and simultaneously introduce a random noise vector, so that the generative model performs multi-scale convolutional encoding on the multi-channel semantic layout map to generate multi-scale layout features. At the same time, it performs instance encoding on each instance based on the instance-level layout set and the relative scale constraint set to generate instance features, and maps the random noise vector into a low-resolution latent variable feature map, wherein the multi-scale layout features include multi-scale layout spatial features and multi-scale layout semantic features.
[0090] S32. The multi-scale layout features are concatenated and convolved with the low-resolution latent variable feature map, and the instance features are cross-attention calculated with the low-resolution latent variable feature map to iteratively update the low-resolution latent variable feature map until the termination condition is met, so as to output the iteratively updated low-resolution latent variable feature map, wherein the iteratively updated low-resolution latent variable feature map is a two-dimensional latent space feature map.
[0091] In this embodiment, as Figure 2 As shown, the multi-channel semantic layout map, instance-level layout set, and relative scale constraint set are input into the pre-trained generative model. Random noise vectors are also introduced to enable the generative model to capture both global layout features and individual instance-specific features. Specifically, the generative model performs multi-scale convolutional encoding on the multi-channel semantic layout map through convolutional layers to generate multi-scale layout features that include both multi-scale spatial and semantic features. Simultaneously, the generative model uses a token layer to encode each instance based on the instance-level layout set and relative scale constraint set, constructing a token sequence, i.e., generating instance features. The instance features can be represented as: T = {t} i} i=1,2…,N Where T represents the token sequence, t i Let t represent the instance characteristics of instance i, and N represent the total number of instances, where t iThe dataset includes at least the following information: instance category, instance mask, pixel scale ratio, and dimensionless relative scale constraint. Random noise vectors are mapped to low-resolution latent variable feature maps. Multi-scale layout features are concatenated and convolved with the low-resolution latent variable feature map, while instance features are cross-attentionally calculated with the low-resolution latent variable feature map to iteratively update the low-resolution latent variable feature map until a termination condition is reached. This iterative update of the low-resolution latent variable feature map is essentially feature fusion of multi-scale layout features and instance features. The termination condition can be a preset iteration threshold or a fusion loss calculation using the cross-entropy loss function on the iteratively updated low-resolution latent variable feature map. Iteration stops when the calculated fusion loss value is lower than the preset threshold. The final iteratively updated low-resolution latent variable feature map is used as a two-dimensional latent space feature map. The pre-trained generative models include: diffusion models, conditional diffusion models, and global flow models.
[0092] At this point, step S3 involves inputting the two-dimensional latent space feature map into a depth prediction network for depth prediction, outputting the predicted depth of each pixel in the two-dimensional latent space feature map, and constructing a depth prior for each pixel based on the instance-level layout set and the relative scale constraint set. The depth prior of each pixel is then weighted and fused with the corresponding predicted depth to obtain a depth probability distribution, including:
[0093] S33. Obtain the preset number of depth layers and the preset relative depth range in three-dimensional space. Input the preset number of depth layers and the preset relative depth range into the logarithmic interval formula to generate a discrete sequence of relative depths. The discrete sequence of relative depths is a discretized depth layer expressed in logarithmic intervals. The logarithmic interval formula is:
[0094] ;
[0095] ;
[0096] in, This represents the k-th depth layer in the discretization. This represents the minimum value within the preset relative depth range. This represents the maximum value within a preset relative depth range. Indicates the preset depth number of layers;
[0097] S34. The depth prediction network uses a multi-scale convolutional structure to enhance the features of the two-dimensional latent space feature map to extract pixel-level spatial features and pixel-level semantic features. Based on the pixel-level spatial features and the pixel-level semantic features, the network outputs the predicted depth of each pixel in each depth layer of the relative depth discrete sequence through a self-attention mechanism.
[0098] S35. Construct a depth prior for each pixel based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set. Then, weight and fuse the depth prior of each pixel with the corresponding predicted depth to obtain the depth probability distribution.
[0099] In this embodiment, as Figure 2 As shown, the preset depth layers and preset relative depth ranges in 3D space are input into the logarithmic interval formula, representing the discretized depth components as a logarithmic interval ratio. This generates a relative depth discrete sequence. According to the logarithmic interval formula, the density of the near depth layer is greater than that of the distant depth component, allowing the near depth layer to capture the depth details of nearby instances while avoiding redundancy in distant depths, thus ensuring the accuracy of subsequent depth predictions. The depth prediction network uses a multi-scale convolutional structure to enhance the features of the 2D spatial feature map output in step S32, extracting pixel-level spatial and semantic features. Based on these pixel-level spatial and semantic features, a self-attention mechanism is used to output the predicted depth of each pixel at each depth layer in the relative depth discrete sequence. The predicted depth is output in logits form, representing the unnormalized log probability of a pixel at each depth layer, denoted as . ,in, Indicates the first The predicted depth at each depth layer, where u represents the x-coordinate of a pixel. Represents the ordinate of a pixel. Representing depth layer The corresponding index. The predicted depth is actually the depth confidence of a pixel in the corresponding depth layer.
[0100] Based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set, a depth prior for each pixel is constructed. The depth prior of each pixel is then weighted and fused with the corresponding predicted depth to obtain the depth probability distribution. In other words, the predicted depth of each pixel on all depth layers is normalized through the depth prior, and the discrete predicted depth is normalized into a continuous predicted depth.
[0101] At this point, step S35 includes:
[0102] S351. Input the pixel scale ratio of each instance and the corresponding dimensionless relative scale constraint into the center formula and variance formula respectively to calculate the depth prior center and depth prior variance of each instance. The center formula is as follows:
[0103] ;
[0104] in, Denotes the depth prior center of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers;
[0105] The variance formula is:
[0106] ;
[0107] in, Let represent the depth prior variance of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers, This represents the variance adjustment coefficient;
[0108] S352. Based on the depth prior center and depth prior variance of each instance, and combined with the instance mask of each instance in the instance-level layout set, the depth prior center and depth prior variance of each instance are discretized to the relative depth discrete sequence through Gaussian distribution, and the depth prior of each pixel on each depth layer in the relative depth discrete sequence is constructed.
[0109] S353. The depth prior of each depth layer is weighted and fused with the corresponding predicted depth input using a weighted fusion formula to obtain the depth probability distribution of each pixel in each depth layer. The weighted fusion formula is as follows:
[0110] ;
[0111] in, Let represent the depth probability distribution of pixel (u,v) at the k-th depth layer, where u represents the x-coordinate of the pixel. Represents the ordinate of a pixel. Representing depth layer The corresponding index, Indicates the first A normalized function at each depth layer Indicates the first Predicted depth at each depth layer Indicates the first Depth priors at each depth layer This represents the prior weight coefficient.
[0112] In this embodiment, as Figure 2 As shown, the pixel scale ratio of each instance and the corresponding dimensionless relative scale constraint are input into the center formula and variance formula, respectively, to calculate the depth prior center and depth prior variance for each instance. The hyperparameters in the center and variance formulas are used to prevent the denominator from being zero. The calculated depth prior center and depth prior variance are instance-level. Since the instance mask marks the association between instances and pixels, it is known all pixels contained in the same instance mask. Therefore, based on the depth prior center and depth prior variance of each instance, combined with the instance mask of each instance in the instance-level layout set, the depth prior center and depth prior variance of each instance are discretized to a relative depth discrete sequence using a Gaussian distribution. That is, the depth prior center and depth prior variance of each instance are assigned to all pixels within the corresponding instance mask, constructing the depth prior for each pixel at each depth layer in the relative depth discrete sequence. In practice, the depth prior is used to constrain the relative scale relationship between different instances, reducing object scale drift during subsequent 3D scene generation. The depth prior at each depth layer is weighted and fused with the corresponding predicted depth input using a weighted fusion formula. Specifically, the depth prior and the predicted depth are logarithmically added together, and then normalized using a normalization function to obtain the depth probability distribution of each pixel at each depth layer. The prior weight coefficient in the weighted fusion formula is used to adjust the influence of the depth prior on the depth probability distribution. When the prior weight coefficient increases, the influence of the depth prior on the depth probability distribution is enhanced; conversely, when the prior weight coefficient decreases, the depth probability distribution is more dependent on the depth prediction.
[0113] S4. Convert all pixel coordinates in the multi-channel semantic layout map into normalized gaze coordinates, and scatter the two-dimensional latent space feature map into the three-dimensional space along the gaze coordinates in a weighted projection manner according to the depth probability distribution to construct three-dimensional features, and render a three-dimensional scene based on the three-dimensional features.
[0114] In this embodiment, all pixel coordinates in the multi-channel semantic layout map are transformed into normalized gaze coordinates, and the two-dimensional latent space feature map is scattered into the three-dimensional space along the gaze coordinates in a weighted projection manner according to the depth probability distribution to construct three-dimensional features. The density and color of the three-dimensional features are decoded by a multilayer perceptron, and the volume rendering of the decoded three-dimensional features is performed to render a three-dimensional scene.
[0115] At this point, step S4, which involves converting all pixel coordinates in the multi-channel semantic layout map into normalized gaze coordinates and scattering the two-dimensional latent space feature map along the gaze coordinates using a weighted projection method according to the depth probability distribution to construct three-dimensional features, includes:
[0116] S41. Determine whether there are calibrated camera parameters. If they exist, obtain the principal point parameter and the equivalent focal length parameter from the camera parameters. If they do not exist, obtain the width pixel number and height pixel number of the multi-channel semantic layout map, calculate the principal point parameter based on the width pixel number and the height pixel number, and obtain the preset reference focal length as the equivalent focal length parameter.
[0117] S42. By using the principal point parameter and the equivalent focal length parameter, all pixel coordinates in the multi-channel semantic layout map are transformed into normalized line-of-sight coordinates.
[0118] S43. Construct projection sampling points in three-dimensional space for each two-dimensional pixel in the two-dimensional latent space feature map according to the depth layer of the line-of-sight coordinates and the depth probability distribution. Perform weighted processing on each two-dimensional pixel in the two-dimensional latent space feature map according to the depth probability distribution to obtain the weighted feature value of each two-dimensional pixel. Use the weighted feature value of each two-dimensional pixel as the three-dimensional feature value of the corresponding projection sampling point. Accumulate the three-dimensional feature value of each projection sampling point into the three-dimensional voxel grid corresponding to the three-dimensional space through the accumulation formula to construct the three-dimensional feature.
[0119] In this embodiment, a flexible approach is used to obtain the principal point parameters and equivalent focal length parameters. If calibrated camera parameters exist, the principal point parameters and equivalent focal length parameters are obtained from them. If they do not exist, the principal point parameters are calculated based on the width and height pixel counts of the multi-channel semantic layout map, and a preset reference focal length is obtained as the equivalent focal length parameter. The preset reference focal length can be a unit value. The specific calculation of the principal point parameters is as follows:
[0120] Input the width in pixels into the formula for the principal point's x-coordinate to calculate the x-coordinate of the principal point parameter. The formula for the principal point's x-coordinate is:
[0121] ;
[0122] in, The x-coordinate of the principal point parameter is represented by W, and the width is represented by the number of pixels.
[0123] Simultaneously, the height in pixels is input into the formula for the principal point's ordinate to calculate the ordinate of the principal point parameter. The formula for the principal point's ordinate is:
[0124] ;
[0125] in, H represents the ordinate of the principal point parameter, and H represents the height in pixels.
[0126] By using principal point parameters and equivalent focal length parameters, all pixel coordinates in the multi-channel semantic layout map are transformed into normalized gaze coordinates. Specifically, this can be achieved using the gaze normalization formula, as follows:
[0127] ;
[0128] in, This represents the x-coordinate of the normalized view coordinates, where n is the normalization flag and u represents the x-coordinate of the pixel. The x-coordinate of the principal point parameter. The x-axis represents the equivalent focal length parameter.
[0129] ;
[0130] in, The ordinate represents the normalized view coordinates, where n is the normalization flag and v represents the ordinate of the pixel. The ordinate of the principal point parameter is represented by its vertical axis. The vertical axis represents the equivalent focal length parameter.
[0131] Based on the depth layer constructed from the gaze coordinates and depth probability distribution, each two-dimensional pixel in the two-dimensional latent space feature map corresponds to a projection sampling point in three-dimensional space, i.e., a three-dimensional sampling point, which can be represented as: ,in Represents pixels ( The projection sampling points on the k-th depth layer, This represents the k-th depth layer in the discretization. The x-coordinate represents the normalized line-of-sight coordinates. The ordinate represents the normalized line-of-sight coordinates. This represents the direction vector of the normalized gaze coordinates. Each 2D pixel in the 2D latent space feature map is weighted according to the depth probability distribution to obtain a weighted feature value for each 2D pixel. This weighted feature value is then used as the 3D feature value of the corresponding projection sampling point. The 3D feature value of each projection sampling point is accumulated into the corresponding 3D voxel mesh using an accumulation formula to construct the 3D feature. The accumulation formula is:
[0132] ;
[0133] in, This represents the 3D feature corresponding to voxel g in a 3D voxel mesh, where u represents the x-coordinate of the pixel. Represents the ordinate of a pixel. Representing depth layer The corresponding index, This represents the depth probability distribution of pixel (u,v) at the k-th depth layer. Indicates the weights of the trilinear interpolation. Represents pixels ( ) Projected sampling points on the k-th depth layer.
[0134] In fact, according to the accumulation formula, the three-dimensional feature value of each projection sampling point is accumulated into the corresponding voxel of the three-dimensional voxel grid in the three-dimensional space. Through pixel-by-pixel and depth-by-depth layer accumulation operation, the three-dimensional feature values of all projection sampling points are integrated into the corresponding voxel of the three-dimensional voxel grid, and finally the three-dimensional feature is constructed.
[0135] In this embodiment, an optimization mechanism is also provided, and the specific steps are as follows:
[0136] The three-dimensional scene is input into a pre-trained semantic prediction network to obtain the scene semantic prediction layout. At the same time, the original semantic layout of the multi-channel semantic layout map is obtained. The original semantic layout and the scene semantic prediction layout are input into a first loss function for calculation to obtain the semantic consistency loss.
[0137] The expected pixel depth of each pixel in the two-dimensional latent space feature map is calculated based on the depth probability distribution. The median of the expected pixel depths of all pixels within the same instance mask is calculated and used as the expected instance depth of the instance corresponding to the instance mask. The expected instance depth of each instance is combined with the corresponding pixel scale ratio to obtain the actual scale of each instance. The actual scales of all instances are input into the second loss function along with the corresponding dimensionless relative scale constraint to obtain the relative scale consistency loss.
[0138] The relative scale consistency loss and the semantic consistency loss are input into the total loss function to calculate the total loss value;
[0139] If the total loss value exceeds the first loss threshold, the total loss value is backpropagated to the generative model and the deep prediction network through the backpropagation algorithm to achieve iterative optimization of the generative model and the deep prediction network, resulting in an iteratively optimized generative model and an iteratively optimized deep prediction network.
[0140] In this embodiment, the rendered 3D scene undergoes subsequent evaluation, and the model / network in the 3D scene generation process is optimized based on the evaluation results. The 3D scene is input into a pre-trained semantic prediction network to obtain a scene semantic prediction layout. The scene semantic prediction layout and the original semantic layout of the multi-channel semantic layout map are then input into a first loss function for calculation. This first loss function can be the cross-entropy loss function, thus obtaining the semantic consistency loss. The expected pixel depth of each pixel in the 2D latent space feature map is calculated based on the depth probability distribution, specifically:
[0141] ;
[0142] in, This represents the expected pixel depth of pixel (u,v). Indicates the preset depth number of layers. This represents the k-th depth layer in the discretization. This represents the depth probability distribution of pixel (u,v) at the k-th depth layer.
[0143] Calculate the median of the expected pixel depths of all pixels within the same instance mask, and use this calculated median as the expected instance depth. The expected instance depth can be expressed as:
[0144] ;
[0145] in, This represents the expected depth of instance i.
[0146] The actual scale of each instance is obtained by combining the expected depth of the instance with the corresponding pixel scale ratio.
[0147] = ;
[0148] in, This represents the actual scale of instance i.
[0149] The actual scale of all instances and their corresponding dimensionless relative scale constraints are input into the second loss function for calculation, yielding the relative scale consistency loss, specifically:
[0150] ;
[0151] ;
[0152] in, This indicates the loss of relative scale consistency. This represents the scale consistency loss of instance i. Indicates the scale loss weight, This represents the loss function.
[0153] The relative scale consistency loss and semantic consistency loss are input into the total loss function for calculation, resulting in the total loss value. The total loss function is as follows:
[0154] ;
[0155] in, This represents the total loss value. This represents the basic training objective of the generative model or deep prediction network. This represents the semantic consistency loss. The weight coefficients representing the semantic consistency loss. This indicates the loss of relative scale consistency. The weighting coefficient represents the loss of relative scale consistency.
[0156] If the total loss exceeds the first loss threshold, the total loss is backpropagated to the generative model and the depth prediction network using the backpropagation algorithm. This iterative optimization of the generative model and the depth prediction network yields iteratively optimized versions, thus forming a closed loop for 3D scene generation, rendering, loss calculation, and model optimization.
[0157] All systems / appliances used in the methods of the above embodiments of the present invention are within the scope of protection of the present invention.
[0158] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0159] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions.
[0160] It should be noted that any reference numerals placed between parentheses in the claims should not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention can be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims that enumerate several means, several of these means may be embodied by the same hardware. The use of the terms first, second, third, etc., is merely for convenience of expression and does not indicate any order. These terms can be understood as part of the component names.
[0161] Furthermore, it should be noted that in the description of this specification, the terms "one embodiment," "some embodiments," "embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0162] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the claims should be interpreted to include both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0163] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, then this invention should also include these modifications and variations.
Claims
1. A method for generating 3D scenes based on probabilistic scattering with semantic layout constraints, characterized in that, include: A two-dimensional vector plane image is obtained, and the two-dimensional vector plane image is normalized and rasterized to obtain a multi-channel semantic layout image. At the same time, an instance mask for each instance is generated based on the instance connected regions in the multi-channel semantic layout image. The pixel scale ratio of the corresponding instance is calculated based on the instance mask. The instance category, instance mask and pixel scale ratio of all instances are integrated to generate an instance-level layout set. For each instance in the instance-level layout set, construct a corresponding dimensionless relative scale constraint to obtain a relative scale constraint set. The multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set are input into a pre-trained generative model to output a two-dimensional latent space feature map. The two-dimensional latent space feature map is then input into a depth prediction network for depth prediction to output the predicted depth of each pixel in the two-dimensional latent space feature map. A depth prior for each pixel is constructed based on the instance-level layout set and the relative scale constraint set. The depth prior for each pixel is then weighted and fused with the corresponding predicted depth to obtain a depth probability distribution. All pixel coordinates in the multi-channel semantic layout map are transformed into normalized gaze coordinates, and the two-dimensional latent space feature map is scattered into the three-dimensional space along the gaze coordinates in a weighted projection manner according to the depth probability distribution to construct three-dimensional features, and a three-dimensional scene is rendered based on the three-dimensional features.
2. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, The normalization and rasterization of the two-dimensional vector plane map to obtain a multi-channel semantic layout map includes: All vector coordinates in the two-dimensional vector plane diagram are uniformly mapped to the same plane coordinate system and normalized to obtain a normalized two-dimensional vector plane diagram; Obtain the target output resolution, and then rasterize the normalized two-dimensional vector plane image according to the target output resolution to obtain a rasterized two-dimensional vector plane image. Finally, split the rasterized two-dimensional vector plane image into channels according to instance categories to obtain a multi-channel semantic layout image.
3. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, The step of calculating the pixel scale ratio of the corresponding instance based on the instance mask includes: Obtain the height in pixels of the multi-channel semantic layout map and the bounding rectangle of each instance mask. Input the pixel height of each bounding rectangle and the height in pixels into a first formula for calculation to obtain the pixel scale ratio of the corresponding instance. The first formula is: ; ; in, H represents the pixel scale ratio of the i-th instance, and H represents the height in pixels. This represents the pixel height of instance i. This represents the maximum pixel coordinate within the instance mask of instance i. This represents the instance mask of instance i. The x-coordinate of the pixel coordinate. The ordinate represents the pixel coordinate. Represents the minimum pixel coordinates within the instance mask of instance i.
4. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, The step of constructing a dimensionless relative scale constraint for each instance in the instance-level layout set includes: Obtain the bounding rectangle of each instance mask, and input the length of the longer side of each bounding rectangle into the second formula for calculation to construct the dimensionless relative scale constraint corresponding to each instance. The second formula is: ; ; in, This represents the dimensionless relative scale constraint for the i-th instance. This represents the preset tolerance hyperparameter. Indicates the dimensionless relative scale of instance i. This represents the length of the longer side of instance i. This indicates the overall reference length corresponding to the multi-channel semantic layout diagram.
5. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, The step of inputting the multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set into a pre-trained generative model to output a two-dimensional latent space feature map includes: The multi-channel semantic layout map, the instance-level layout set, and the relative scale constraint set are input into a pre-trained generative model. At the same time, a random noise vector is introduced, so that the generative model performs multi-scale convolutional encoding on the multi-channel semantic layout map to generate multi-scale layout features. Simultaneously, based on the instance-level layout set and the relative scale constraint set, each instance is encoded to generate instance features. The random noise vector is mapped to a low-resolution latent variable feature map, wherein the multi-scale layout features include multi-scale layout spatial features and multi-scale layout semantic features. The multi-scale layout features are concatenated and convolved with the low-resolution latent variable feature map. At the same time, the instance features are cross-attentioned with the low-resolution latent variable feature map to achieve iterative updates of the low-resolution latent variable feature map until the termination condition is met, so as to output the iteratively updated low-resolution latent variable feature map, which is a two-dimensional latent space feature map.
6. A method for generating 3D scenes based on probabilistic scattering with semantic layout constraints, characterized in that, The process involves inputting the two-dimensional latent space feature map into a depth prediction network for depth prediction, outputting the predicted depth of each pixel in the two-dimensional latent space feature map, constructing a depth prior for each pixel based on the instance-level layout set and the relative scale constraint set, and then weighting and fusing the depth prior of each pixel with the corresponding predicted depth to obtain a depth probability distribution, including: Obtain a preset number of depth layers and a preset relative depth range in three-dimensional space. Input the preset number of depth layers and the preset relative depth range into a logarithmic interval formula to generate a discrete sequence of relative depths. The discrete sequence of relative depths is a discretized depth sequence expressed in logarithmic intervals. The logarithmic interval formula is as follows: ; ; in, This represents the k-th depth layer in the discretization. This represents the minimum value within the preset relative depth range. This represents the maximum value within a preset relative depth range. Indicates the preset depth number of layers; The depth prediction network uses a multi-scale convolutional structure to enhance the features of the two-dimensional latent space feature map to extract pixel-level spatial features and pixel-level semantic features. Based on the pixel-level spatial features and the pixel-level semantic features, the network outputs the predicted depth of each pixel in each depth layer of the relative depth discrete sequence through a self-attention mechanism. The depth prior of each pixel is constructed based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set. The depth prior of each pixel is then weighted and fused with the corresponding predicted depth to obtain the depth probability distribution.
7. The method for generating a 3D scene based on semantic layout constraints as described in claim 6, characterized in that, The process involves constructing a depth prior for each pixel based on the pixel scale ratio of each instance in the instance-level layout set and the dimensionless relative scale constraint of each instance in the relative scale constraint set. The depth prior for each pixel is then weighted and fused with the corresponding predicted depth to obtain a depth probability distribution, including: The pixel scale ratio of each instance and the corresponding dimensionless relative scale constraint are input into the center formula and variance formula, respectively, to calculate the depth prior center and depth prior variance for each instance. The center formula is as follows: ; in, Denotes the depth prior center of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers; The variance formula is: ; in, Let represent the depth prior variance of the i-th instance. This represents the dimensionless relative scale constraint for the i-th instance. Let represent the maximum value in the dimensionless relative scale constraint for the i-th instance. Let represent the minimum value in the dimensionless relative scale constraint for the i-th instance. This represents the pixel scale ratio of the i-th instance. Hyperparameters representing positive numbers, This represents the variance adjustment coefficient; Based on the depth prior center and depth prior variance of each instance, and combined with the instance mask of each instance in the instance-level layout set, the depth prior center and depth prior variance of each instance are discretized to the relative depth discrete sequence through Gaussian distribution, and the depth prior of each pixel on each depth layer in the relative depth discrete sequence is constructed. The depth prior at each depth layer is weighted and fused with the corresponding predicted depth input using a weighted fusion formula to obtain the depth probability distribution of each pixel at each depth layer. The weighted fusion formula is as follows: ; in, Let represent the depth probability distribution of pixel (u,v) at the k-th depth layer, where u represents the x-coordinate of the pixel. Represents the ordinate of a pixel. Representing depth layer The corresponding index, Indicates the first A normalized function at each depth layer Indicates the first Predicted depth at each depth layer Indicates the first Depth priors at each depth layer This represents the prior weight coefficient.
8. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, The step of converting all pixel coordinates in the multi-channel semantic layout map into normalized gaze coordinates, and scattering the two-dimensional latent space feature map into three-dimensional space along the gaze coordinates using a weighted projection according to the depth probability distribution to construct three-dimensional features includes: Determine whether there are calibrated camera parameters. If they exist, obtain the principal point parameters and equivalent focal length parameters from the camera parameters. If they do not exist, obtain the width and height pixel counts of the multi-channel semantic layout map, calculate the principal point parameters based on the width and height pixel counts, and obtain a preset reference focal length as the equivalent focal length parameters. The principal point parameter and the equivalent focal length parameter are used to transform all pixel coordinates in the multi-channel semantic layout map into normalized line coordinates. Based on the line-of-sight coordinates and the depth layer of the depth probability distribution, a projection sampling point corresponding to each two-dimensional pixel in the two-dimensional latent space feature map is constructed in three-dimensional space. Each two-dimensional pixel in the two-dimensional latent space feature map is weighted according to the depth probability distribution to obtain a weighted feature value for each two-dimensional pixel. The weighted feature value of each two-dimensional pixel is used as the three-dimensional feature value of the corresponding projection sampling point. The three-dimensional feature value of each projection sampling point is accumulated into the three-dimensional voxel grid corresponding to the three-dimensional space through an accumulation formula to construct the three-dimensional feature.
9. The method for generating a 3D scene based on semantic layout constraints as described in claim 1, characterized in that, Also includes: The three-dimensional scene is input into a pre-trained semantic prediction network to obtain the scene semantic prediction layout. At the same time, the original semantic layout of the multi-channel semantic layout map is obtained. The original semantic layout and the scene semantic prediction layout are input into a first loss function for calculation to obtain the semantic consistency loss. The expected pixel depth of each pixel in the two-dimensional latent space feature map is calculated based on the depth probability distribution. The median of the expected pixel depths of all pixels within the same instance mask is calculated and used as the expected instance depth of the instance corresponding to the instance mask. The expected instance depth of each instance is combined with the corresponding pixel scale ratio to obtain the actual scale of each instance. The actual scales of all instances are input into the second loss function along with the corresponding dimensionless relative scale constraint to obtain the relative scale consistency loss. The relative scale consistency loss and the semantic consistency loss are input into the total loss function to calculate the total loss value; If the total loss value exceeds the first loss threshold, the total loss value is backpropagated to the generative model and the deep prediction network through the backpropagation algorithm to achieve iterative optimization of the generative model and the deep prediction network, resulting in an iteratively optimized generative model and an iteratively optimized deep prediction network.