Method and device for generating three-dimensional scene based on planar graph, equipment and medium
Through multimodal feature extraction, modal alignment and autoregressive transformer architecture optimization, the accuracy and efficiency problems of generating three-dimensional scenes from planar images are solved, and high-precision three-dimensional scenes are generated for application in medical training and financial technology virtual scenes.
Patent Information
- Application Number
- CN202510701698.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-09
AI Technical Summary
The existing technology for generating three-dimensional scenes based on planar drawings has problems with accuracy and efficiency, especially in the fields of healthcare and financial technology, where details are lost and the accuracy and efficiency of virtual scenes are insufficient.
Through multimodal feature extraction and preprocessing, modal alignment, and the use of autoregressive transformer architecture for three-dimensional scene prediction, high-precision three-dimensional scenes are generated through loss calculation and parameter optimization iteration.
It achieves high-precision generation of three-dimensional scenes, improves the accuracy and efficiency of virtual scenes, and is suitable for virtual scene applications in medical training and financial technology.
Smart Images

Figure CN120612426A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection technology, and in particular to a method, device, equipment and medium for generating a three-dimensional scene based on a plane image. Background Art
[0002] Generating a 3D scene from a planar image is the process of transforming a 2D planar image into a realistic 3D environment using computer technology, algorithms, models, and tools. For example, using a large model to generate a 3D scene from a planar image involves applying a variety of techniques to transform a 2D image into a realistic 3D environment.
[0003] In the healthcare sector, using large models to generate three-dimensional scenes from floor plans can assist in the construction of virtual medical training environments. For example, using a two-dimensional layout of a hospital department and two-dimensional designs of related medical equipment, large models can generate highly realistic three-dimensional scenes. Medical staff can conduct surgical simulation training in this generated virtual scene, familiarizing themselves with the surgical environment, equipment location, and operating procedures in advance, thereby improving surgical proficiency and accuracy. However, because traditional 3D generation relies on continuous parameter representation, the fine structures of medical equipment such as surgical instruments are lost in the generated three-dimensional scenes, making it difficult to meet the requirements of high-fidelity training.
[0004] In the fintech sector, large models can be used to generate 3D virtual scenes based on 2D floor plans of bank branches for online display and customer guidance. Customers can navigate the virtual branch through an online platform, gaining a pre-visualized understanding of the branch layout, business area distribution, and the location of self-service equipment, enhancing the customer service experience. However, a single image input makes it difficult to clearly define specific business semantics. This results in the generated scene's decoration materials and lighting effects being inconsistent with the actual business location, reducing the accuracy and efficiency of the virtual scene.
[0005] Therefore, there are problems of accuracy and efficiency in generating three-dimensional scenes based on plane maps in the existing technology that need to be solved urgently. Summary of the Invention
[0006] The present invention provides a method, device, computer equipment and medium for generating a three-dimensional scene based on a plane map, so as to solve the problems of accuracy and efficiency in the prior art of generating a three-dimensional scene based on a plane map.
[0007] In a first aspect, a method for generating a three-dimensional scene based on a planar graph is provided, comprising:
[0008] Acquire a data set including planar image and three-dimensional scene data, and perform multimodal feature extraction and preprocessing on the data set to obtain structured data;
[0009] Performing modal alignment on the structured data to obtain fused data in a unified feature space;
[0010] Performing three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result;
[0011] Performing loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model;
[0012] The optimized model is used to construct a preset plan into a three-dimensional scene.
[0013] In a second aspect, a device for generating a three-dimensional scene based on a planar graph is provided, comprising:
[0014] A feature extraction module is used to obtain a data set containing planar image and three-dimensional scene data, perform multimodal feature extraction and preprocessing on the data set to obtain structured data;
[0015] A modality alignment module, configured to perform modality alignment on the structured data to obtain fused data in a unified feature space;
[0016] A scene prediction module, configured to perform three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result;
[0017] an iterative optimization module, configured to perform loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model;
[0018] An image construction module is used to construct a preset plan view into a three-dimensional scene using the optimization model.
[0019] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned method for generating a three-dimensional scene based on a planar graph are implemented.
[0020] In a fourth aspect, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above-mentioned method for generating a three-dimensional scene based on a planar graph are implemented.
[0021] In the above-mentioned scheme implemented by the method, device, computer equipment and storage medium for generating three-dimensional scenes based on planar maps, the data set is converted into structured data through multimodal feature extraction and preprocessing, which can integrate multiple information such as images and text, break the limitations of a single modality, and provide the model with richer and more comprehensive information; the modal alignment operation maps different modal data to a unified feature space, eliminating the semantic gap and data differences between modalities, so that the fused data has consistency and coherence; the autoregressive transformer architecture, with its sequence modeling capability, can gradually generate three-dimensional scenes in the order of spatial structure, and at the same time use the attention mechanism to accurately capture the dependency relationship between the features of each part of the scene; by performing loss calculation and parameter optimization iteration based on the prediction results, the difference between the predicted scene and the real scene can be quantified, driving the adjustment of model parameters to make the generated results more in line with reality; finally, the preset planar map is constructed into a three-dimensional scene using the optimization model, which solves the accuracy and efficiency problems in generating three-dimensional scenes based on planar maps in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0023] Figure 1 This is a schematic diagram of an application environment of a method for generating a three-dimensional scene based on a plane graph in one embodiment of the present invention;
[0024] Figure 2 This is a flow chart of a method for generating a three-dimensional scene based on a planar graph in one embodiment of the present invention;
[0025] Figure 3 This is a flow chart of a specific implementation of step S1;
[0026] Figure 4 This is a flow chart of a specific implementation of step S2;
[0027] Figure 5 This is a flow chart of a specific implementation of step S3;
[0028] Figure 6 This is a schematic structural diagram of a device for generating a three-dimensional scene based on a plane diagram in one embodiment of the present invention;
[0029] Figure 7 is a structural diagram of a computer device in one embodiment of the present invention;
[0030] Figure 8FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0032] The method for generating a three-dimensional scene based on a plane map provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the client communicates with the server through a network. The server can convert the data set into structured data through multimodal feature extraction and preprocessing, and can integrate multiple information such as images and texts, breaking the limitations of a single modality and providing the model with richer and more comprehensive information; the modal alignment operation maps different modal data to a unified feature space, eliminating the semantic gap and data differences between modalities, so that the fused data has consistency and coherence; the autoregressive transformer architecture, with its sequence modeling capability, can gradually generate a three-dimensional scene in accordance with the spatial structure sequence, while using the attention mechanism to accurately capture the dependency relationship between the features of each part of the scene; by performing loss calculation and parameter optimization iteration based on the prediction results, it can quantify the difference between the predicted scene and the real scene, drive the adjustment of the model parameters, and make the generated results more in line with reality; finally, the preset plan view is constructed into a three-dimensional scene using the optimization model, solving the accuracy and efficiency problems in generating three-dimensional scenes based on plan views in the prior art. Among them, the client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0033] See also Figure 2 As shown, Figure 2 A schematic flow chart of a method for generating a three-dimensional scene based on a planar graph provided in an embodiment of the present invention includes the following steps:
[0034] S1. Obtain a data set containing planar map and three-dimensional scene data, perform multimodal feature extraction and preprocessing on the data set to obtain structured data.
[0035] In the example of the present invention, the dataset is a multimodal paired dataset for training and verifying the mapping relationship between two-dimensional plan design drawings and three-dimensional scenes, including strictly aligned two-dimensional design drawing data (such as architectural floor plans, hospital interior layout drawings, industrial design sketches, etc., covering features such as geometric contours, semantic annotations and lighting prompts), three-dimensional scene data (including structural data such as point clouds, mesh models, parametric models, and rendering data such as texture maps, PBR materials, and lighting configurations) and verification data (including three-dimensional scene-related data consistent with the training data format, specifically covering the real category labels, precise three-dimensional coordinates, accurate dimensions and actual postures of scene elements, etc., as a reference standard for evaluating the accuracy of model prediction results.); in terms of data organization, each two-dimensional image is associated with the corresponding three-dimensional scene through a unique identifier (such as a digital code), and may include element-level ID mapping. The element-level ID mapping refers to assigning the same ID number to the same semantic element (such as a table, window, etc.) in the two-dimensional image and the three-dimensional scene to establish a fine-grained correspondence across modalities.
[0036] In the present invention, see Figure 3 As shown, the multimodal feature extraction and preprocessing of the data set to obtain structured data includes:
[0037] S11, tokenizing and representing the three-dimensional scene data in the dataset to obtain a token sequence;
[0038] S12, extracting image features of the planar graph in the data set using a preset image feature extraction network;
[0039] S13: Aggregate the tag sequence and the image features to obtain structured data.
[0040] In the examples of the present invention, the tokenized representation refers to the process of converting the continuous geometric structure, semantic information and spatial relationship in three-dimensional scene data (such as point cloud, mesh, etc.) into an ordered sequence composed of finite discrete tags through an algorithm.
[0041] In an embodiment of the present invention, the tokenizing and characterizing the three-dimensional scene data in the dataset to obtain a token sequence includes:
[0042] Extracting local geometric features from the three-dimensional scene data in the data set;
[0043] Initializing a query vector, mapping the query vector to the same feature space as the local geometric feature, to obtain an alignment query vector;
[0044] Performing a multi-head attention mechanism on the local geometric features and the alignment query vector to obtain a dense embedding;
[0045] The dense embedding is processed by a vector quantization model based on optimal transmission theory to obtain a label sequence.
[0046] In the examples of the present invention, the extraction of local geometric features in the three-dimensional scene data in the data set refers to dividing the three-dimensional scene data using the KNN clustering algorithm and grouping it into multiple local areas based on the distance between points. For each local area, the point cloud data in the area is processed using spherical harmonic function encoding, and the direction information, geometric shape, etc. in the three-dimensional space are converted into mathematical representations to generate key (K) and value (V) vectors with local geometric characteristics, namely local geometric features. For example, in the field of financial technology for real estate value assessment, the three-dimensional point cloud data of the house is divided to identify the local geometric features of different functional areas (such as living rooms and bedrooms), such as space size and wall shape, so as to generate a more accurate house valuation model to assist banks in approving mortgage loan amounts.
[0047] Among them, the local geometric features refer to the feature vectors obtained by analyzing and encoding the local area of the three-dimensional scene data, which are used to characterize the geometric structure and attribute information of the area. The kNN clustering algorithm (k-NearestNeighbors Clustering) is a classification / clustering method based on the nearest neighbor idea. Given a data point, by calculating its distance from other points in the data set (such as Euclidean distance), the nearest k points are selected, and the attributes of the point are inferred based on the category (classification task) or mean (clustering task) of these k points. For example, points with similar distances in space are divided into the same object or area. The spherical harmonic function encoding is that the spherical harmonic function is the solution of the Laplace equation in the spherical coordinate system, which can be used to represent functions on the spherical surface in three-dimensional space (such as direction, color, illumination, etc.). By encoding three-dimensional directions or surface attributes through spherical harmonic coefficients, efficient rotation invariant representation can be achieved.
[0048] In the examples of the present invention, initializing a query vector and mapping the query vector to the same feature space as the local geometric features to obtain an aligned query vector refers to randomly generating an initial query vector (Q) with a suitable dimension (for example, if the hidden layer dimension of the target Transformer model is 512, then a 512-dimensional random vector is generated). Then, a linear projection layer is used to linearly transform the query vector (Q) and map it to the same dimensional space as the local set features to obtain an aligned query vector.
[0049] In the examples of the present invention, the multi-head attention mechanism is used to process the local geometric features and the query vector to obtain a dense embedding. This means that when processing the local geometric features and the query vector, the query vector (Q), key (K), and value (V) are processed by multiple parallel computing units (i.e., "heads"). Each computing unit first performs a linear transformation on the input vector, then calculates the similarity and normalizes it to generate an attention weight. The value vector is then weighted and aggregated based on the weight to obtain the output results of each computing unit. These results are then concatenated and dimensionalized through linear transformation to ultimately form a dense embedding representation that matches the dimension of the input value vector.
[0050] The dense embedding vectors contain both the inherent properties of the data itself (such as the semantics of words and the local structure of geometric features), the implicit relationships between data (such as spatial correlations across regions in a three-dimensional scene), and the aggregation of global contextual information through feature interaction mechanisms (such as multi-head attention) to enhance the integrity of the vectors. For example, in the field of financial technology, when generating a three-dimensional model for a company's factory floor plan, the dense embedding vector can integrate inherent properties such as the factory's structural layout and fire protection facilities, as well as the implicit relationship between the factory and its surrounding environment.
[0051] In the example of the present invention, the dense embedding is processed by a vector quantization model based on optimal transmission theory to obtain a tag sequence, which means taking the dense embedding as input, using a predefined trainable codebook in the vector quantization model, calculating the optimal matching relationship between the embedding vector and each atom in the codebook through optimal transmission theory, finding the nearest neighbor atom of each embedding vector and indexing it as a discrete tag, and finally arranging all the tags in sequence to form a tag sequence of fixed length.
[0052] Among them, the vector quantization model based on optimal transmission theory is a discretization method that combines optimal transmission theory with vector quantization ideas. First, a learnable codebook is defined, which contains multiple atomic vectors as candidates for discrete labels. Then, for each input continuous embedding vector, the optimal transmission algorithm is used to calculate the "optimal transmission path" between it and each atomic vector in the codebook, forming a transmission matrix. Each row of the matrix represents the probability or weight that an input vector should be "assigned" to each atomic vector. Discrete labels are then generated through hard allocation or soft allocation. The hard allocation directly selects the index of the atomic vector with the highest transmission probability as the discrete label of the input vector. The soft allocation uses the probability distribution in the transmission matrix as the weight during training, performs weighted summation on the codebook atoms, and generates a "soft discretized" vector representation. The atomic vectors in the codebook are adjusted through backpropagation to minimize the transmission cost, so that the generated discrete labels can more accurately represent the original continuous vector.
[0053] Specifically, the tag sequence refers to the quantification and serialization of the geometric structure, semantic information, spatial relationship and other contents in the three-dimensional scene data through a pre-defined symbol system (tag set), forming a one-dimensional sequence composed of a finite number of discrete tags arranged in sequence. For example, in the field of medical health, areas such as operating rooms, wards, and pharmacies are converted into sequences using semantic tags (such as "operating room-sterile area-tag A" and "ward-intensive care unit-tag B"), geometric tags (such as area length, width and height size coding) and spatial relationship tags (such as "pharmacy-adjacent-inpatient department-tag C").
[0054] In this example, the preset image feature extraction network is the Vision Transformer (ViT). ViT is a model that applies the Transformer architecture to computer vision tasks. It divides the image into fixed-size patches, linearly projects these patches, adds positional encodings, and then feeds them into the Transformer encoder, thereby capturing global dependencies and structural information in the image.
[0055] In an embodiment of the present invention, the step of extracting the image features of the planar image in the dataset using a preset image feature extraction network includes:
[0056] Preprocessing the two-dimensional plane image in the data set to obtain a preprocessed image;
[0057] Segmenting and embedding the preprocessed image to obtain an embedding vector;
[0058] Performing position encoding and classification on the embedded vector to obtain a feature vector of the two-dimensional plane graph;
[0059] The feature vector is hierarchically processed using a preset encoder to obtain image features of the two-dimensional plane image.
[0060] In the example of the present invention, the preprocessing of the two-dimensional plane image in the data set to obtain a preprocessed image refers to scaling it to the size required by the model (such as 224×224 pixels), and then performing a normalization operation, that is, normalizing the pixel values, dividing the RGB values by 255 and mapping them to the [0,1] interval to obtain the preprocessed image.
[0061] In the example of the present invention, the segmentation and embedding of the preprocessed image to obtain the embedding vector refers to segmenting the preprocessed image into patches of a fixed size (for example, 16×16 pixels), flattening each patch into a one-dimensional vector, and mapping it to the embedding space through a linear projection layer to generate a patch embedding vector.
[0062] In the example of the present invention, the position encoding and classification of the embedding vector to obtain the feature vector of the two-dimensional plane graph refers to adding a learnable positional encoding (Positional Embedding) to each patch embedding vector, adding a special classification token ([CLS] token, i.e., [Class] token) at the front end of the vector, and obtaining the feature vector of the two-dimensional plane graph.
[0063] In this embodiment of the present invention, hierarchical processing of the feature vectors using a preset encoder to obtain the image features of the two-dimensional plane image involves inputting a sequence of feature vectors containing positional codes and [CLS] tokens into an encoder. The encoder is composed of multiple identical layers, each of which sequentially calculates the association weights between vectors using a multi-head self-attention mechanism. The output is then fed into a feedforward neural network for nonlinear transformation. After multiple layers of processing, the output vector corresponding to the [CLS] token is ultimately extracted as the image features of the two-dimensional plane image.
[0064] Specifically, the preset encoder is a Transformer encoder, which is composed of multiple identical layers stacked together, each layer containing a multi-head self-attention mechanism (Multi-Head Self-Attention) and a feedforward neural network (Feed Forward Network). In the multi-head self-attention mechanism, each patch is embedded by calculating three matrices: query, key, and value, to capture the dependencies between different regions in the image and realize the interaction of global information. The feedforward neural network is a nonlinear transformation component that performs two layers of linear transformation on the output of the multi-head self-attention mechanism. The first layer uses a nonlinear activation function to introduce nonlinear characteristics and map the input to a higher-dimensional space; the second layer projects the features back to the original dimension to restore the dimensional consistency of the input; the two layers are connected by residual connections and layer normalization to keep the gradient stable.
[0065] In the examples of the present invention, the aggregating processing of the discrete tag sequence and the image features to obtain structured data refers to mapping the tag sequence and the image features of the two-dimensional plane image to the same feature space through linear projection, using a cross-attention mechanism to establish the association between the tag sequence and the image features, and using a self-attention mechanism to integrate the structural information between the tags; finally, a nonlinear transformation is performed on the fused features through a multi-layer perceptron to output structured data (such as a scene graph or an attribute matrix).
[0066] Among them, the cross-attention mechanism is an attention mechanism that realizes information interaction between sequences from different sources. Taking the labeled sequence as the query vector, the image features as the key vector and the value vector, the attention weight is generated by calculating the similarity between the query and the key and normalizing it. By weighted aggregation of the value vector, the query sequence can obtain the relevant information of the key-value sequence, realizing cross-modal or cross-sequence information fusion and interaction modeling. The self-attention mechanism is to generate query, key, and value vectors for each element in the sequence through linear projection, and generate attention weights by calculating the similarity between the query and the key and normalizing it. The weight reflects the strength of the association between the elements. By weighted aggregation of the value vector, the representation of each element contains the global information of the sequence, thereby capturing the long-range dependency relationship within the sequence. The multi-layer perceptron is a nonlinear transformation module composed of multiple layers of fully connected layers. Its logical process is: the input features are first mapped to a higher-dimensional space by the first fully connected layer and complex interactions are captured by a nonlinear activation function, and then projected to the target dimension by the second fully connected layer to realize feature abstraction and compression. The layers are trained through residual connections and normalization optimization, and finally output the features after nonlinear transformation.
[0067] In the examples of the present invention, by tokenizing and representing three-dimensional scene data, the spatial structure can be converted into an ordered sequence of semantic symbols, retaining geometric and semantic details; the image feature extraction network is used to capture the visual features of the two-dimensional plane (such as edges, textures, and semantic areas), thereby achieving efficient abstraction of plane information; the tag sequence and image features are aggregated to fuse the three-dimensional spatial semantics with the two-dimensional visual clues, construct a structured representation of cross-modal associations, and make up for the information loss of a single modality; this process improves data consistency through standardization, reduces the heterogeneity of multi-source data, provides hierarchical and highly complementary input features for subsequent models, and enhances the model's ability to understand the scene structure; at the same time, the formation of structured data facilitates feature interaction and global modeling, laying a data foundation for generating high-precision three-dimensional scenes.
[0068] S2. Perform modal alignment on the structured data to obtain fused data in a unified feature space.
[0069] In the examples of the present invention, the modal alignment refers to an optimization process that eliminates the heterogeneity of multimodal data (such as three-dimensional label sequences and two-dimensional image features) at the feature representation level through mathematical modeling and machine learning methods, so that it can simultaneously meet the requirements of geometric consistency and semantic equivalence in a unified high-dimensional vector space.
[0070] In the present invention, see Figure 4 As shown, the modality alignment of the structured data to obtain fused data in a unified feature space includes:
[0071] S21, analyzing feature representation differences of different modal data in the structured data;
[0072] S22, calibrating the features of the different modal data according to the feature representation differences to obtain modality-aligned features;
[0073] S23: Fusing the modality-aligned features to obtain fused data in a unified feature space.
[0074] In the examples of the present invention, the analysis of the feature representation differences of different modal data in the structured data refers to extracting the original feature representation (such as text, image or numerical features) from each modal data, and then calculating the distribution differences of the features of each modality (such as Wasserstein distance) or using contrastive learning methods (such as alignment knowledge based on pre-trained contrastive language-image models) to identify the feature differences of the same entity in different modalities. Finally, statistical analysis is used to quantify the semantic consistency and noise interference degree of each feature dimension, and obtain the feature representation differences, that is, the inconsistency of different modalities in semantic expression, dimensional weights and distribution characteristics.
[0075] Among them, the calculation of the distribution difference of each modal feature refers to treating the original feature vector extracted from each modality as a probability distribution, and by constructing an optimal transmission scheme, solving the minimum "transportation cost" required to convert one distribution into another, so as to quantify the degree of difference in feature distribution between modalities. The contrastive language-image model is used to map images and texts to the same high-dimensional space to achieve cross-modal matching. The use of the contrastive learning method is to use the semantic association ability learned by the contrastive language-image model on large-scale image and text data to map the features of different modalities to a shared semantic space, and by constructing positive and negative sample pairs (multimodal features of the same entity are positive samples, and different entities are negative samples), minimize the distance between positive samples and maximize the distance between negative samples, thereby identifying the feature differences of the same entity in different modalities. The statistical analysis quantifies the feature dimensions of each modality. By calculating statistics such as mean, variance, and covariance, the stability and dispersion of the feature dimensions are evaluated. Correlation coefficients are calculated in conjunction with semantic labels to quantify the contribution of each dimension to semantic expression. Furthermore, through methods such as outlier detection and noise density estimation, the degree of noise interference in the feature dimensions is determined. Ultimately, the inconsistencies between different modalities in semantic expression, dimension weights, and distribution characteristics are determined. Quantification is a quantitative description of feature differences.
[0076] In the examples of the present invention, the features of different modalities are calibrated according to the feature representation differences to obtain modality-aligned features, which means that the features of each modality are projected into a shared latent space by constructing a nonlinear mapping function (such as MLP), and then the distribution differences are reduced by minimizing the Wasserstein distance. An attention mechanism is introduced to assign adaptive weights to the feature dimensions, suppress noise and enhance semantically consistent components. Then, the alignment knowledge of the contrastive language-image model is used to constrain the feature distances of the same entity in different modalities through contrastive learning. Then, normalization technology is applied to eliminate statistical differences to stabilize the feature distribution. Finally, the original and aligned features are fused through residual connections to achieve semantic alignment while retaining modality-specific information.
[0077] For example, in the healthcare field, a patient's X-ray planar image and magnetic resonance imaging (MRI) image can be used as different modal data. By constructing a nonlinear mapping function (such as MLP), the features of the X-ray image and the MRI image are projected into a shared latent space. The distribution difference between the two modal data is narrowed by minimizing the Wasserstein distance to ensure feature comparability. An attention mechanism is introduced to focus on the features of the bone lesion area, weakening the interference of irrelevant tissues. With the help of pre-trained model alignment knowledge, comparative learning is used to make the features of the same lesion site in the two modalities closer. Normalization technology eliminates statistical differences caused by different imaging principles and stabilizes the feature distribution. Finally, residual connections are used to fuse the original and aligned features. This not only retains the clear display of bone contours by X-rays and the accurate presentation of soft tissue lesions by MRI, but also achieves cross-modal semantic alignment to assist doctors in determining the extent and scope of bone lesions.
[0078] In this example, the modality-aligned features are fused to generate fused data in a unified feature space. This is achieved by extracting complementary information and suppressing redundancy between modalities through a dynamic weight allocation mechanism. A cross-modal interaction module is then constructed to capture fine-grained semantic associations. A hierarchical feature integration strategy is then used to gradually fuse multi-source information from a local to a global perspective. A randomized training mechanism is then introduced to enhance the model's adaptability to input uncertainty. Finally, a feature optimization path is used to preserve modality specificity and ensure information integrity, resulting in a unified fused data space.
[0079] Specifically, the dynamic weight allocation mechanism adopts a gated fusion mechanism to dynamically calculate the weights of each modal feature through a learnable gated unit (such as a fully connected layer activated by Sigmoid), suppressing redundant information and retaining complementary features. The construction of the cross-modal interaction module utilizes the cross-attention mechanism in the Transformer architecture to enable different modal features to pay attention to each other and capture fine-grained semantic associations, and then designs a hierarchical fusion strategy, first performing splicing and linear transformation at the feature level, and then integrating global context information through a multi-head self-attention mechanism. The feature optimization path is to optimize feature propagation through residual connections and layer normalization to ensure lossless transmission of information during the fusion process and generate a unified feature representation containing multimodal complementary information.
[0080] In the examples of the present invention, modality alignment of structured data and acquisition of fused data can effectively solve the heterogeneity problem of multimodal data caused by different feature representation forms, dimensions and semantic spaces. Analyzing the differences in feature representation can accurately identify the essential differences between modal data and provide a basis for subsequent calibration; based on the difference calibration features, different modal data are adjusted to the same semantic and metric space through mapping, conversion and other operations to achieve feature standardization and compatibility; the calibrated features are fused to form data in a unified feature space, which not only retains the unique information of each modality, but also gives full play to the complementary advantages of multimodal data, so that the data contains richer semantic information and contextual associations, providing higher quality and more comprehensive input for subsequent model training, thereby improving the model's generalization ability, accuracy and ability to handle complex tasks, and avoiding information loss or model performance degradation due to modality inconsistency.
[0081] S3. Use a preset autoregressive transformer architecture to perform three-dimensional scene prediction on the fused data to obtain a three-dimensional scene prediction result.
[0082] In this example, the autoregressive transformer architecture is based on a Transformer-based generative model, employing an encoder-decoder structure. The encoder maps fused data into latent space features, while the decoder sequentially generates scene elements using an autoregressive, masked multi-head attention mechanism, with the output of each time step serving as the conditional input for the next moment. Simultaneously, the spatial structure of the scene is captured by adding three-dimensional position encoding, constructing positive and negative sample pairs and introducing a contrastive loss function to enhance semantic consistency learning. Furthermore, with the aid of a scene constraint module that integrates geometric and semantic constraints, the generated scene is ensured to conform to physical and semantic logic, achieving end-to-end prediction of structured three-dimensional scenes from multimodal fusion information.
[0083] In the present invention, see Figure 5 As shown, the 3D scene prediction is performed on the fused data using a preset autoregressive transformer architecture to obtain a 3D scene prediction result, including:
[0084] S31, converting the fused data into an autoregressive embedding vector;
[0085] S32, using a preset autoregressive transformer architecture to perform three-dimensional scene element prediction on the autoregressive embedding vector to obtain three-dimensional scene element prediction information;
[0086] S33: Construct a three-dimensional scene prediction result based on the three-dimensional scene element prediction information.
[0087] In the example of the present invention, the conversion of the fused data into an autoregressive embedding vector is to perform a linear projection on the fused data, map it to a feature dimension space suitable for autoregressive model processing through a learnable weight matrix, form an initial embedding vector, and add a three-dimensional spatial position code to the initial embedding vector to obtain an autoregressive embedding vector. The three-dimensional spatial position code can reflect the coordinate position relationship of the scene elements in the three-dimensional space (such as the position information corresponding to the x, y, and z axis coordinates). The code is generated by a preset function (such as a sine cosine function). The autoregressive embedding vector contains both the semantic and geometric information of the fused data and the position representation of the three-dimensional spatial structure.
[0088] In an example of the present invention, the use of a preset autoregressive transformer architecture to predict three-dimensional scene elements on the autoregressive embedding vector to obtain three-dimensional scene element prediction information refers to inputting the autoregressive embedding vector into the autoregressive transformer architecture, and using a multi-head self-attention mechanism in the decoder part to sequentially predict the category labels (such as table, chair), geometric parameters (such as three-dimensional coordinates, size, orientation) and semantic relationships (such as adjacency, support) of the three-dimensional scene elements. The prediction output of each time step is used as the conditional input for prediction at the next moment; by constructing positive and negative sample pairs (such as real scene element combinations and random combinations), cross-modal feature alignment is optimized, and a beam search algorithm is used to retain multiple high-probability prediction branches during the decoding process, and structured prediction information containing scene element attributes and association relationships is output.
[0089] For example, in the healthcare sector, multimodal fusion data, such as ward floor plans, medical equipment dimensions, and ergonomic standards, is converted into an embedded vector via linear projection and encoded in the ward's three-dimensional spatial position. This encoded vector is then fed into an autoregressive transformer architecture. The decoder, utilizing a multi-head self-attention mechanism, sequentially predicts the category, placement coordinates, dimensions, and relative positional relationships of three-dimensional scene elements, such as beds, medical cabinets, and monitors. For example, the bed's position is first determined, and then the angle and distance between adjacent medical cabinets are predicted based on this position. By constructing positive and negative sample pairs combining real-world ward layouts with random combinations, the alignment accuracy of features from different modalities is optimized. A beam search algorithm is then used to retain multiple high-probability prediction branches, ultimately outputting structured prediction information containing the attributes and relationships of each element within the ward, ensuring a rational medical equipment layout that meets operational procedures and patient needs.
[0090] In the examples of the present invention, the 3D scene prediction results constructed based on the 3D scene element prediction information are physically plausible based on parameters such as category, 3D coordinates, size, and orientation within the 3D scene element prediction information. A spatial geometry constraint algorithm is used to verify the physical plausibility of collision and occlusion relationships between elements. Simultaneously, a scene topology is constructed based on semantic relationships such as "adjacency," "containment," and "support" to define hierarchical structures and logical relationships. Gridding / voxelization techniques are used to map the calibrated elements to a unified 3D spatial coordinate system, forming a geometric model foundation. Finally, texture mapping and rendering optimization are performed in conjunction with lighting parameters and material properties, resulting in a 3D structured scene that integrates physical constraints, semantic topology, and visual representation, achieving a complete construction from discrete element prediction to spatial semantic modeling.
[0091] In the examples of the present invention, the autoregressive transformer architecture is used to predict three-dimensional scene elements, fully utilizing its sequence modeling capabilities. Combined with a multi-head self-attention mechanism, it captures the complex semantics and spatial associations between elements, enabling accurate inference from local element attributes to global scene structure. Integrating the predicted information into a three-dimensional scene prediction result effectively connects scattered element information to construct a complete, structured scene model, ensuring the consistency and rationality of each element in terms of geometric position and semantic relationships. This process not only efficiently handles the heterogeneity of multimodal fusion data, but also mines the underlying logic of the data through the autoregressive characteristics of the model. The output three-dimensional scene prediction results can be directly applied to visualization, spatial planning, and decision analysis, providing reliable data support and scene references for fields such as architectural design and virtual simulation, thereby improving the practicality and accuracy of scene generation.
[0092] S4. Perform loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model.
[0093] In an embodiment of the present invention, the loss calculation and parameter optimization iteration of the autoregressive transformer architecture are performed according to the three-dimensional scene prediction result to obtain an optimized model, including:
[0094] Calculate the loss value between the three-dimensional scene prediction result and the verification data in the dataset using a preset loss function;
[0095] Calculate the parameter gradient using a preset optimization algorithm according to the loss value;
[0096] Parameters of the autoregressive transformer architecture are iteratively optimized according to the parameter gradients to obtain an optimized model.
[0097] In this example, the use of a preset loss function to calculate the loss between the 3D scene prediction result and the validation data in the dataset is to use a cross-entropy loss function to calculate the category probability distribution error for the scene element's category, 3D coordinates, size, posture, and other prediction information, and a mean square error (MSE) loss function to measure the deviation of numerical parameters. At the same time, a structured loss function is introduced to construct geometric and semantic constraint loss terms by quantifying the differences in the scene's topological structure. Ultimately, these various loss terms are linearly combined according to preset weights to obtain a comprehensive loss value.
[0098] In the examples of the present invention, the parameter gradient is calculated based on the loss value using a preset optimization algorithm. This involves using a backpropagation algorithm. According to the chain rule, the gradient of the loss value output by the loss function with respect to the network output is backpropagated layer by layer to the parameters of each layer of the autoregressive transformer architecture. Automatic differentiation techniques are then used to calculate the partial derivatives of the activation function, weight matrix, and bias term of each layer to obtain the parameter gradient. Simultaneously, in combination with optimization algorithms such as stochastic gradient descent (SGD) and its variants (such as Adam and Adagrad), the gradient is adaptively adjusted through strategies such as learning rate decay and momentum terms, optimizing the update direction and step size, and ultimately obtaining a gradient for iterative parameter optimization.
[0099] In the example of the present invention, the parameters of the autoregressive transformer architecture are iteratively optimized according to the parameter gradient to obtain an optimized model, which adopts optimization algorithms such as stochastic gradient descent (SGD) and its variants (such as Adam, Adagrad), and updates the model parameters (such as attention weights, position encoding parameters) in the opposite direction of the parameter gradient. In each round of iteration, batch training data is input into the model, and the parameters are updated after forward propagation, loss calculation and back propagation. At the same time, the learning rate is dynamically adjusted through a learning rate scheduling strategy (such as cosine annealing), combined with an early stopping mechanism to prevent overfitting, and training is stopped when the loss function converges or the performance of the validation set meets the standard, and finally an optimized model is obtained.
[0100] In this embodiment of the present invention, the accuracy and robustness of the model's three-dimensional scene predictions can be systematically improved by iterating the autoregressive transformer architecture through loss calculation and parameter optimization. For example, in the design of wards in the healthcare field, after using the autoregressive transformer architecture to generate ward three-dimensional scene prediction results, the design accuracy can be significantly improved through loss calculation and parameter optimization. For example, a loss function is used to calculate the difference between the predicted ward layout (bed, medical equipment location, etc.) and the actual medical needs and spatial specifications, quantifying the loss value caused by problems such as insufficient bed spacing and narrow equipment aisles. Based on this loss value, an optimization algorithm is used to calculate the model parameter gradient, locating parameter deviations in aspects such as spatial relationship modeling and equipment size prediction. Subsequently, the model parameters are iteratively optimized based on the parameter gradient, allowing the model to more accurately plan the layout of various elements in the ward in subsequent predictions, such as reasonably adjusting the orientation of the bed to facilitate medical operations and optimizing the placement of medical equipment to improve rescue efficiency. The iteratively optimized model can output a 3D ward scene solution that better fits the medical process and patient needs.
[0101] S5. Utilize the optimization model to construct the preset plan into a three-dimensional scene.
[0102] Among them, the preset floor plan is a two-dimensional graphic that requires three-dimensional scene generation and is used to represent the spatial layout and object position relationship, and can be a design floor plan, etc.
[0103] In an embodiment of the present invention, constructing a preset plan view into a three-dimensional scene using the optimization model includes:
[0104] Extracting image features of the preset plan view using a preset image feature extraction network;
[0105] Performing three-dimensional prediction on the image features using the optimization model to obtain three-dimensional prediction information;
[0106] A three-dimensional scene is constructed according to the three-dimensional prediction information.
[0107] In an example of the present invention, the image features of the preset plan view are extracted using a preset image feature extraction network, which is to divide the preset plan view into image blocks (Patch) of fixed size, and each image block is converted into an embedding vector by linear projection and position encoding is added. The processed embedding vector sequence is then used to capture the global dependency between image blocks through a multi-head self-attention mechanism. After processing by a multi-layer Transformer encoder, the classification tag ([CLS]token) or the feature representation of each image block is extracted. Finally, the extracted features are converted into feature vectors of fixed dimension through feature fusion to form image features that can be input into the optimization model.
[0108] In an example of the present invention, the optimization model is used to perform three-dimensional prediction on the image features to obtain three-dimensional prediction information, which is to input the extracted image features into the optimized autoregressive transformer architecture. The model first maps the image features to the latent space through the encoder, and then the decoder generates three-dimensional scene elements one by one in an autoregressive manner, including object categories, three-dimensional coordinates, sizes, postures and other parameters. The prediction output of each time step is used as the conditional input of the next moment; at the same time, the model uses the scene constraint module to ensure that the generated elements conform to the geometric and semantic logic, and finally outputs the three-dimensional prediction information containing the attributes and association relationships of the scene elements.
[0109] In the example of the present invention, the construction of a three-dimensional scene based on the three-dimensional prediction information is to parse the object category, three-dimensional coordinates, size, posture and other parameters in the three-dimensional prediction information, and create corresponding three-dimensional basic geometric bodies (such as cubes and cylinders) as scene elements based on this; then, adjust the relative positions according to the semantic relationships between the elements (such as adjacency and support) to construct the scene topology structure; then, generate a fine geometric model through triangulation or voxelization processing, and assign corresponding material texture to each element; finally, combine lighting calculation and rendering engine to integrate discrete three-dimensional elements into a complete three-dimensional scene with realism.
[0110] In the examples of the present invention, using an optimization model to construct a preset floor plan into a three-dimensional scene can effectively improve design efficiency and accuracy. For example, in hospital design in the medical and health field, image features such as wall distribution, door and window positions in the preset floor plan of the ward are extracted through an image feature extraction network; these features are then input into the optimized model, and its precise three-dimensional prediction capabilities are used to perform three-dimensional predictions on the number of beds, the placement of medical equipment, the direction of medical passages, etc., to obtain three-dimensional prediction information containing the categories of each element, spatial coordinates, and their relationships; finally, based on the three-dimensional prediction information, a realistic three-dimensional scene of the ward is quickly constructed, intuitively presenting the ward space layout, assisting hospital managers in evaluating the practicality of the ward in advance, and ensuring that the design plan meets the medical process and patient needs.
[0111] It can be seen that in the above scheme, for the target result business, a data set is obtained, and multimodal feature extraction and preprocessing are performed on the data set to obtain structured data; modal alignment is performed on the structured data to obtain fused data in a unified feature space; a preset autoregressive transformer architecture is used to perform three-dimensional scene prediction on the fused data to obtain a three-dimensional scene prediction result; loss calculation and parameter optimization iteration are performed on the autoregressive transformer architecture based on the three-dimensional scene prediction result to obtain an optimization model; the optimization model is used to construct the preset plan view into a three-dimensional scene, which solves the accuracy and efficiency problems of generating three-dimensional scenes based on plan views in the existing technology.
[0112] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0113] In one embodiment, a device for generating a three-dimensional scene based on a plane map is provided. The device for generating a three-dimensional scene based on a plane map corresponds one-to-one to the method for generating a three-dimensional scene based on a plane map in the above embodiment. Figure 6 As shown, the device for generating a three-dimensional scene based on a plane image includes a feature extraction module 101, a modality alignment module 102, a scene prediction module 103, an iterative optimization module 104, and an image construction module 105. The functional modules are described in detail as follows:
[0114] The feature extraction module 101 is used to obtain a data set including planar image and three-dimensional scene data, perform multimodal feature extraction and preprocessing on the data set to obtain structured data;
[0115] A modality alignment module 102 is configured to perform modality alignment on the structured data to obtain fused data in a unified feature space;
[0116] A scene prediction module 103 is configured to perform three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result;
[0117] an iterative optimization module 104 for performing loss calculation and parameter optimization iterations on the autoregressive transformer architecture according to the three-dimensional scene prediction results to obtain an optimized model;
[0118] The image construction module 105 is configured to construct the preset plan view into a three-dimensional scene using the optimization model.
[0119] In one embodiment, the feature extraction module 101 performs multimodal feature extraction and preprocessing on the data set to obtain structured data for:
[0120] Tokenizing the three-dimensional scene data in the dataset to obtain a token sequence;
[0121] Using a preset image feature extraction network to extract image features of the planar graph in the data set;
[0122] Aggregation processing is performed on the tag sequence and the image features to obtain structured data.
[0123] In one embodiment, the feature extraction module 101 tokenizes the three-dimensional scene data in the dataset to obtain a token sequence, specifically for:
[0124] Extracting local geometric features from the three-dimensional scene data in the data set;
[0125] Initializing a query vector, mapping the query vector to the same feature space as the local geometric feature, to obtain an alignment query vector;
[0126] Performing a multi-head attention mechanism on the local geometric features and the alignment query vector to obtain a dense embedding;
[0127] The dense embedding is processed by a vector quantization model based on optimal transmission theory to obtain a label sequence.
[0128] In one embodiment, the modality alignment module 102 performs modality alignment on the structured data to obtain fused data in a unified feature space, which is used to:
[0129] Analyzing differences in feature representations of data of different modalities in the structured data;
[0130] calibrating the features of the different modal data according to the feature representation differences to obtain modality-aligned features;
[0131] The modality-aligned features are fused to obtain fused data in a unified feature space.
[0132] In one embodiment, the scene prediction module 103 performs 3D scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a 3D scene prediction result, which is used to:
[0133] Converting the fused data into an autoregressive embedding vector;
[0134] Using a preset autoregressive transformer architecture to perform three-dimensional scene element prediction on the autoregressive embedding vector to obtain three-dimensional scene element prediction information;
[0135] A three-dimensional scene prediction result is constructed based on the three-dimensional scene element prediction information.
[0136] In one embodiment, the iterative optimization module 104 performs loss calculation and parameter optimization iterations on the autoregressive transformer architecture based on the three-dimensional scene prediction results to obtain an optimization model for:
[0137] Calculate the loss value between the three-dimensional scene prediction result and the verification data in the dataset using a preset loss function;
[0138] Calculate the parameter gradient using a preset optimization algorithm according to the loss value;
[0139] Parameters of the autoregressive transformer architecture are iteratively optimized according to the parameter gradients to obtain an optimized model.
[0140] In one embodiment, the image construction module 105 constructs the preset plan view into a three-dimensional scene using the optimization model, and is used to:
[0141] Extracting image features of the preset plan view using a preset image feature extraction network;
[0142] Performing three-dimensional prediction on the image features using the optimization model to obtain three-dimensional prediction information;
[0143] A three-dimensional scene is constructed according to the three-dimensional prediction information.
[0144] The present invention provides a device for generating a three-dimensional scene based on a plane graph. For a target result business, a data set is acquired, multimodal feature extraction and preprocessing are performed on the data set to obtain structured data; modal alignment is performed on the structured data to obtain fused data in a unified feature space; three-dimensional scene prediction is performed on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result; loss calculation and parameter optimization iteration are performed on the autoregressive transformer architecture based on the three-dimensional scene prediction result to obtain an optimization model; and the optimization model is used to construct the preset plane graph into a three-dimensional scene.
[0145] Regarding the specific definition of the device for generating a three-dimensional scene based on a plane map, please refer to the definition of the method for generating a three-dimensional scene based on a plane map above, which will not be repeated here. The various modules in the above-mentioned device for generating a three-dimensional scene based on a plane map can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0146] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a method for generating a three-dimensional scene based on a plan view.
[0147] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 8As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a method for generating a three-dimensional scene based on a planar map.
[0148] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0149] Acquire a data set including planar image and three-dimensional scene data, and perform multimodal feature extraction and preprocessing on the data set to obtain structured data;
[0150] Performing modal alignment on the structured data to obtain fused data in a unified feature space;
[0151] Performing three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result;
[0152] Performing loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model;
[0153] The optimized model is used to construct a preset plan into a three-dimensional scene.
[0154] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0155] Acquire a data set including planar image and three-dimensional scene data, and perform multimodal feature extraction and preprocessing on the data set to obtain structured data;
[0156] Performing modal alignment on the structured data to obtain fused data in a unified feature space;
[0157] Performing three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result;
[0158] Performing loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model;
[0159] The optimized model is used to construct a preset plan into a three-dimensional scene.
[0160] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0161] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0162] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0163] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.
[0164] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A method for generating a three-dimensional scene based on a planar graph, characterized in that: include: Acquire a data set including planar image and three-dimensional scene data, and perform multimodal feature extraction and preprocessing on the data set to obtain structured data; Performing modal alignment on the structured data to obtain fused data in a unified feature space; Performing three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result; Performing loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model; The optimized model is used to construct a preset plan into a three-dimensional scene.
2. The method for generating a three-dimensional scene based on a planar graph according to claim 1, wherein: The multimodal feature extraction and preprocessing of the data set to obtain structured data includes: Tokenizing the three-dimensional scene data in the dataset to obtain a token sequence; Using a preset image feature extraction network to extract image features of the planar graph in the data set; Aggregation processing is performed on the tag sequence and the image features to obtain structured data.
3. The method for generating a three-dimensional scene based on a planar graph according to claim 2, wherein: The tokenizing and characterizing the three-dimensional scene data in the data set to obtain a token sequence includes: Extracting local geometric features from the three-dimensional scene data in the data set; Initializing a query vector, mapping the query vector to the same feature space as the local geometric feature, to obtain an alignment query vector; Performing a multi-head attention mechanism on the local geometric features and the alignment query vector to obtain a dense embedding; The dense embedding is processed by a vector quantization model based on optimal transmission theory to obtain a label sequence.
4. The method for generating a three-dimensional scene based on a planar graph according to claim 1, wherein: The modality alignment of the structured data to obtain fused data in a unified feature space includes: Analyzing differences in feature representations of data of different modalities in the structured data; calibrating the features of the different modal data according to the feature representation differences to obtain modality-aligned features; The modality-aligned features are fused to obtain fused data in a unified feature space.
5. The method for generating a three-dimensional scene based on a planar graph according to claim 1, wherein: The method of performing three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result includes: Converting the fused data into an autoregressive embedding vector; Using a preset autoregressive transformer architecture to perform three-dimensional scene element prediction on the autoregressive embedding vector to obtain three-dimensional scene element prediction information; A three-dimensional scene prediction result is constructed based on the three-dimensional scene element prediction information.
6. The method for generating a three-dimensional scene based on a planar graph according to claim 1, wherein: The step of performing loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model includes: Calculate the loss value between the three-dimensional scene prediction result and the verification data in the dataset using a preset loss function; Calculate the parameter gradient using a preset optimization algorithm according to the loss value; Parameters of the autoregressive transformer architecture are iteratively optimized according to the parameter gradient to obtain an optimized model.
7. The method for generating a three-dimensional scene based on a planar graph according to claim 1, wherein: The method of constructing a preset plan into a three-dimensional scene using the optimization model includes: Extracting image features of the preset plan view using a preset image feature extraction network; Performing three-dimensional prediction on the image features using the optimization model to obtain three-dimensional prediction information; A three-dimensional scene is constructed according to the three-dimensional prediction information.
8. A device for generating a three-dimensional scene based on a plane map, characterized in that: include: A feature extraction module is used to obtain a data set containing planar image and three-dimensional scene data, perform multimodal feature extraction and preprocessing on the data set to obtain structured data; A modality alignment module, configured to perform modality alignment on the structured data to obtain fused data in a unified feature space; A scene prediction module, configured to perform three-dimensional scene prediction on the fused data using a preset autoregressive transformer architecture to obtain a three-dimensional scene prediction result; an iterative optimization module, configured to perform loss calculation and parameter optimization iteration on the autoregressive transformer architecture according to the three-dimensional scene prediction result to obtain an optimized model; An image construction module is used to construct a preset plan view into a three-dimensional scene using the optimization model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for generating a three-dimensional scene based on a planar graph as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating a three-dimensional scene based on a planar graph as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Task processing method and device based on optimal transmission, equipment and medium
CN121236528A
Production line layout optimization method based on constraint perception Transform and soft permutation matrix
CN122047152A