Ship cabin furniture three-dimensional arrangement automatic generation method
Through the genre-text dual-channel fusion recognition and large language model enhancement technology, combined with knowledge retrieval and specification fusion, efficient automated three-dimensional modeling of ship cabin furniture is achieved, solving the problems of inefficiency and poor compliance in traditional methods, and is suitable for large ship designs with complex structures.
Patent Information
- Application Number
- CN202510775414.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
The three-dimensional modeling of traditional ship cabin furniture is inefficient and highly dependent on the experience of designers. It is difficult to adapt to the large-scale and high-precision cabin layout needs of modern ships, and there is a lack of efficient mechanisms to organically integrate drawing information with design knowledge.
The CAD drawing information is extracted by the primitive-text dual-channel fusion recognition model, combined with the large language model to enhance semantic analysis and knowledge retrieval, and the automatic generation of furniture models is achieved through the joint shape-layout generation network, combined with the technical guidance and layout logic of knowledge retrieval and standardized fusion technology, and the conditional diffusion model is used to ensure the fidelity and diversity of shape generation.
It realizes high-precision and automated three-dimensional layout of cabin furniture, improves modeling efficiency and compliance, reduces the burden of manual design, and is suitable for large-scale ship design tasks with complex structures.
Smart Images

Figure CN120296884A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of ship design and computer-aided design, and particularly relates to a method for automatically generating a three-dimensional layout of ship cabin furniture. Background Art
[0002] As an important part of the internal functional areas of a ship, the three-dimensional modeling and furniture layout of ship living cabins (such as crew bedrooms, kitchens, dining rooms, and office areas) face multiple challenges such as complex spatial structures, diverse drawing expressions, and intensive layout logics. Traditional methods rely on manual identification of CAD drawings, selection of furniture templates from tables, and manual adjustment of model sizes and positions. These methods are not only inefficient but also highly dependent on the experience of designers, prone to errors, and difficult to meet the requirements of large-scale and high-precision cabin layouts in modern ships.
[0003] On the other hand, the layout of cabin furniture is not only a geometric modeling problem but also involves the precise invocation of heterogeneous knowledge such as classification society codes, enterprise standards, and design experience. Currently, there is still a lack of an efficient mechanism to organically integrate drawing information with complex design knowledge, which has become the core difficulty in automated modeling.
[0004] Recently, in the field of three-dimensional scene synthesis, significant progress has been made in the joint layout-shape generation method based on graph structures, and large language models (LLMs) have demonstrated powerful capabilities in graph semantic understanding and knowledge generation. By introducing a large model enhancement mechanism, it is possible to globally understand the graph elements, annotations, and structures in cabin drawings and achieve semantic alignment with design specifications. At the same time, through a knowledge retrieval mechanism, constraint conditions can be dynamically extracted from the ship design rule library to guide the automated generation and reasonable layout of furniture models.
[0005] Therefore, there is an urgent need for an automated method that integrates drawing parsing and large model-enhanced knowledge retrieval to achieve high-precision parsing of cabin drawings, intelligent understanding of design knowledge, and automatic generation of three-dimensional models, breaking through the bottlenecks in the efficiency, intelligence, and compliance of current ship cabin furniture layouts. Summary of the Invention
[0006] Aiming at the problems and deficiencies in the prior art, the purpose of the present invention is to provide a method for automatically generating a three-dimensional layout of ship cabin furniture.
[0007] To achieve the object of the invention, the technical solution adopted by the present invention is as follows: The first aspect of the present invention provides a method for automatically generating a three-dimensional layout of ship cabin furniture, including the following steps: S1: Extract the graph element information from the CAD drawings of the ship cabin, and construct an initial scene graph of the ship cabin according to the graph element information G; The graphic element information includes the boundary information of the ship cabin, graphic element type information, furniture category information, furniture layout position information, furniture geometric information, and annotation information; S2: Perform semantic enhancement processing on the initial scene graph G to obtain an enhanced scene graph , ; ; S3: Input the scene graph into a pre-trained joint shape-layout generation network, which consists of a scene encoder and a scene decoder; the scene graph is first subjected to feature extraction by the scene encoder to obtain the full graph feature vector of the scene graph and output it. The layout decoder decodes the full graph feature vector to obtain the mean layout position , mean rotation angle , layout position variance , rotation angle variance of the furniture object, and splices them to obtain the distribution parameter ; Sample the distribution parameter to obtain the latent layout vector , and use the set of the latent layout vectors to replace the furniture layout position information in the scene graph to obtain the scene graph ; S4: Use the scene encoder of the pre-trained joint shape-layout generation network to perform feature extraction on the scene graph to obtain the full graph feature vector of the scene graph , and input the full graph feature vector into the layout decoder and shape decoder of the scene decoder respectively for decoding processing to obtain the spatial position information and three-dimensional shape information of each furniture object in the scene graph ; Input the spatial position information and three-dimensional shape information of each furniture object output by the scene decoder into the VQ-VAE decoder to reconstruct the three-dimensional layout diagram of the ship cabin furniture.
[0008] According to the above method for automatically generating the three-dimensional layout of ship cabin furniture, preferably, the layout decoder consists of two groups of parallel stacked multi-layer perceptrons. The first group of multi-layer perceptrons is used to predict the three-dimensional bounding box information b of each furniture object; the second group of multi-layer perceptrons is used to predict the rotation angle of each furniture object in the vertical axis direction; the layout position information generated by the layout decoder is , and the symbol represents vector splicing.
[0009] According to the above method for automatically generating the 3D layout of ship cabin furniture, preferably, the shape decoder consists of a conditional diffusion model and a set of multi-layer perceptrons; the multi-layer perceptrons are used to extract the shape latent variable encoding from the full-map feature vector c and output it. The input of the multi-layer perceptrons is the full-map feature vector output by the scene encoder; the conditional diffusion model uses the shape latent variable encoding c output by the multi-layer perceptrons as the guiding condition to generate the true 3D layout-shape encoding of the furniture object in the cabin .
[0010] According to the above method for automatically generating the 3D layout of ship cabin furniture, preferably, the conditional diffusion model is the Latent Diffusion Model.
[0011] More preferably, the Latent diffusion model uses the shape latent variable encoding c output by the multi-layer perceptrons as the guiding condition to generate the true 3D layout-shape encoding of the furniture object in the cabin . The specific operation is as follows: (1) Gradually add random noise to the shape latent variable encoding c to obtain the noise representation at any time step : , where , , is the diffusion scheduling sequence; represents the original signal retention coefficient at each step, which is used to describe the retention ratio of the original signal at the k -th step; represents the added random noise; (2) Starting from the Gaussian noise , use the 3D-UNet network to restore it to obtain the true 3D layout-shape encoding .
[0012] According to the above method for automatically generating the 3D layout of ship cabin furniture, preferably, the scene encoder is the Graph Convolutional Network (GCN).
[0013] According to the above method for automatically generating a three-dimensional layout of ship cabin furniture, preferably, the specific operation of step S1 is as follows: Input the CAD drawing of the ship cabin into a pre-trained primitive-text dual-channel fusion recognition model, use the primitive-text dual-channel fusion recognition model to recognize and extract the primitive information of the ship cabin in the CAD drawing, and then construct an initial scene graph of the ship cabin according to the primitive information by using the primitive-text dual-channel fusion recognition model. G , , where, is the furniture object recognized in the CAD drawing, is the spatial constraint relationship between furniture objects (the spatial constraint relationship includes "close to", "aligned with", "perpendicular to the wall"), and each furniture object and the spatial constraint relationship are embedded as initial feature vectors: , where represents the feature dimension, is the number of furniture objects, M is the number of spatial constraint relationship types.
[0014] According to the above method for automatically generating a three-dimensional layout of ship cabin furniture, preferably, the primitive-text dual-channel fusion recognition model is composed of a text encoder, a visual encoder, a large language model, a hidden layer, and a large model Head prediction layer; among them, the text encoder is used to convert the text information in the CAD drawing of the ship cabin into vector data recognizable by a computer, the input of the text encoder is the text information in the CAD drawing of the ship cabin, and the output is numerical vector data recognizable by a computer; the visual encoder is used to convert the primitive information in the CAD drawing of the ship cabin into vector data recognizable by a computer, the input of the visual encoder is the primitive information in the CAD drawing of the ship cabin, and the output is numerical vector data recognizable by a computer; the numerical vector data recognizable by a computer output by the text encoder and the visual encoder is processed by vector addition to obtain a combined vector; the large language model is used to unify the combined vector into the same semantic space to obtain a semantic vector and output it, the input of the large language model is the combined vector, and the output is the semantic vector; the hidden layer is used to extract features from the semantic vector output by the large language model to obtain an intermediate feature numerical vector and output it, the input of the hidden layer is the semantic vector, and the output is the intermediate feature numerical vector; the large model Head prediction layer is used to complete classification prediction on the intermediate feature numerical vector, the input of the large model Head prediction layer is the intermediate feature numerical vector output by the hidden layer, and the output is the primitive information in the CAD drawing of the ship cabin. More preferably, the large prediction model is the large language model GPT-4o.
[0015] According to the above-mentioned automatic generation method for the three-dimensional layout of ship cabin furniture, preferably, the specific operation of step S2 is as follows: S201: According to the initial scene diagram , retrieve from the standardized knowledge base composed of specification documents related to ship interior design, and obtain the design specifications and empirical constraints related to the cabin type, furniture category, and spatial constraint relationships between furniture objects corresponding to the initial scene diagram , and define the semantic knowledge of each design specification and empirical constraint obtained in the form of a triple, and obtain the triple semantic knowledge of each design specification and empirical constraint , , where represents the subject object , represents the object object, represents the predicate or / and relationship; the triple semantic knowledge of all design specifications and empirical constraints obtained through retrieval is aggregated to obtain the scene standard knowledge set of the initial scene diagram G ;
[0016] S202: Extract text information from each triple semantic knowledge in the scene standard knowledge set, and convert the extracted text information into a text in the form of natural language subject-predicate-object and input it into the visual language model for text encoding processing, and obtain the semantic vector embedding of each triple semantic knowledge ( ); where represents the semantic vector of the overall relationship of the triple semantic knowledge , represents the semantic vector of the subject, represents the semantic vector of the object; S203: Take all the triple semantic knowledge in the scene standard knowledge set as the structured input into the large language model, and at the same time take the initial scene diagram as the structured text input into the large language model, and use the large language model to extract semantic features to obtain the scene-level global vector embedding of the initial scene diagram and output; S204: According to the semantic vector embedding of each triple semantic knowledge obtained in step S202 ( ) and the scene-level global vector embedding obtained in step S203 to perform semantic enhancement on the initial scene diagram to obtain an enhanced scene diagram , , where represents each furniture object and the relationship are embedded as the initial feature vectors, is represented as the furniture layout position information embedded extracted from the CAD drawings, is represented as the semantic vector embedding obtained by encoding the text using a vision - language model, is represented as the semantic vector embedding of the overall relationship obtained by the text encoder of the vision - language model, represents the initial scene graph scene - level global vector embedding, represents the scene graph the set of initial feature vectors of the relationships in, and the symbol refers to vector concatenation.
[0017] More preferably, in step S202, the vision - language model is the CLIP vision - language model; in step S203, the large - language model is GPT - 4o.
[0018] According to the above - mentioned method for automatically generating the three - dimensional layout of ship cabin furniture, preferably, in step S1, the furniture geometric information includes furniture size and furniture shape information.
[0019] According to the above - mentioned method for automatically generating the three - dimensional layout of ship cabin furniture, preferably, the scene encoder is used to extract features from the enhanced scene graph, obtain the full - graph feature vector of the enhanced scene graph and output it; the scene decoder is used to decode the full - graph feature vector output by the scene encoder to generate the spatial position and three - dimensional shape of the furniture objects in the enhanced scene graph in the cabin.
[0020] According to the above - mentioned method for automatically generating the three - dimensional layout of ship cabin furniture, preferably, the training of the joint shape - layout generation network includes the training of the scene encoder and the training of the scene decoder. Among them, the scene encoder is the Graph Convolutional Network (GCN), and the Graph Convolutional Network GCN is trained through the KL - divergence loss for training, and the KL - divergence loss is calculated as follows:
[0021] where is the posterior distribution under the given input; prior distribution; represents the Kullback - Leibler divergence.
[0022] The training objective of the graph neural network GCN is to minimize the posterior distribution and the prior distribution The KL (Kullback-Leibler) divergence between them.
[0023] The training of the scene decoder is as follows: The layout decoder and the shape decoder are trained through the layout loss , shape loss respectively. Among them, The calculation formula of is as follows
[0024] In the formula, is the layout reconstruction loss, is the regularization loss based on the intersection over union, is the hyperparameter that balances the two loss terms. By introducing the regularization loss based on the intersection over union (IoU) in the present invention , regularize The loss avoids layout overlap and improves the rationality of the layout by encouraging the similarity between the predicted layout and the true layout, and solves the problem that when the layout decoder is trained using only the reconstruction loss , the layout decoder can usually only handle simple scenes, and the spatial layout prediction for complex scenes is not realistic enough.
[0025] The calculation formula of is as follows:
[0026] In the formula, is the layout position information of the object, is the th three-dimensional bounding box of the object, is the rotation angle, is the number of divisions of the rotation space; The first term of the layout reconstruction loss is the three-dimensional bounding box regression loss, and the second term is the rotation classification loss.
[0027] The calculation formula of is as follows: .
[0028] Shape loss The calculation formula of is as follows:
[0029] In the formula, t is a time step, which is sampled from the set of discrete values ; represents the true noise in the diffusion process, denotes the noise predicted by the denoising model (such as a neural network) according to the current state , time step t and three-dimensional layout-shape guidance conditions , and denotes the noise signal after diffusion in the th step.
[0030] Finally, the overall training loss of the joint shape-layout generation network is calculated as follows:
[0031] In the formula, , and are the weight coefficients of the , , losses respectively.
[0032] The present invention adopts this joint training method so that the model can simultaneously learn a reasonably structured layout and a realistic object shape, thereby generating a more realistic 3D scene.
[0033] In a second aspect of the present invention, there is provided a computer device. Preferably, the computer device includes: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the three-dimensional layout automatic generation method of ship cabin furniture as described in the first aspect above.
[0034] In a third aspect of the present invention, there is provided a computer-readable storage medium. Preferably, the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the three-dimensional layout automatic generation method of ship cabin furniture as described in the first aspect above.
[0035] Compared with the prior art, the positive and beneficial effects achieved by the present invention are as follows: (1) Based on the primitives and annotation contents in the CAD drawings, the present invention uses a primitive-text dual-channel fusion recognition model to extract key parameters such as the type, size, and orientation of furniture in the CAD drawings, and performs dynamic matching and reconstruction with the parametric furniture model library to realize the automatic generation and flexible configuration of furniture models. Moreover, the present invention introduces a large language model (LLM) into the primitive-text dual-channel fusion recognition model to enhance the drawing semantic parsing and knowledge invocation capabilities, and generates a hierarchical semantic representation of the cabin scene by virtue of its global understanding ability of the graph structure and text content. Therefore, through the primitive-text dual-channel fusion recognition and the automatic extraction technology of cabin boundaries, a high-precision cabin scene graph can be constructed.
[0036] (2) The present invention combines knowledge retrieval and specification fusion technologies to dynamically obtain constraint rules and design logics from a vast amount of design specifications and enterprise standards, which are used to guide furniture type selection, layout logic judgment, and spatial relationship constraints, significantly improving the compliance and rationality of the automatic layout results.
[0037] (3) Through the dual-channel fusion recognition of graphic elements and text and the automatic extraction technology of cabin boundaries, the present invention can construct a high-precision cabin scene graph; relying on the knowledge retrieval and specification extraction mechanism, it automatically invokes classification society standards, industry regulations, and experience rules to form structured layout constraints; with the help of CLIP and large language models, semantic enhancement at the node level, edge level, and global level is performed on the scene graph, comprehensively improving the integrity and context relevance of the graph structure expression; on this basis, the joint generation of the three-dimensional furniture position and shape is realized through the joint shape-layout generation network, and the conditional diffusion model and VQ-VAE structure are introduced to ensure the fidelity and diversity of shape generation.
[0038] (4) The method for automatically generating the three-dimensional layout of furniture in ship cabins of the present invention innovatively integrates CAD drawing semantic parsing, graph structure modeling, design specification knowledge retrieval, and large language model enhancement mechanisms, constructing a closed-loop linkage and automated modeling process of "drawing - knowledge - model" linkage, effectively breaking through the technical bottlenecks in traditional furniture modeling processes, such as heavy dependence on design experience, low efficiency of specification invocation, and poor layout compliance. It has strong adaptability and high scalability, and is particularly suitable for large ship design tasks with a large number of cabins, complex structures, and inconsistent drawing specifications.
[0039] (5) The present invention significantly improves the automation level, modeling accuracy, compliance reliability of the cabin furniture layout modeling, reduces the manual design burden, is applicable to various engineering scenarios such as new ship design, cabin renovation, digital twin system construction, and virtual simulation, and has good technical versatility, system intelligence, and industrial application prospects. Description of the Drawings
[0040] Figure 1 It is the flowchart of semantic knowledge enhancement of the scene graph based on the large language model and visual language model in the present invention; Figure 2 It is the prompt template for generating scene-level global vector embedding using the large language model in the present invention in the present invention; Figure 3 It is the schematic diagram of the architecture of the joint shape-layout generation network in the present invention; Figure 4 It is the schematic diagram of the network architecture of the dual-channel fusion recognition model of graphic elements and text in the present invention; Figure 5The 3D layout rendering of ship cabin furniture generated according to the 3D layout automatic generation method of ship cabin furniture described in Embodiment 1 of the present invention based on the CAD drawings of the ship cabin. Detailed implementation manners
[0041] Embodiment 1:
[0042] A 3D layout automatic generation method for ship cabin furniture, the specific steps are as follows: S1: Input the CAD drawings of the ship cabin into a pre-trained dual-channel fusion recognition model of graphic elements and text, and use the dual-channel fusion recognition model of graphic elements and text to recognize and extract the graphic element information in the CAD drawings; the graphic element information includes the boundary information of the ship cabin, graphic element type information, furniture category information, furniture layout position information, furniture geometric information and annotation information; among them, the furniture layout position information is expressed as , is the 3D bounding box information of the furniture, is the furniture orientation angle, is the embedded representation of the furniture layout position information in the CAD drawings by the dual-channel fusion recognition model of graphic elements and text, represents the feature dimension, is the number of furniture objects; the furniture geometric information includes furniture size and furniture shape information; S2: According to the ship cabin boundary information, furniture category, furniture layout position information and furniture geometric information extracted by the dual-channel fusion recognition model of graphic elements and text, use the dual-channel fusion recognition model of graphic elements and text to construct the initial scene graph of the ship cabin , , where is the furniture object recognized in the CAD drawings, is the spatial constraint relationship between furniture objects (the spatial constraint relationship includes "close to", "aligned", "perpendicular to the wall"), and each furniture object and the spatial constraint relationship are embedded as the initial feature vectors: , where represents the feature dimension, is the number of furniture objects, M is the number of spatial constraint relationship types; S3: According to the initial scene graph , retrieve from the standardized knowledge base composed of specification documents related to ship interior design (as Figure 1 shown), and obtain the information related to the initial scene graph Design specifications and empirical constraints related to the spatial constraint relationships among corresponding cabin types, furniture categories, and furniture objects, and define the semantic knowledge of each design specification and empirical constraint obtained in the form of a triple, obtaining the triple semantic knowledge of each design specification and empirical constraint , , where represents the subject object (starting node), such as the furniture object chair in the scene graph; represents the object object (target node), such as the furniture object table in the scene graph; represents the predicate or / and relationship (edge), for example: next to...; The triple semantic knowledge of all design specifications and empirical constraints obtained through retrieval set, obtaining the initial scene graph scene standard knowledge set of , , M is the number of spatial constraint relationship types; S4: Extract text information from each triple semantic knowledge in the scene standard knowledge set and convert the extracted text information into a natural language subject-predicate-object form of text input into the CLIP vision-language model for text encoding processing, obtaining the semantic vector embedding of each triple semantic knowledge ( ), where represents the semantic vector of the overall relationship of the triple semantic knowledge , represents the semantic vector of the subject, represents the semantic vector of the object; S5: Take all the triple semantic knowledge in the scene standard knowledge set as structured input into the large language model GPT-4o, and at the same time take the initial scene graph as structured text input into the large language model, and use the large language model GPT-4o for semantic feature extraction to generate the scene-level global vector embedding of the initial scene graph and output, represents the feature dimension; among them, use the large language model GPT-4o for semantic feature extraction to generate the scene-level global vector embedding of the initial scene graph Figure 2 shown; S6: According to the semantic vector embedding of each triple semantic knowledge obtained in step S4 and the scene-level global vector embedding obtained in step S5 For the initial scene graph perform semantic enhancement to obtain an enhanced scene graph , , where represents each furniture object and relationship are embedded as initial feature vectors represents the furniture layout position information embedding extracted from the CAD drawing represents the object semantic vector embedding obtained by encoding text using the CLIP vision-language model represents the semantic vector embedding of the overall relationship of the object obtained by the text encoder of the CLIP vision-language model represents the scene-level global vector embedding of the initial scene graph represents the set of initial feature vectors of the relationships in the scene graph , and the symbol refers to vector concatenation; S7: Input the enhanced scene graph into a pre-trained joint shape-layout generation network (as shown in Figure 3 ); The joint shape-layout generation network consists of a scene encoder and a scene decoder. Among them, the scene encoder is used to extract features from the enhanced scene graph, obtain and output the full graph feature vector of the enhanced scene graph, and the scene decoder is used to decode the full graph feature vector output by the scene encoder to generate the spatial position and three-dimensional shape of the furniture objects in the cabin in the enhanced scene graph; The scene decoder consists of a layout decoder and a shape decoder; The layout decoder is used to decode the input full graph feature vector to generate the layout position information of each furniture object in the cabin; The shape decoder is used to decode the input full graph feature vector to generate the three-dimensional shape of each furniture object in the cabin; The enhanced scene graph input into the joint shape-layout generation network is first subjected to feature extraction by the scene encoder to obtain the full graph feature vector of the enhanced scene graph and output. The layout decoder decodes the full graph feature vector output by the scene encoder to obtain the layout position mean , rotation angle mean , layout position variance , rotation angle variance , and after concatenation, the distribution parameter is obtained, and then the latent layout vector is sampled from the distribution parameter through the reparameterization trick , the set of potential layout vectors is , , , denoting the set of potential layout vectors; using to replace the enhanced scene graph in , to obtain the scene graph , ; S8: Using the scene encoder of the pre-trained joint shape-layout generation network to extract features from the scene graph to obtain the full-image feature vector of the scene graph , and inputting the full-image feature vector into the layout decoder and shape decoder of the scene decoder respectively for decoding processing to obtain the spatial position information and three-dimensional shape information of each furniture object in the scene graph ; Inputting the spatial position information and three-dimensional shape information of each furniture object output by the scene decoder into the pre-trained VQ-VAE decoder to reconstruct the three-dimensional layout drawing of the ship cabin furniture.
[0043] Among them, the graphic-character dual-channel fusion recognition model (as shown in Figure 4 ) consists of a text encoder, a visual encoder, the large language model GPT-4o, a hidden layer, and a large model Head prediction layer; among them, the text encoder is used to convert the text information in the ship cabin CAD drawing into vector data recognizable by a computer, the input of the text encoder is the text information in the ship cabin CAD drawing, and the output is the numerical vector recognizable by a computer; the visual encoder is used to convert the graphic information in the ship cabin CAD drawing into vector data recognizable by a computer, the input of the visual encoder is the graphic information in the ship cabin CAD drawing, and the output is the numerical vector recognizable by a computer; the numerical vector data recognizable by the computer output by the text encoder and the visual encoder are processed by vector addition to obtain a combined vector; the large language model GPT-4o is used to unify the combined vector into the same semantic space to obtain and output a semantic vector, the input of the large language model GPT-4o is the combined vector, and the output is the semantic vector; the hidden layer is used to extract features from the semantic vector output by the large language model GPT-4o to obtain and output an intermediate feature numerical vector, the input of the hidden layer is the semantic vector, and the output is the intermediate feature numerical vector; the large model Head prediction layer is used to complete classification prediction on the intermediate feature numerical vector, the input of the large model Head prediction layer is the intermediate feature numerical vector output by the hidden layer, and the output is the category information corresponding to the circle, point, arc, and line in the ship cabin CAD drawing.
[0044] The layout decoder consists of two groups of parallel stacked multi-layer perceptrons (MLPs). The first group of MLPs is used to predict the three-dimensional bounding box information of each furniture object. b , b , Among them, represents the central coordinates of the furniture object, represent the width, length, and height of the furniture object respectively; the second group of MLPs is used to predict the rotation angle of each furniture object in the vertical axis direction. ; The layout position information generated by the layout decoder is , and the symbol represents vector concatenation.
[0045] The shape decoder consists of a conditional diffusion model and a group of multi-layer perceptrons (MLPs) (denoted as ); The multi-layer perceptron is used to extract the shape latent variable encoding c from the full-image feature vector and output it. The input of the multi-layer perceptron is the full-image feature vector output by the scene encoder; The conditional diffusion model uses the shape latent variable encoding c output by the multi-layer perceptron as the guiding condition to generate the true three-dimensional layout-shape encoding of the furniture object in the cabin.
[0046] Preferably, the conditional diffusion model is the Latent Diffusion Model. The specific operation of the Latent diffusion model using the shape latent variable encoding c output by the multi-layer perceptron as the guiding condition to generate the true three-dimensional layout-shape encoding of the furniture object in the cabin is (refer to the paper: "High-Resolution Image Synthesis with Latent Diffusion Models"): (1) Gradually add random noise to the shape latent variable encoding c to obtain the noise representation at any time step :
[0047] Among them, , , is the diffusion scheduling sequence; represents the original signal retention coefficient at each step, which is used to describe the retention ratio of the original signal in the k step; represents the added random noise; (2) Starting from Gaussian noise Start with using a 3D-UNet network for restoration to obtain the true three-dimensional layout - shape encoding .
[0048] According to the above method for automatically generating the three-dimensional layout of ship cabin furniture, preferably, the training of the joint shape-layout generation network includes the training of the scene encoder and the scene decoder. Among them, the scene encoder is a Graph Convolutional Network (GCN). The graph convolutional network GCN is trained through the KL divergence loss and the calculation formula of the KL divergence loss is as follows:
[0049] where, is the posterior distribution under the given input; is the prior distribution; represents the Kullback-Leibler divergence.
[0050] The training objective of the graph neural network GCN is to minimize the KL (Kullback-Leibler) divergence between the posterior distribution and the prior distribution . Among them, the prior distribution and the posterior distribution are both well-known professional terms in the field; among them, the prior distribution represents the prior knowledge of the parameters before there is data and is obtained through real indoor scene data; the posterior distribution can be understood as the distribution of the predicted shape and layout given the observed data.
[0051] The training of the scene decoder is as follows: The layout decoder and the shape decoder are respectively trained through the layout loss and the shape loss . Among them, The calculation formula of is as follows In the formula, is the layout reconstruction loss, is the regularization loss based on the intersection over union, is the hyperparameter for balancing the two loss terms. In the present invention, by introducing the regularization loss based on the intersection over union (IoU), the regularization loss encourages the similarity between the predicted layout and the true layout, avoids layout overlap, improves the rationality of the layout, and solves the defect that when the layout decoder is trained using only the reconstruction loss , the layout decoder can usually only handle simple scenes and the spatial layout prediction for complex scenes is not realistic enough.
[0052] The calculation formula is as follows:
[0053] In the formula, is the layout position information of the object, is the th three-dimensional bounding box of the object, is the rotation angle, is the number of divisions of the rotation space; the first term of the layout reconstruction loss is the three-dimensional bounding box regression loss, and the second term is the rotation classification loss.
[0054] The calculation formula of .
[0055] Shape loss The calculation formula is as follows:
[0056] In the formula, is a time step, which is sampled from the discrete value set ; represents the true noise in the diffusion process, represents the noise predicted by the denoising model (such as a neural network) according to the current state , time step t and the three-dimensional layout-shape guidance condition , represents the noise signal after diffusion at the th step.
[0057] Finally, the overall training loss of the joint shape-layout generation network is calculated as follows:
[0058] In the formula, , and are the weight coefficients of , , losses respectively.
[0059] The three-dimensional layout rendering of the ship cabin furniture reconstructed according to the three-dimensional layout automatic generation method of the ship cabin furniture described in the above Embodiment 1 is as shown in Figure 5 . Figure 5 The upper and lower figures in Figure 5It can be seen that the automatic generation method for three-dimensional layout of ship cabin furniture according to the present invention can achieve the automatic generation of the layout of ship cabin furniture, with high modeling accuracy and high fidelity.
[0060] Example 2: A computer device, comprising: a processor and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the automatic generation method for three-dimensional layout of ship cabin furniture as described in the above-mentioned Example 1.
[0061] Example 3: A computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the automatic generation method for three-dimensional layout of ship cabin furniture as described in the above-mentioned Example 1.
Claims
1. A method for automatically generating a three-dimensional layout of ship cabin furniture, characterized in that, Including the following steps: S1: Extract primitive information from the CAD drawings of the ship's cabin, and construct an initial scene graph of the ship's cabin according to the primitive information. G The primitive information includes boundary information, primitive type information, furniture category information, furniture layout position information, furniture geometry information, and annotation information of the ship's cabin. S2: Perform semantic enhancement processing on the initial scene graph G to obtain an enhanced scene graph , ; ; S3: Scene graph Input is the pre-trained joint shape-layout generation network, which consists of a scene encoder and a scene decoder; the scene graph First, the scene encoder is used to extract features and obtain the scene graph The full-image feature vector And output, the layout decoder for the full image feature vector Decode and get the layout position mean of furniture objects , mean rotation angle , Layout Position Variance , rotation angle variance , splicing to get the distribution parameters ; Sample the distribution parameters to obtain a latent layout vector , and use the set of latent layout vectors to replace the furniture layout position information in the scene graph to obtain the scene graph ; ; S4: Scene encoder using a pre-trained joint shape-layout generation network to generate scene graphs Perform feature extraction to obtain the scene graph The full-image feature vector , the whole image feature vector Input into the layout decoder and shape decoder of the scene decoder respectively, decode and process to obtain the scene graph The spatial position information and three-dimensional shape information of each furniture object in the scene decoder are input into the VQ-VAE decoder to reconstruct the three-dimensional layout of the ship cabin furniture.
2. The three-dimensional layout automatic generation method for ship cabin furniture according to claim 1, characterized in that The layout decoder consists of two groups of multi-layer perceptrons stacked in parallel. The first group of multi-layer perceptrons is used to predict the three-dimensional bounding box information of each furniture object b ; The second group of multi-layer perceptrons is used to predict the rotation angle of each furniture object in the vertical axis direction ; The layout position information generated by the layout decoder is , the symbol represents vector concatenation.
3. The three-dimensional layout automatic generation method for ship cabin furniture according to claim 1, characterized in that The shape decoder consists of a conditional diffusion model and a set of multi-layer perceptrons; the multi-layer perceptrons are used to extract the shape latent variable encoding from the full-image feature vector c and output, and the input of the multi-layer perceptrons is the full-image feature vector output by the scene encoder; the conditional diffusion model uses the shape latent variable encoding c output by the multi-layer perceptrons as the guiding condition to generate the true 3D layout-shape encoding of the furniture object in the cabin .
4. The three-dimensional layout automatic generation method for ship cabin furniture according to claim 3, characterized in that The conditional diffusion model is the Latent Diffusion Model; the diffusion model uses the shape latent variable encoded by the multi-layer perceptron output c as the guiding condition to generate the true three-dimensional layout-shape encoding of the furniture object in the cabin The specific operation is as follows: (1)Gradually add random noise to the shape latent variable encoding c to obtain the noise representation at any time step : , Among them, , , is the diffusion scheduling sequence; represents the original signal retention coefficient at each step, used to describe the retention ratio of the original signal in the k step; represents the added random noise; (2) Starting from Gaussian noise use the 3D-UNet network to restore and obtain the true three-dimensional layout-shape encoding .
5. The three-dimensional layout automatic generation method for ship cabin furniture according to claim 1, characterized in that The specific operation of step S1 is as follows: input the CAD drawing of the ship's cabin into a pre-trained primitive-character dual-channel fusion recognition model, use the primitive-character dual-channel fusion recognition model to recognize and extract the primitive information of the ship's cabin in the CAD drawing, and then construct an initial scene graph of the ship's cabin using the primitive-character dual-channel fusion recognition model according to the primitive information G。 6. The automatic generation method for three-dimensional layout of ship cabin furniture according to claim 5, characterized in that The primitive-text dual-channel fusion recognition model consists of a text encoder, a visual encoder, a large language model, a hidden layer, and a large model Head prediction layer; among them, the text encoder is used to convert the text information in the ship cabin CAD drawing into vector data recognizable by a computer. The input of the text encoder is the text information in the ship cabin CAD drawing, and the output is numerical vector data recognizable by a computer; the visual encoder is used to convert the primitive information in the ship cabin CAD drawing into vector data recognizable by a computer. The input of the visual encoder is the primitive information in the ship cabin CAD drawing, and the output is numerical vector data recognizable by a computer; the numerical vector data recognizable by the computer output by the text encoder and the visual encoder is processed by vector addition to obtain a combined vector; the large language model is used to unify the combined vector into the same semantic space to obtain and output a semantic vector. The input of the large language model is the combined vector, and the output is the semantic vector; the hidden layer is used to extract features from the semantic vector output by the large language model to obtain and output an intermediate feature numerical vector. The input of the hidden layer is the semantic vector, and the output is the intermediate feature numerical vector; the large model Head prediction layer is used to complete classification prediction on the intermediate feature numerical vector. The input of the large model Head prediction layer is the intermediate feature numerical vector output by the hidden layer, and the output is the primitive information in the ship cabin CAD drawing.
7. The method for automatically generating the three-dimensional layout of ship cabin furniture according to claim 1, wherein The specific operation of step S2 is: S201: According to the initial scene diagram , retrieve from the standardized knowledge base composed of specification documents related to ship interior design, and obtain the design specifications and empirical constraints related to the cabin type, furniture category, and spatial constraint relationship between furniture objects corresponding to the initial scene diagram . Define the semantic knowledge of each design specification and empirical constraint obtained in the form of triples, and obtain the triple semantic knowledge of each design specification and empirical constraint , , where represents the subject object , represents the object object, represents the predicate or / and relationship; Aggregate the triple semantic knowledge of all design specifications and empirical constraints obtained through retrieval to obtain the scene standard knowledge set of the initial scene diagram G ; ; S202: Extract text information from each triple semantic knowledge in the scene standard knowledge set and convert the extracted text information into a text in the form of natural language subject-predicate-object, and input the text into the vision-language model for text encoding processing to obtain the semantic vector embedding ( ) of each triple semantic knowledge; where represents the semantic vector of the overall relationship of the triple semantic knowledge , represents the semantic vector of the subject, and represents the semantic vector of the object; S203: Take all the triple semantic knowledge in the scene standard knowledge set as the structured input to the large language model. At the same time, take the initial scene graph as the structured text input to the large language model, and use the large language model to extract semantic features to obtain the scene-level global vector embedding of the initial scene graph and output it; S204: According to the semantic vectors of each triple semantic knowledge obtained in step S202 and the scene-level global vector embedding obtained in step S203 , perform semantic enhancement on the initial scene graph to obtain an enhanced scene graph , , where , represents that each furniture object and the relationship are embedded as initial feature vectors, represents the embedding of the furniture layout position information extracted from the CAD drawing, represents the semantic vector embedding obtained by text encoding using a vision-language model, represents the semantic vector embedding of the overall relationship obtained by using a vision-language model for text encoder, represents the scene-level global vector embedding of the initial scene graph , represents the set of initial feature vectors of the relationships in the scene graph , and the symbol refers to vector concatenation.
8. The three-dimensional layout automatic generation method of ship cabin furniture according to claim 7, characterized in that The visual language model is the CLIP visual language model; the large language model is GPT-4o.
9. A computer device, characterized in that, Including: A processor and a memory communicatively connected to the processor; The memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method for automatically generating a three-dimensional layout of ship cabin furniture according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method for automatically generating a three-dimensional layout of ship cabin furniture according to any one of claims 1 to 8.
Citation Information
Patent Citations
Multi-modal scene generation method based on relation and style perception
CN117496025A
Scene graph generation method, system and equipment based on open vocabulary and medium
CN119649381A
Automatic 3D modeling method, system and device based on 2D ship CAD drawings
CN119783261A
Vector-quantized transformable bottleneck networks
US20230316454A1
Contextual augmentation using scene graphs
WO2021263018A1
Cited By
Large language model-based pension building function topological relation aging-suitable generation method
CN121479883A
Three-dimensional indoor scene layout generation method and system based on scene graph control
CN121564225A