A method, apparatus, electronic device and storage medium for generating three-dimensional twin scenes

CN122574199APending Publication Date: 2026-08-14NINGBO TELIAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]有鉴于此,本申请实施例提供了一种三维孪生场景生成方法、装置、电子设备及存储介质,以解决现有技术中,三维孪生场景生成存在空间幻觉,导致生成效果差的问题

Benefits of technology

[0009]本申请实施例与现有技术相比存在的有益效果是:本申请实施例中的方法通过获取计算机辅助设计源数据、计算机辅助设计图像数据以及文本指令;对计算机辅助设计源数据进行解析,提取计算机辅助设计源数据的线段特征,对计算机辅助设计图像数据进行视觉特征提取,生成图像特征,对文本指令进行编码,生成文本特征;将线段特征、图像特征和文本特征输入预先训练的孪生多模态大模型,孪生多模态大模型用于基于线段特征、图像特征和文本特征的空间位置关联关系,对线段特征、图像特征和文本特征进行融合,得到结构化孪生数据;其中,结构化孪生数据包括:用于指示对象类型的标识位、中心点坐标、尺寸参数;解析结构化孪生数据,依据中心点坐标和尺寸参数确定空间位置基准,并根据标识位和空间位置基准,动态生成对应的三维实体模型,以渲染生成三维孪生场景。本申请通过提取CAD源数据中的线段特征,并基于空间位置关联关系将多模态数据进行对齐融合,为孪生多模态大模型的生成推理提供了精确的底层几何坐标约束。这确保了后续生成的三维实体模型能够与原始CAD图纸在空间位置和尺寸上实现精确对齐,有效避免了空间幻觉。且本申请通过线段特征、图像特征和文本特征之间的相互补充与交叉印证,大幅增强了系统对复杂输入数据的容错能力和整体解析准确度。本申请通过孪生多模态大模型输出包含标识位、中心点坐标、尺寸参数的结构化孪生数据,将场景解耦为轻量化的参数集合。这不仅大幅降低了数据处理的算力开销,还使得最终生成的三维场景支持基于参数修改的无缝二次编辑。本申请在场景渲染阶段,并非直接生成三维实体模型,而是预先依据中心点坐标和尺寸参数确立空间位置基准(即物理占位边界)。随后再根据该基准和标识位动态生成对应的三维实体模型,为模型的实例化提供了严格的物理边界约束,从根本上规避了模型在三维空间中组装时发生的干涉和穿模错位,保障了孪生场景的渲染质量,进而避免了现有技术中,三维孪生场景生成存在空间幻觉,导致生成效果差的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574199A_ABST
    Figure CN122574199A_ABST
Patent Text Reader

Abstract

This application relates to the field of large-scale model technology, and provides a method, apparatus, electronic device, and storage medium for generating three-dimensional twin scenes. The method acquires computer-aided design (CAD) source data, CAD image data, and text instructions; it parses the CAD source data, extracts line segment features, extracts visual features from the CAD image data to generate image features, and encodes the text instructions to generate text features; it inputs the line segment features, image features, and text features into a twin multimodal large-scale model, which, based on the spatial positional relationships of the line segment features, image features, and text features, obtains structured twin data; it parses the structured twin data, determines the spatial position reference based on the center point coordinates and size parameters, and dynamically generates the corresponding three-dimensional solid model based on the identifier and spatial position reference to render and generate a three-dimensional twin scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of large model technology, and in particular to a method, apparatus, electronic device and storage medium for generating three-dimensional twin scenes. Background Technology

[0002] With the rapid development of digital twin, virtual reality, and metaverse technologies, the rapid construction of high-fidelity 3D twin scenes has become a core fundamental requirement in fields such as architectural design, intelligent manufacturing, and smart cities.

[0003] In recent years, with the development of artificial intelligence technology, large-model-based 3D content generation technology (AIGC 3D) has gradually emerged. However, those skilled in the art have found in practical engineering applications that existing 3D twin scene generation methods suffer from severe "spatial illusions." Existing generative large models have inherent strong randomness. Although they can generate visually harmonious scenes, they cannot guarantee that the generated results will achieve a precise and verifiable one-to-one correspondence with the original design drawings (such as CAD engineering drawings) in geometry and space. For example, the model can understand the semantics of "there should be a door here," but often the position drifts, failing to ensure that the object is placed on the precise coordinates specified in the drawing. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for generating three-dimensional twin scenes, in order to solve the problem that spatial illusion exists in the generation of three-dimensional twin scenes in the prior art, resulting in poor generation effects.

[0005] A first aspect of this application provides a method for generating a three-dimensional twin scene. The method includes: acquiring computer-aided design source data, computer-aided design image data, and text instructions; parsing the computer-aided design source data to extract line segment features, extracting visual features from the computer-aided design image data to generate image features, and encoding the text instructions to generate text features; inputting the line segment features, image features, and text features into a pre-trained twin multimodal large model, whereby the twin multimodal large model is used to fuse the line segment features, image features, and text features based on their spatial positional relationships to obtain structured twin data; wherein the structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters; parsing the structured twin data, determining a spatial position reference based on the center point coordinates and size parameters, and dynamically generating a corresponding three-dimensional entity model based on the identifier bit and spatial position reference to render and generate a three-dimensional twin scene.

[0006] A second aspect of this application provides a three-dimensional twin scene generation apparatus, comprising: an acquisition module for acquiring computer-aided design source data, computer-aided design image data, and text instructions; a feature module for parsing the computer-aided design source data, extracting line segment features from the computer-aided design source data, extracting visual features from the computer-aided design image data to generate image features, and encoding the text instructions to generate text features; a fusion module for inputting the line segment features, image features, and text features into a pre-trained twin multimodal large model, wherein the twin multimodal large model is used to fuse the line segment features, image features, and text features based on the spatial positional relationship between the line segment features, image features, and text features to obtain structured twin data; wherein the structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters; and a generation module for parsing the structured twin data, determining a spatial position reference based on the center point coordinates and size parameters, and dynamically generating a corresponding three-dimensional entity model based on the identifier bit and spatial position reference to render and generate a three-dimensional twin scene.

[0007] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.

[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.

[0009] The beneficial effects of this application embodiment compared with the prior art are as follows: The method in this application embodiment acquires computer-aided design source data, computer-aided design image data, and text instructions; it parses the computer-aided design source data to extract line segment features, extracts visual features from the computer-aided design image data to generate image features, and encodes the text instructions to generate text features; it inputs the line segment features, image features, and text features into a pre-trained twin multimodal large model, which is used to fuse the line segment features, image features, and text features based on the spatial positional relationship of the line segment features, image features, and text features to obtain structured twin data; wherein, the structured twin data includes: an identifier bit for indicating the object type, center point coordinates, and size parameters; it parses the structured twin data, determines the spatial position reference based on the center point coordinates and size parameters, and dynamically generates the corresponding three-dimensional entity model based on the identifier bit and spatial position reference to render and generate a three-dimensional twin scene. This application extracts line segment features from CAD source data and aligns and fuses multimodal data based on spatial positional relationships, providing precise underlying geometric coordinate constraints for the generation and inference of twin multimodal large models. This ensures that the subsequently generated 3D solid model can be accurately aligned with the original CAD drawings in terms of spatial position and size, effectively avoiding spatial illusions. Furthermore, this application significantly enhances the system's fault tolerance and overall parsing accuracy for complex input data through the mutual complementarity and cross-verification of line segment features, image features, and text features. This application outputs structured twin data containing identifiers, center point coordinates, and size parameters through the twin multimodal large model, decoupling the scene into a lightweight parameter set. This not only significantly reduces the computational cost of data processing but also enables seamless secondary editing of the final generated 3D scene based on parameter modifications. In the scene rendering stage, this application does not directly generate 3D solid models but pre-establishes spatial positional benchmarks (i.e., physical occupancy boundaries) based on center point coordinates and size parameters. Then, based on the benchmark and the identifier, the corresponding 3D entity model is dynamically generated, which provides strict physical boundary constraints for the instantiation of the model. This fundamentally avoids interference and clipping misalignment that occur when the model is assembled in 3D space, ensuring the rendering quality of the twin scene. In turn, it avoids the problem of spatial illusion in the generation of 3D twin scenes in the existing technology, which leads to poor generation effect. Attached Figure Description

[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart illustrating a method for generating a three-dimensional twin scene provided in an embodiment of this application; Figure 2 This is a schematic diagram of the overall architecture logic of a three-dimensional twin scene generation method provided in an embodiment of this application; Figure 3 This is a schematic diagram of the basic structure of a twin multimodal large model provided in an embodiment of this application; Figure 4 This is a schematic diagram of the structure of structured twin data for a twin multimodal large model provided in an embodiment of this application; Figure 5 This is a schematic diagram of another method for generating a three-dimensional twin scene provided in the embodiments of this application; Figure 6 This is a schematic diagram of the structure of a three-dimensional twin scene generation device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0012] In the following description, specific details such as particular device structures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known devices, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] The following will describe in detail, with reference to the accompanying drawings, a method and apparatus for generating a three-dimensional twin scene according to an embodiment of this application.

[0014] Figure 1 This application provides a method for generating a three-dimensional twin scene, such as... Figure 1 As shown, the method includes: S101. Obtain computer-aided design source data, computer-aided design image data, and text instructions; The methods provided in this application can be executed by a terminal, a server, or a combination of both. In other words, the technical solutions of this application can be deployed on a single device or run in a distributed manner. Terminals include, but are not limited to, smartphones, tablets, personal computers (PCs), or IoT devices; servers include, but are not limited to, independent physical servers, server clusters, or cloud servers providing cloud computing services. For ease of explanation, subsequent embodiments will use a server as the execution entity.

[0015] In this embodiment, the server first acquires multimodal input data. This includes computer-aided design (CAD) source data (such as vector engineering drawings) containing precise underlying geometric primitive information; computer-aided design image data, which are two-dimensional images visually corresponding to the CAD source data, providing an intuitive global topology view; and text commands containing prior control information input by the user through natural language or preset scripts. These text commands include, but are not limited to: scene style descriptions, global generation intentions, model attribute settings (such as specifying specific materials or default dimensions), layer semantic mapping rules (such as specifying the entity type corresponding to a specific layer), and local editing constraints for specific areas. These three elements together constitute the basic materials for generating the 3D scene.

[0016] S102. Parse the computer-aided design source data, extract the line segment features of the computer-aided design source data, extract visual features from the computer-aided design image data, generate image features, and encode the text instructions to generate text features. The server employs a multi-track feature extraction mechanism to process data of different modalities. Specifically, for CAD source data, the server first parses the underlying structure of the CAD source data (e.g., reading the ENTITIES primitive segments in the DXF file) and extracts the parameters of each basic geometric primitive in the drawing (e.g., start coordinates, end coordinates, line type, layer, etc.). Subsequently, the server vectorizes these discrete geometric parameters using a feature extraction algorithm to generate line segment features that can represent absolute geometric constraints in space.

[0017] For CAD image data, considering that pure vector data lacks a macroscopic representation of the overall spatial layout, the server performs depth feature extraction on the rendered 2D image corresponding to the CAD source data. Through a visual feature extraction network, the model can "view" the spatial topology of the entire drawing with a global receptive field, transforming pixel-level spatial distribution into high-dimensional image features. This feature provides intuitive visual context and strong fault tolerance for the subsequent generation process.

[0018] For text instructions, the server uses a text encoder (such as a tokenizer and word embedding model) from the field of natural language processing to segment and map the acquired text instructions into a high-dimensional space. This step transforms the discrete natural language description input by the user into a continuous semantic vector, thereby generating text features that contain prior control information (such as style and attribute settings).

[0019] S103. Input the line segment features, image features, and text features into the pre-trained twin multimodal large model. The twin multimodal large model is used to fuse the line segment features, image features, and text features based on the spatial positional relationship of the line segment features, image features, and text features to obtain structured twin data. The structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters. In this embodiment, the server inputs the three types of features extracted through multi-track processing into a pre-trained twin multimodal large model for joint inference.

[0020] First, the twin multimodal large model performs cross-modal spatial alignment and feature fusion. Since line segment features contain absolute physical coordinates, while image features contain corresponding pixel spatial distributions, there is an objective mapping relationship between the two. The large model utilizes this spatial positional association (e.g., by constructing a unified 2D / 3D spatial positional encoding mechanism) to deeply bind the geometric constraints of line segments with the same spatial position to visual local features; simultaneously, it combines prior semantic instructions from text features (such as "generate a load-bearing wall here") to complete the multimodal fusion of image, text, and line segments in the feature space. This fusion mechanism takes into account both global semantics and establishes insurmountable underlying geometric constraints.

[0021] Subsequently, the twin multimodal large model performs parameterized decoding and output based on the fused multimodal features. Unlike traditional generative models that directly output uneditable 3D meshes or point clouds, the twin multimodal large model in this application is trained to output highly semantic parameter sequences. Through its internal inference network, the twin multimodal large model parses the attributes of each independent entity within the target scene from the fused features, thereby outputting structured twin data.

[0022] Structured twin data is essentially a highly lightweight parameter dictionary that specifies at least three core pieces of information: first, an identifier to indicate the type of the generated object (e.g., identifying whether the object is a wall, door, window, or a specific piece of equipment based on the identifier); second, center point coordinates to provide the object's coordinate anchor point in three-dimensional virtual space; and third, size parameters (such as length and width) to define the object's spatial boundaries and geometric contours in three-dimensional space.

[0023] It is understandable that this twin multimodal large model is obtained by pre-training or fine-tuning based on massive amounts of real business data.

[0024] For example, during the training phase of the twin multimodal large model, the server pre-constructs a training set containing a large number of paired samples. The input of each sample consists of historical line segment features (obtained based on historical CAD source data), image features (obtained based on historical CAD image data), and text features (obtained based on historical text instructions); while the corresponding supervised label data (Ground Truth) is a set of structured entity parameters with absolute spatial accuracy (i.e., real identifiers, real center point coordinates, and real dimension parameters) that has been manually calibrated or exported from a real BIM (Building Information Modeling) system.

[0025] During training, the network parameters of the Siamese multimodal large model are iteratively optimized using a joint loss function. Specifically, for discrete semantic classification tasks such as "identifier bits," the Siamese multimodal large model can use classification losses (such as cross-entropy loss) to constrain its output to correctly identify the object type. For continuous spatial values ​​such as "center point coordinates" and "size parameters," the Siamese multimodal large model uses regression losses (such as L1 loss, mean squared error (MSE) loss, or spatial bounding box (IoU) loss) to severely penalize the spatial deviation between the model's output coordinates and the true physical coordinates. Through this end-to-end training with multi-task joint constraints, the Siamese multimodal large model converges and acquires the ability to accurately infer high-semantic parameter sequences with engineering-grade absolute geometric accuracy from multimodal inputs.

[0026] S104. Analyze the structured twin data, determine the spatial position reference based on the center point coordinates and size parameters, and dynamically generate the corresponding three-dimensional solid model based on the identifier and spatial position reference to render and generate a three-dimensional twin scene.

[0027] In this embodiment, after the twin multimodal large model outputs structured twin data, the server's scene rendering module (e.g., an assembly module built based on WebGL, Unity, Unreal or other self-developed 3D graphics engines) will automatically parse the structured twin data.

[0028] First, the server establishes physical space constraints. The rendering engine reads the "center point coordinates" and "size parameters" from the structured twin data and calculates the spatial footprint of the object to be generated in a preset 3D coordinate system. Specifically, the engine uses the center point coordinates as the base reference and extends outwards based on dimensions such as length and width to define a "spatial position reference" (essentially equivalent to a virtual space boundary or anchoring area) with absolute scale and position in the 3D coordinate system. This step ensures that the generated 3D solid model has strict physical boundaries and will not experience arbitrary coordinate drift.

[0029] Secondly, the server performs parametric entity generation and placement. It reads the "identifier bits" from the structured twin data and matches them with the corresponding object type (e.g., "load-bearing wall," "glass door," or "specific mechanical equipment") from a pre-defined model library or parametric construction script. Subsequently, the server invokes the corresponding basic geometry construction logic, strictly constraining the generated geometry within the established "spatial location benchmark." In this way, the 3D solid model is dynamically "created" and precisely positioned at the absolute coordinates specified in the drawing.

[0030] Finally, as all entity parameters in the structured twin data sequence are parsed and instantiated one by one, massive 3D entity models are precisely assembled and combined in virtual 3D space. The rendering engine performs final calculations on the lighting and materials of the global model, thereby rendering and outputting a complete 3D twin scene that is geometrically precisely aligned 1:1 with the input CAD source data and highly conforms to the text semantic control expectations.

[0031] According to the solution provided in this application, computer-aided design (CAD) source data, CAD image data, and text instructions are acquired. The CAD source data is parsed to extract line segment features. Visual features are extracted from the CAD image data to generate image features. Text instructions are encoded to generate text features. The line segment features, image features, and text features are input into a pre-trained twin multimodal large model. The twin multimodal large model is used to fuse the line segment features, image features, and text features based on their spatial positional relationships to obtain structured twin data. The structured twin data includes: an identifier indicating the object type, center point coordinates, and size parameters. The structured twin data is parsed, and a spatial positional reference is determined based on the center point coordinates and size parameters. Based on the identifier and spatial positional reference, a corresponding 3D solid model is dynamically generated to render a 3D twin scene. This application provides precise underlying geometric coordinate constraints for the generation and inference of the twin multimodal large model by extracting line segment features from the CAD source data and aligning and fusing the multimodal data based on spatial positional relationships. This ensures that the subsequently generated 3D solid model can be precisely aligned with the original CAD drawings in terms of spatial position and size, effectively avoiding spatial illusions. Furthermore, this application significantly enhances the system's fault tolerance and overall parsing accuracy by complementing and cross-verifying line segment features, image features, and text features. This application outputs structured twin data containing identifiers, center point coordinates, and size parameters through a twin multimodal large model, decoupling the scene into a lightweight set of parameters. This not only significantly reduces the computational cost of data processing but also enables seamless secondary editing of the final generated 3D scene based on parameter modifications. In the scene rendering stage, this application does not directly generate 3D solid models but pre-establishes a spatial position benchmark (i.e., physical occupancy boundary) based on the center point coordinates and size parameters. Then, it dynamically generates the corresponding 3D solid model based on this benchmark and identifiers, providing strict physical boundary constraints for model instantiation. This fundamentally avoids interference and clipping misalignment that occur during model assembly in 3D space, ensuring the rendering quality of the twin scene and thus avoiding the spatial illusions and poor generation results found in existing technologies.

[0032] In some examples, the source data of computer-aided design is parsed to extract the line segment features of the source data. This includes: extracting the feature information of line segments in the source data of computer-aided design through a parallel Transformer encoder and residual network architecture to generate line segment features. The feature information includes: line segment type, feature data of the line segment itself, and spatial relationship between line segments.

[0033] Specifically, CAD parsing involves reading and decoding the precise coordinate data and layer attributes of geometric primitives (especially line segments) stored in open format files such as DXF. DXF files are essentially structured text databases that record all entity information constituting the drawing. By parsing specific ENTITIES primitive segments layer by layer, the start and end coordinates of each line segment, as well as its layer and line type, can be precisely extracted. This transforms the visual content of the drawing into a lossless and quantitative set of pure vector data that can be understood and processed by a computer. The fundamental purpose of this is to establish a unique and precise spatial reference for subsequent pixel-level accurate generation, essentially providing an insurmountable geometric constraint for the generation process. This ensures that every object unit in the final generated twin scene is strictly aligned with the original design, laying the data foundation for achieving a 1:1 high-fidelity conversion.

[0034] After obtaining the basic line segment vectors mentioned above, the server needs to perform in-depth feature extraction. In traditional CAD analysis tasks, a single convolutional neural network (CNN) is limited by its local receptive field, making it difficult to capture cross-regional primitive relationships; while a simple sequence model is prone to losing the underlying geometric accuracy. To address this, this embodiment introduces a parallel dual-track feature extraction architecture of "Transformer encoder + Residual Network (ResNet)" to achieve perfect decoupling and fusion of local and global features.

[0035] On one hand, the server inputs the basic line segment vectors into the Residual Network (ResNet) branch. ResNet, with its deep convolutional capabilities and residual connection mechanism, focuses on extracting "line segment type" (such as solid, dashed, or dotted lines) and "feature data of the line segment itself" (such as line length, local curvature, thickness, and layer color attributes). This branch ensures that the absolute physical accuracy of the underlying geometric constraints is not lost during feature reduction.

[0036] On the other hand, the server inputs the same sequence of line segments in parallel into the Transformer encoder branch. The Transformer encoder, utilizing its core self-attention mechanism, can directly calculate the attention weights between any two line segments without considering the absolute physical distance on the drawing paper. In this way, the branch efficiently extracts the spatial relationships between line segments (e.g., parallel, perpendicular, intersecting, collinear, or the topological dependency between the line segment containing a "door" and the line segment containing a "wall").

[0037] Finally, the server performs dimensional alignment and concatenation, or feature addition, between the local self-feature data extracted by ResNet and the global spatial relationship features extracted by the Transformer encoder, thereby generating high-dimensional line segment features. This parallel architecture design ensures that the final generated line segment features possess both microscopic self-properties and macroscopic spatial relationships, laying an extremely solid data foundation for subsequent complex 3D spatial reasoning in multimodal large models.

[0038] In some examples, visual feature extraction is performed on computer-aided design image data to generate image features, including: extracting visual features from computer-aided design image data using a visual Transformer architecture to obtain line segment information in the computer-aided design image data to generate image features; wherein, the line segment information includes at least: line segment start coordinates, line segment end coordinates, line segment type, and line segment width.

[0039] It's understandable that relying solely on the underlying vector data in CAD source data has inherent limitations. Real industrial or architectural design drawings are often riddled with numerous non-standard drawing issues, such as incompletely closed wall lines (broken lines), overlapping line segments, and a large number of redundant annotation symbols. In this "dirty data" situation, simple vector analysis can easily get stuck in local logical dead ends, preventing the large model from correctly deducing the complete spatial topology (e.g., failing to recognize that an unclosed area is actually a room).

[0040] To address this technical pain point of low fault tolerance, this application introduces a visual extraction branch as a powerful supplement. The server pre-renders the CAD source data into a two-dimensional raster image (i.e., computer-aided design image data) and inputs it into the Vision Transformer (ViT) architecture. Unlike traditional convolutional neural networks, the ViT architecture first divides the entire drawing image into multiple fixed-size image patches, and then uses a global self-attention mechanism to serialize these image patches.

[0041] Through this global perspective, the ViT architecture can extract global topological semantics directly from the macroscopic distribution of pixels, much like a human designer looking at blueprints. Even if the underlying vector lines break at some point, ViT can still keenly capture macroscopic line information (including visual line start coordinates, end coordinates, line type, and intuitive line width) through pixel-level visual coherence.

[0042] Ultimately, the image features generated by this visual Transformer provide a highly robust global visual context for the large twin multimodal model. This forms a perfectly complementary dual-track mechanism with the extremely precise but relatively fragile vector line segment features described in the previous embodiments: the vector branch provides absolute geometric dimensions, while the visual branch provides macroscopic topological correction. This multimodal cross-verification enables the large model to generate accurate 3D twin parameters even when faced with highly irregular sketches.

[0043] In some examples, text instructions are encoded to generate text features, including: semantic segmentation and high-dimensional word embedding mapping of text instructions using a pre-trained natural language encoding model, extracting attribute settings, entity constraints and scene style semantics contained in the text instructions to generate continuous high-dimensional vectors as text features.

[0044] Specifically, in the construction of digital twin scenes, although CAD vector data and image data can perfectly lock the physical coordinates and geometric topology of the 3D model, they cannot express the designer's higher-level intentions (such as material replacement, local fine-tuning, or overall stylization). In order to achieve zero-code parametric intervention in the 3D scene without modifying the underlying drawings, this embodiment introduces a powerful text encoding branch.

[0045] When the server receives discrete natural language instructions from the user (e.g., "Replace the exterior walls of all meeting rooms with transparent glass and set the default height to 3 meters", "Generate a modern industrial-style office area"), it first uses a tokenizer to segment the long text sequence into fine-grained semantic tokens. Then, using a pre-trained text encoder (such as the BERT architecture or the embedding layer of a large language model), these tokens are mapped to a continuous high-dimensional semantic vector space.

[0046] During this process, the text encoder is able to keenly capture and extract the core control elements in the instructions, namely: specific "attribute settings" (such as a height of 3 meters), local "entity constraints" (such as the exterior wall of the conference room), and global "scene style semantics" (such as modern industrial style).

[0047] The resulting textual features will serve as a powerful prior control constraint, injected into the graphical and vector features during the subsequent large-scale model fusion stage (S103). This design breaks away from the rigid logic of traditional 3D modeling, which relies solely on rote copying of images, and endows the large model with extremely high flexibility and interactivity. Users only need to input natural language to precisely fine-tune the materials, dimensions, and style of the twin scene at the semantic level, based on the geometric hard constraints.

[0048] In some examples, the spatial position reference is determined based on the center point coordinates and size parameters, including: extracting the center point coordinates and using them as the coordinate anchor point of the 3D solid model in the preset 3D coordinate system; parsing the length and width in the size parameters, and using the coordinate anchor point as the origin to calculate the spatial boundary coordinates of the 3D solid model in the 3D coordinate system; and combining the coordinate anchor point with the spatial boundary coordinates to establish the spatial position reference when generating the 3D solid model.

[0049] Specifically, in the underlying logic of computer graphics and 3D rendering engines (such as WebGL, Unity, UE, etc.), it is impossible to accurately define the volume and footprint of a 3D object based solely on an isolated center point. Without strict boundary constraints, during automated batch assembly of models, it is highly likely that adjacent solid models will overlap and collide (i.e., "clipping") or become disproportionate. To provide absolute physical constraints for the dynamic generation in the backend, this embodiment pre-constructs a rigorous spatial bounding box calculation logic before rendering and assembly.

[0050] First, the server extracts the center point coordinates from the structured twin data and projects them directly into a preset virtual 3D coordinate system (such as the world coordinate system or a specific local coordinate system) as the coordinate anchor point of the 3D entity model. This coordinate anchor point is equivalent to an unshakeable absolute position stake for the entity in the virtual world.

[0051] Subsequently, the server parses the size parameters (i.e., length and width values) in the parameter dictionary. Using the established coordinate anchor point as the geometric origin, and combining the length and width values ​​of the entity, mathematical calculations are performed on the coordinate axes around it to obtain the spatial boundary coordinates of the 3D entity model on the 3D coordinate system (especially the target ground plane) (e.g., the coordinate sequence of the four corner points of the target occupant rectangle).

[0052] Finally, the server combines and binds the coordinate anchor points representing the absolute position with the spatial boundary coordinates representing the occupied area. This combination essentially defines an invisible "occupying area" (i.e., a spatial location reference) in three-dimensional space with precise length and width boundaries. When a specific entity model (e.g., a load-bearing wall 3 meters long and 0.2 meters wide) is subsequently dynamically generated based on the identifier, the geometric mesh of that entity will be strictly constrained, or even automatically aligned, within this pre-calculated reference framework.

[0053] This application successfully and rigorously transforms the one-dimensional numerical parameters output by the twin multimodal large model into the three-dimensional spatial geometric constraints (spatial position reference) required by the graphics rendering engine through the above-described method. By establishing its physical boundaries and occupancy range in advance before the twin multimodal large model is formally instantiated, this solution fundamentally eliminates the mutual interference and spatial disorder problems that are prone to occur during the automated assembly of massive three-dimensional entities, giving AIGC-generated content true engineering-grade (not just visual-grade) high fidelity.

[0054] In some examples, corresponding 3D solid models are dynamically generated based on identifier bits and spatial location references to render a 3D twin scene. This includes: determining the object type of the 3D solid model to be generated based on the identifier bits; dynamically constructing a 3D solid model with a corresponding 3D contour based on the object type and size parameters; performing spatial pose transformation on the 3D solid model based on the spatial location reference to anchor the bottom center point of the 3D solid model to the 3D coordinates corresponding to the spatial location reference, thus completing the spatial positioning of the 3D solid model; and combining and rendering multiple 3D solid models that have completed spatial positioning to generate a 3D twin scene.

[0055] Specifically, in traditional graphics rendering engines (such as Unity, Unreal Engine, or self-developed engines based on WebGL), the default local coordinate origin of a 3D entity is usually located at its geometric center. This means that if a wall with a height of 3 meters is to be placed on the target ground plane (e.g., an elevation plane with Z=0), the rendering engine must perform an additional Z-axis (height direction) offset determinant calculation for the model (i.e., shift upwards by 1.5 meters). When assembling massive industrial twin scenes with hundreds of thousands of primitives, this height offset calculation for each individual entity not only consumes a great deal of computing resources, but also easily leads to rendering errors such as the model being half-submerged underground (clipping) or floating in the air if the height parameter is slightly adjusted.

[0056] To address the aforementioned issues of spatial illusion and assembly misalignment, this embodiment employs a bottom center point anchoring mechanism during the dynamic generation and rendering placement phases.

[0057] First, the rendering system parses the "identifier bits" and matches the corresponding "object type" (such as walls, columns, doors and windows) in the asset library or parametric generation script. Based on the extracted "size parameters" (length, width), it dynamically stretches or generates a solid geometric mesh with the corresponding three-dimensional contour.

[0058] Subsequently, during the critical spatial pose transformation stage, this application redefines or aligns the origin of the local coordinate system of the 3D solid model to the bottom center point of its bounding box. When performing the spatial translation matrix, the bottom center point is directly anchored to the 3D coordinates of a pre-established spatial position reference (e.g., directly aligned with the elevation plane of the target floor).

[0059] By employing the steps described above, this application eliminates the need for offset calculations of the model's Z-axis height at the underlying mathematical logic level. Regardless of how the height of the solid model is parametrically changed by text commands (e.g., instantly stretching the wall height from 3 meters to 5 meters), because its bottom anchor point is firmly fixed to the spatial reference, the model will only naturally grow upwards and will never penetrate the ground plane downwards.

[0060] Finally, based on this anchoring mechanism that requires no bias calculation and has absolute physical boundary constraints, thousands of independent 3D solid models are rapidly and accurately combined in virtual space. The rendering engine then adds lighting and materials to them, efficiently rendering a 1:1 seamless, high-fidelity 3D twin scene without clipping artifacts.

[0061] In some examples, structured twin data also includes: the rotation angle of the 3D solid model determined based on the world coordinate system; and the material type that specifies the surface rendering properties of the 3D solid model.

[0062] Specifically, in complex 3D space, there exist entities whose orientation cannot be fully determined solely by the center point coordinates and size parameters (e.g., whether a 3-meter-long wall is oriented east-west or north-south). Therefore, in this embodiment, the twin multimodal large model further decodes and outputs the rotation angle when outputting structured parameters. In particular, this rotation angle is strictly limited to an absolute yaw angle or Euler angle determined based on the world coordinate system. The advantage of this setting is that it avoids the matrix multiplication calculation errors caused by using local coordinate systems in complex hierarchical nested models, ensuring that the engine can accurately rotate the model to a position completely parallel to the CAD base drawing during assembly using a globally unique directional reference.

[0063] Furthermore, to achieve the leap from geometric white models to high-fidelity twins, material types (e.g., concrete, frosted glass, wood grain finish) are decoupled and output from the structured twin data. These material types not only originate from mapping specific layers in the CAD drawings but also deeply integrate the user's high-level natural language commands contained in the "text features" from the aforementioned steps.

[0064] During the rendering generation stage, the downstream 3D graphics engine can automatically call the corresponding Physically Based Rendering (PBR) material sphere or texture library by parsing the material type field and assign it to the surface of the 3D solid model whose spatial pose has just been established.

[0065] By introducing rotation angle and material type, this application fully constructs a "lightweight structural twin dictionary" with high-dimensional semantics (e.g., a parameter set in JSON format: <identifier, center point coordinates, size, rotation angle, material type>). This highly parameterized data structure completely abandons the traditional generative model's practice of directly outputting bloated and difficult-to-modify meshes, making the twin scene highly editable. Users only need to modify a single line of values ​​in the parameter dictionary to instantly change the orientation or material of a wall in the scene, greatly reducing the threshold and computational cost of secondary editing of digital twin scenes.

[0066] In some examples, after dynamically generating corresponding 3D solid models based on identifier bits and spatial location benchmarks to render and generate a 3D twin scene, the method also includes: automatically comparing the spatial location information of each 3D solid model in the 3D twin scene with the geometric constraints in the computer-aided design source data to generate spatial verification results; when the spatial verification results indicate that there is a positional deviation or physical collision, the center point coordinates, size parameters or rotation angles of the 3D solid models with deviations are automatically adjusted in the dimension of the structured twin data, and a re-rendering update is triggered to complete the fine-tuning of the 3D twin scene.

[0067] Specifically, in complex digital twin projects, although large multimodal twin models have achieved extremely high generation accuracy, tiny numerical truncation errors may still occur when faced with extremely dense primitives or extremely stringent industrial tolerances. This can lead to positional shifts or slight bounding box interference (collisions) in a very small number of models. To achieve truly industrial-grade delivery, this embodiment introduces a post-processing automated consistency check and closed-loop fine-tuning mechanism after rendering.

[0068] First, the server performs spatial verification. The verification module extracts the actual spatial location information (such as 3D bounding box coordinates) of each entity model in the rendered scene, maps it in reverse, and aligns it with the absolute geometric constraints (such as the start and end coordinates of the original line segments) obtained from the initial parsing of the CAD source data. Simultaneously, the system uses the physics engine's collision detection algorithms (such as AABB or OBB bounding box intersection tests) to check for illogical physical collisions in the scene. The results of the comparison and investigation are then summarized to generate a "spatial verification result".

[0069] Once the spatial verification results trigger a deviation threshold alarm (i.e., a wall is found to deviate from the baseline of the original drawing, or the models of two devices overlap), it will immediately enter the automatic fine-tuning stage.

[0070] Unlike traditional operations that heavily consume computational resources by modifying the vertices of the underlying mesh, the fine-tuning in this application is performed at an extremely high speed at the level of the parameter dictionary. With minimal computational overhead, a correction algorithm is invoked to automatically correct the <center point coordinates> (for translation correction), <size parameters> (for scaling correction), or <rotation angle> (for attitude correction) of the entity model with deviation.

[0071] After the parameters are modified, the updated lightweight data is re-injected into the rendering engine, triggering a local re-render update. This closed-loop mechanism enables this application to have self-verification and self-correction capabilities, greatly reducing the cost of later manual review and model conversion.

[0072] To better understand this application, this embodiment provides a more specific example for illustration.

[0073] This application defines a new model architecture, a data structure for model output, and a twin scene generation method to ensure that the generated twin scene achieves a 1:1 reproduction in terms of position and structure.

[0074] The overall architectural logic of this application is as follows: Figure 2 As shown, this application supports text, CAD source files (DXF source files), and CAD exported images (CAD image data) as data types. The data is used to extract features and infer through a twin multimodal large model, and the structured data of the twin scene is output, such as twin data describing a door, twin data describing a wall, etc. Finally, the structured data is processed and rendered to present the twin scene through the twin scene assembly method.

[0075] CAD parsing involves reading and decoding the precise coordinate data and layer attributes of geometric primitives (especially line segments) stored in open format files like DXF. DXF files are essentially structured text databases that record all entity information that makes up the drawing. By parsing its specific "ENTITIES" segments layer by layer, the program can accurately extract the start and end coordinates of each line segment, as well as its layer, linetype, and other information, thus losslessly and quantitatively converting the visual content of the drawing into a pure data set that the computer can understand and process.

[0076] The fundamental purpose of this approach is to establish a unique and precise spatial reference for subsequent "pixel-level accurate generation." By acquiring this basic line segment data, the model can accurately determine the precise position and outline of every wall and door in the original design on the two-dimensional plane. This is equivalent to providing an insurmountable geometric "hard constraint" for the generation process, ensuring that every object unit in the final generated twin scene is strictly aligned in space with its corresponding element in the original design drawings, thus laying the data foundation for achieving a 1:1 high-fidelity conversion from design drawings to the twin model.

[0077] The basic structure of twin multimodal large models, such as Figure 3 As shown, the structure includes: 1) Line Encoder: Line segments are extracted in parallel using the Transformer Encoder and ResNet architecture to extract feature information between CAD line segments, including line segment type, feature data of the line segments themselves, and spatial relationships between line segments; 2) Image Encoder: Image encoding uses the VisionTransformer architecture to extract feature information from the image, especially line segment information in the image, such as: line segment start coordinates, end coordinates, line segment type, line segment width, etc.; 3) Tokenizer / Embedding: The text tokenizer and embedding mainly convert text into machine-understandable continuous numbers, providing instruction information for subsequent model decoding; 4) Features: We use image features + line segment features + text features to concatenate the three types of data features, and record the coordinate position when concatenating them; 5) Decoder: The Transformer Decoder architecture is used to infer the final structured twin data one by one. The structured twin data output by the twin multimodal large model is represented by a specific structure. Its core design purpose is to achieve semantic and editable generation results, providing support for subsequent scene installations and twin scene rendering.

[0078] like Figure 4 As shown, the structure of this structured twin data includes: Identifier bit: Used to specify the type of the currently generated twin data, for example: 1 represents a wall, 2 represents a window, etc.; Center point: Specifies the location of the center point of the 3D solid model. The center point is generally the bottom center point of the smallest rectangular bounding box of the model. Length / Width: Specifies the size of the 3D solid model; Rotation angle: Specifies the rotation value of the 3D solid model based on the world map; Material type: Specifies the material of the 3D solid model, such as: metal material, solid color material, glass material, etc.; Other information: Additional attribute values ​​used to expand the model; Structured twin data generated from a large multimodal twin model is used to render a 3D twin scene through standard data parsing and device methods. The data parsing process must adhere to the specifications of the large model's output data structure. The twin scene device dynamically generates the twin model using the parsed twin data. For example, when generating a wall, the corresponding wall model is generated based on the wall's center point position, rotation angle, length, width, and custom materials.

[0079] For example, such as Figure 5 As shown, the process includes: 1) Preparing the CAD source file of DXF or exporting the CAD image. Generally, simple processing such as deleting some irrelevant information in the CAD source file can improve the accuracy and efficiency of generating twin data from large models. 2) Combine user instructions and CAD data to perform inference using a twin multimodal large model to obtain structured twin data (twin scene data). 3) Use a scene renderer to render the generated structured twin data into a 3D twin scene; 4) Detect whether there are errors between the generated 3D twin scene and the CAD source data, and select and edit the model in the scene using the mouse.

[0080] All of the above-mentioned optional technical solutions can be combined in any way to form the optional embodiments of this application, and will not be described in detail here.

[0081] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0082] Based on the same concept, this application also provides a three-dimensional twin scene generation device, such as... Figure 6 As shown, the 3D twin scene generation device includes: The acquisition module 601 is used to acquire computer-aided design source data, computer-aided design image data, and text instructions; The feature module 602 is used to parse the computer-aided design source data, extract the line segment features of the computer-aided design source data, extract visual features of the computer-aided design image data, generate image features, and encode text instructions to generate text features. The fusion module 603 is used to input line segment features, image features, and text features into a pre-trained twin multimodal large model. The twin multimodal large model is used to fuse line segment features, image features, and text features based on the spatial positional relationship of the features to obtain structured twin data. The structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters. The generation module 604 is used to parse the structured twin data, determine the spatial position reference based on the center point coordinates and size parameters, and dynamically generate the corresponding three-dimensional solid model based on the identifier and spatial position reference to render and generate a three-dimensional twin scene.

[0083] In some examples, the source data of computer-aided design is parsed to extract the line segment features of the source data. This includes: extracting the feature information of line segments in the source data of computer-aided design through a parallel Transformer encoder and residual network architecture to generate line segment features. The feature information includes: line segment type, feature data of the line segment itself, and spatial relationship between line segments.

[0084] In some examples, visual feature extraction is performed on computer-aided design image data to generate image features, including: extracting visual features from computer-aided design image data using a visual Transformer architecture to obtain line segment information in the computer-aided design image data to generate image features; wherein, the line segment information includes at least: line segment start coordinates, line segment end coordinates, line segment type, and line segment width.

[0085] In some examples, structured twin data also includes: the rotation angle of the 3D solid model determined based on the world coordinate system; and the material type that specifies the surface rendering properties of the 3D solid model.

[0086] In some examples, the spatial position reference is determined based on the center point coordinates and size parameters, including: extracting the center point coordinates and using them as the coordinate anchor point of the 3D solid model in the preset 3D coordinate system; parsing the length and width in the size parameters, and using the coordinate anchor point as the origin to calculate the spatial boundary coordinates of the 3D solid model in the 3D coordinate system; and combining the coordinate anchor point with the spatial boundary coordinates to establish the spatial position reference when generating the 3D solid model.

[0087] In some examples, corresponding 3D solid models are dynamically generated based on identifier bits and spatial location references to render a 3D twin scene. This includes: determining the object type of the 3D solid model to be generated based on the identifier bits; dynamically constructing a 3D solid model with a corresponding 3D contour based on the object type and size parameters; performing spatial pose transformation on the 3D solid model based on the spatial location reference to anchor the bottom center point of the 3D solid model to the 3D coordinates corresponding to the spatial location reference, thus completing the spatial positioning of the 3D solid model; and combining and rendering multiple 3D solid models that have completed spatial positioning to generate a 3D twin scene.

[0088] In some examples, based on the identifier and spatial location reference, the corresponding 3D solid model is dynamically generated. After rendering and generating the 3D twin scene, the device is also used to: automatically compare the spatial location information of each 3D solid model in the 3D twin scene with the geometric constraints in the computer-aided design source data to generate spatial verification results; when the spatial verification results indicate that there is a positional deviation or physical collision, the device automatically adjusts the center point coordinates, size parameters or rotation angle of the 3D solid model with deviation in the dimension of the structured twin data, and triggers re-rendering and updating to complete the fine-tuning of the 3D twin scene.

[0089] According to the solution provided in this application, computer-aided design (CAD) source data, CAD image data, and text instructions are acquired. The CAD source data is parsed to extract line segment features. Visual features are extracted from the CAD image data to generate image features. Text instructions are encoded to generate text features. The line segment features, image features, and text features are input into a pre-trained twin multimodal large model. The twin multimodal large model is used to fuse the line segment features, image features, and text features based on their spatial positional relationships to obtain structured twin data. The structured twin data includes: an identifier indicating the object type, center point coordinates, and size parameters. The structured twin data is parsed, and a spatial positional reference is determined based on the center point coordinates and size parameters. Based on the identifier and spatial positional reference, a corresponding 3D solid model is dynamically generated to render a 3D twin scene. This application provides precise underlying geometric coordinate constraints for the generation and inference of the twin multimodal large model by extracting line segment features from the CAD source data and aligning and fusing the multimodal data based on spatial positional relationships. This ensures that the subsequently generated 3D solid model can be precisely aligned with the original CAD drawings in terms of spatial position and size, effectively avoiding spatial illusions. Furthermore, this application significantly enhances the system's fault tolerance and overall parsing accuracy by complementing and cross-verifying line segment features, image features, and text features. This application outputs structured twin data containing identifiers, center point coordinates, and size parameters through a twin multimodal large model, decoupling the scene into a lightweight set of parameters. This not only significantly reduces the computational cost of data processing but also enables seamless secondary editing of the final generated 3D scene based on parameter modifications. In the scene rendering stage, this application does not directly generate 3D solid models but pre-establishes a spatial position benchmark (i.e., physical occupancy boundary) based on the center point coordinates and size parameters. Then, it dynamically generates the corresponding 3D solid model based on this benchmark and identifiers, providing strict physical boundary constraints for model instantiation. This fundamentally avoids interference and clipping misalignment that occur during model assembly in 3D space, ensuring the rendering quality of the twin scene and thus avoiding the spatial illusions and poor generation results found in existing technologies.

[0090] Figure 7 This is a schematic diagram of the electronic device 7 provided in an embodiment of this application. Figure 7As shown, the electronic device 7 of this embodiment includes a processor 701, a memory 702, and a computer program 703 stored in the memory 702 and executable on the processor 701. When the processor 701 executes the computer program 703, it implements the steps in the various method embodiments described above. Alternatively, when the processor 701 executes the computer program 703, it implements the functions of each module / unit in the various device embodiments described above.

[0091] Electronic device 7 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 7 may include, but is not limited to, processor 701 and memory 702. Those skilled in the art will understand that... Figure 7 This is merely an example of electronic device 7 and does not constitute a limitation on electronic device 7. It may include more or fewer components than shown, or different components.

[0092] The processor 701 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0093] The memory 702 can be an internal storage unit of the electronic device 7, such as a hard disk or RAM of the electronic device 7. The memory 702 can also be an external storage device of the electronic device 7, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., equipped on the electronic device 7. The memory 702 can also include both internal and external storage units of the electronic device 7. The memory 702 is used to store computer programs and other programs and data required by the electronic device.

[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0095] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in a computer-readable medium can be appropriately added or removed according to regional requirements and patent practice requirements. For example, in some regions, according to regional requirements and patent practice, a computer-readable medium may not include electrical carrier signals and telecommunication signals.

[0096] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating a three-dimensional twin scene, characterized in that, The method includes: Acquire source data, image data, and text instructions for computer-aided design; The computer-aided design source data is parsed to extract line segment features, the computer-aided design image data is subjected to visual feature extraction to generate image features, and the text instructions are encoded to generate text features. The line segment features, image features, and text features are input into a pre-trained twin multimodal large model. The twin multimodal large model is used to fuse the line segment features, image features, and text features based on the spatial positional relationship between them to obtain structured twin data. The structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters. The structured twin data is parsed, a spatial position reference is determined based on the center point coordinates and the size parameters, and a corresponding three-dimensional entity model is dynamically generated based on the identifier and the spatial position reference to render and generate a three-dimensional twin scene.

2. The method according to claim 1, characterized in that, The computer-aided design source data is parsed to extract line segment features, including: By using a parallel Transformer encoder and residual network architecture, feature information of line segments in the computer-aided design source data is extracted to generate the line segment features; wherein, the feature information includes: line segment type, feature data of the line segment itself, and spatial relationship between line segments.

3. The method according to claim 1, characterized in that, Visual feature extraction is performed on the computer-aided design image data to generate image features, including: Visual features are extracted from the computer-aided design image data using a visual Transformer architecture to obtain line segment information in the computer-aided design image data, thereby generating the image features; wherein, the line segment information includes at least: line segment start coordinates, line segment end coordinates, line segment type, and line segment width.

4. The method according to claim 1, characterized in that, The structured twin data also includes: the rotation angle of the three-dimensional solid model determined based on the world coordinate system; and the material type specifying the surface rendering attributes of the three-dimensional solid model.

5. The method according to claim 1, characterized in that, Determining a spatial position reference based on the center point coordinates and the size parameters includes: Extract the coordinates of the center point and use the coordinates of the center point as the coordinate anchor point of the three-dimensional solid model in the preset three-dimensional coordinate system; Analyze the length and width in the size parameters, and calculate the spatial boundary coordinates of the three-dimensional solid model in the three-dimensional coordinate system with the coordinate anchor point as the origin. The coordinate anchor point is combined with the spatial boundary coordinates to establish the spatial position reference when generating the three-dimensional solid model.

6. The method according to claim 1, characterized in that, Based on the identifier and the spatial location reference, a corresponding 3D entity model is dynamically generated to render a 3D twin scene, including: The object type of the three-dimensional solid model to be generated is determined based on the identifier bit; Based on the object type and the size parameters, a three-dimensional solid model with a corresponding three-dimensional contour is dynamically constructed; Based on the spatial position reference, the three-dimensional solid model is subjected to spatial pose transformation to anchor the bottom center point of the three-dimensional solid model on the three-dimensional coordinates corresponding to the spatial position reference, thereby completing the spatial positioning of the three-dimensional solid model. The multiple 3D entity models that have completed spatial positioning are combined and rendered to generate the 3D twin scene.

7. The method according to claim 1, characterized in that, Based on the identifier and the spatial location reference, a corresponding 3D solid model is dynamically generated. After rendering and generating a 3D twin scene, the method further includes: The spatial position information of each three-dimensional entity model in the three-dimensional twin scene is automatically compared with the geometric constraints in the computer-aided design source data to generate spatial verification results. When the spatial verification result indicates the presence of positional deviation or physical collision, the center point coordinates, size parameters, or rotation angles corresponding to the 3D entity model with deviation are automatically adjusted in the dimension of the structured twin data, and a re-rendering update is triggered to complete the fine-tuning of the 3D twin scene.

8. A three-dimensional twin scene generation device, characterized in that, The device includes: The acquisition module is used to acquire computer-aided design source data, computer-aided design image data, and text instructions; The feature module is used to parse the computer-aided design source data, extract the line segment features of the computer-aided design source data, extract visual features of the computer-aided design image data to generate image features, and encode the text instructions to generate text features. The fusion module is used to input the line segment features, image features, and text features into a pre-trained twin multimodal large model. The twin multimodal large model is used to fuse the line segment features, image features, and text features based on the spatial positional relationships of the line segment features, image features, and text features to obtain structured twin data. The structured twin data includes: an identifier bit indicating the object type, center point coordinates, and size parameters. The generation module is used to parse the structured twin data, determine the spatial position reference based on the center point coordinates and the size parameters, and dynamically generate the corresponding three-dimensional entity model based on the identifier and the spatial position reference to render and generate a three-dimensional twin scene.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.