A digital twin scene rapid construction method based on cross-modal generation

By using a cross-modal generation framework and digital twin construction method, the problem of low efficiency in building digital twin scenarios is solved, enabling rapid construction in data-scarce environments. This method is applicable to application scenarios such as complex product manufacturing units and smart blocks.

CN119206062BActive Publication Date: 2025-11-21SOUTHEAST UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411258603.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2025-11-21
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing technologies lack sufficient methods for constructing digital twin scenarios, resulting in high costs and time requirements, especially in environments with scarce data or no samples, where rapid construction is difficult.

Method used

Employing a cross-modal generation framework, which includes a text generation module, an image generation module, and a model remodeling module, and combining a large language model, a stable diffusion model, and a camera viewpoint control method, it achieves the conversion from single-view images to multi-view images, and rapidly constructs high-quality 3D models through digital twin model construction, assembly, and fusion methods.

Benefits of technology

It accelerates the construction efficiency of digital twin scenarios in low-sample or no-sample environments, provides the possibility of object datafication in resource-constrained environments, and reduces the cost and time requirements of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206062B_ABST
    Figure CN119206062B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of based on cross-modal generation digital twin scene fast construction method.This application is composed of cross-modal data generation framework and digital twin construction framework.Cross-modal data generation framework includes text generation module, image generation module and model reconstruction module;Digital twin construction framework includes digital twin model construction method, digital twin model assembly method, digital twin model fusion method, digital twin model verification and calibration method.The method defines the fast virtual modeling strategy in low sample application scene, provides the basis for virtual experiment and virtual-real symbiosis, speeds up the construction efficiency of twin scene, and provides new possibility for the data of scene in resource-constrained environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for rapidly constructing digital twin scenes based on cross-modal generation, belonging to the field of digital twin modeling technology. Background Technology

[0002] With the continuous development of information and digital technologies, digital twin technology has become a hot research topic due to its significant potential in equipment monitoring, predictive maintenance, and operation optimization. The term "digital twin" was first proposed by American scholar Grieves to describe physical and virtual products and the relationship between them, initially used for the health maintenance and assurance of aerospace vehicles. In recent years, it has received widespread attention from scholars both domestically and internationally and has been widely applied in many industrial fields. In 2012, NASA provided a conceptual description of digital twins: Digital twins refer to a simulation process that fully utilizes physical models, sensors, operational history, and other data to integrate multiple disciplines and scales. By creating a virtual copy of a physical entity, it can simulate its performance in real time, predict future states, and provide decision support.

[0003] Digital twins, as a crucial enabling pathway for digital transformation and intelligent upgrading, have garnered significant attention across various industries and have moved from theoretical research to practical application. Leading companies in many sectors have begun experimenting with digital twins to identify problems and improve efficiency. However, much work remains to be done to realize the full potential of digital twins. For instance, in application scenarios, each model is built from scratch, and current research on scenario construction methods and standards is insufficient. Furthermore, collecting various modal data from a large number of sensors is challenging, incurring extremely high costs and time requirements. Therefore, data issues and modeling challenges are significant in the process of constructing digital twin scenarios. In conclusion, further research on methods for constructing digital twin scenarios is urgently needed in the field of digital twins. Summary of the Invention

[0004] This invention addresses the problems existing in the prior art by providing a rapid construction method for digital twin scenarios based on cross-modal generation. It includes a cross-modal data generation framework and a digital twin construction framework. This method defines a rapid construction approach for digital twin scenarios in data-scarce or sample-less engineering environments, accelerating the construction efficiency of digital twin scenarios and providing new possibilities for object datafication in resource-constrained environments.

[0005] To achieve the above objectives, the technical solution of this invention is as follows: a rapid construction method for digital twin scenes based on cross-modal generation. This method consists of a cross-modal generation framework and a digital twin scene construction framework. The cross-modal generation framework includes three modules: a text generation module, an image generation module, and a model remodeling module. The digital twin scene construction framework includes four parts: a digital twin model construction method, a digital twin model assembly method, a digital twin model fusion method, and a digital twin model verification and calibration method. The specific construction process and steps are as follows:

[0006] Step 1: Based on the twin requirements and current status of the project scenario, construct the necessary generation instructions for the image generation module using the text generation module. Then, within the image generation module, generate the target image set based on the image generation instructions exported from the text generation module. Next, input the image set into the model remodeling module to complete the generation of images from a single perspective to multiple perspectives. Based on the multiple perspective images, complete the 3D model reconstruction to obtain a high-quality mesh model and rendering.

[0007] Step 2: Input the model file and rendering generated in Step 1 into the digital twin scene construction framework. Based on the digital twin model construction method, complete the setting of the geometric model, physical model, behavioral model, and rule model. Then, using the digital twin model assembly method, assemble the twin model from a unit-level model to a system-level scene in the spatial dimension according to the knowledge graph and engineering concept. Next, according to the digital twin model fusion method, achieve multi-domain fusion of the models involved in the project through a multi-domain algorithm library and neural network technology. Finally, after model construction, assembly, and fusion, verification is required to ensure correctness and effectiveness. Verify that the model output is consistent with the actual output to ensure accuracy. Verify the unit-level model first, then verify the assembled or fused model. If it meets the actual project requirements, it can be applied; if not, it needs correction. Iterate the verification and correction until the project requirements are met.

[0008] Furthermore, the text generation module, image generation module, and model remodeling module in step 1 are specifically implemented as follows:

[0009] Step 11: The core of the text generation module is a large language model. This invention uses a deep learning model pre-trained on large-scale data. The underlying architecture of this model adopts a converter neural network, consisting of an encoder and a decoder, and supports a self-attention mechanism. By setting a template and inputting a clearly defined data generation task description, the text generation module can generate appropriate image generation instructions. Simultaneously, to increase the effectiveness of the generated text and alleviate semantic discrepancies, the text generation module uses in-context learning methods to assist in executing the instruction generation task, improving the model's learning ability in low-shot and zero-shot environments.

[0010] Step 12: The core of the image generation module is a stable diffusion model. The model used in this invention employs a step-by-step and iterative forward and backward diffusion process, exhibiting strong generalization performance. In the forward diffusion stage, the model progressively adds noise to the latent space, transforming the image into a random noise distribution. In the backward diffusion stage, the model estimates the image noise in the latent space using a noise predictor and progressively removes the noise to restore a clear image. Depending on the input content of the specific application scenario, the constructed image generation module can mainly accept three forms of modal data: text input and combined text-image input.

[0011] Step 13: Since the image sets output in Step 12 are all single-viewpoint image data, a mechanism for controlling camera extrinsic parameters other than those captured in the photographs needs to be added to the model remodeling group. The core of the model remodeling group consists of two parts: a camera viewpoint control method and a 3D model reconstruction method.

[0012] 1) The camera viewpoint control method enables the model to acquire a general mechanism for controlling the camera viewpoint by fine-tuning a pre-trained diffusion model. Based on this method, the model's ability to perform novel view synthesis can be unlocked. The camera viewpoint control method can achieve the conversion from single-view image data to multi-view image data.

[0013] 2) The 3D model reconstruction method references the Score Jacobian Chaining open-source framework, a primary algorithm used in zero-shot development and 3D generation. This framework treats the 3D model as a single point and optimizes it using stochastic gradient descent. Within the 3D model reconstruction method, the classifier-free guidance value is set higher than the conventional value, thereby improving the fidelity of the reconstruction.

[0014] Furthermore, the digital twin model construction method, digital twin model assembly method, digital twin model fusion method, and digital twin model verification and calibration method in step 2 are specifically implemented as follows:

[0015] Step 21: The goal of the digital twin model construction method is to complete the setting of the geometric model, physical model, behavioral model, and rule model based on the mesh data generated and rendered in Step 1. The digital twin model construction method completes the setting of the object's "geometry-physical-behavioral-rule" characteristics from multiple dimensions and domains. It can be adjusted according to actual needs to construct the core multi-domain dimensional model.

[0016] Step 22: The goal of the digital twin model assembly method is to assemble the twin model from a unit-level model to a system-level scene in the spatial dimension. This method includes the following: First, constructing the hierarchical relationships within the twin scene and defining their assembly order; second, adding spatial constraints that conform to physical properties during the model assembly process; and finally, assembling the twin model based on the constructed constraints and assembly order.

[0017] Step 23: The goal of the digital twin model fusion method is to construct a multi-disciplinary, multi-domain integrated digital twin model based on the subject areas involved in the modeling object. The digital twin model fusion method includes multi-domain algorithm libraries covering mechanical, electrical, and hydraulic systems, as well as neural network technologies (AI motion control, path optimization algorithms, image vision detection technology, etc.) to achieve multi-domain fusion of the models involved in the project.

[0018] Step 24: After steps 21, 22, and 23, the constructed digital twin model needs to be validated to ensure its correctness and effectiveness. The validation and calibration criteria are based on project requirements to verify whether the model's output matches the output of the physical object. Model validation and calibration is an iterative process. By selecting model calibration parameters and constructing the objective function, the model's accuracy is continuously optimized, making the digital twin model construction method described in Step 2 better adaptable to different application needs, conditions, and scenarios until it meets project requirements.

[0019] This technical solution explores a method for constructing a digital twin model based on object-oriented thinking, which meets the needs of digital twin modeling in workshops and can provide guidance for the digital twin transformation of production workshops. At the same time, the method defines the modeling mechanism of digital twins under different conditions, clarifies the modeling elements of digital twins in different application scenarios, standardizes the digital twin modeling process, and reduces efficiency losses caused by unclear application scenarios and unclear assembly methods.

[0020] Beneficial Effects: Compared with the prior art, the present invention has the following significant advantages: (1) In traditional application scenarios, the construction of digital twin scenarios is generally based on building each model from scratch. It is difficult to collect various modal data from a large number of sensors, resulting in extremely high costs and time requirements. The present invention addresses the problems existing in the prior art by providing a method for rapid construction of digital twin scenarios based on cross-modal generation, which includes a cross-modal data generation framework and a digital twin construction framework. This method defines a rapid construction method for twin scenarios in data-scarce or sample-less engineering environments, accelerates the construction efficiency of twin scenarios, and provides new possibilities for object datafication in resource-constrained environments; (2) The present invention constructs a cross-modal twin model generation framework, which includes three modules: a text generation module, an image generation module, and a model remodeling module. It can construct the generation instructions required for the image generation module through the text generation module in the case of no or low sample. Then, the target image group is generated according to the image generation instructions in the image generation module; and the three-dimensional model is reconstructed based on the single-view image in the model remodeling group, realizing the rapid generation of model data in the state of no sample or low sample; (3) Based on cross-modal generation data, this invention proposes a subsequent digital twin scene construction framework, which includes four parts: digital twin model construction method, digital twin model assembly method, digital twin model fusion method, and digital twin model verification and calibration method. According to the knowledge graph and engineering concept, and combined with the database and neural network, the rapid virtual modeling of low sample application scenarios is realized. Attached Figure Description

[0021] Figure 1 This is a flowchart of a method for rapidly constructing a digital twin scene based on cross-modal generation according to the present invention;

[0022] Figure 2 This is a flowchart illustrating the construction process of a twin scenario for a complex product manufacturing unit in the embodiment.

[0023] Figure 3 This is a flowchart illustrating the construction process of a twin scene for a smart street scenario in the embodiment. Detailed Implementation

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] Example 1: A twin scenario of a complex product manufacturing unit is constructed according to the present invention. Taking a simulated satellite complex product manufacturing unit scenario as an example, the construction process is as follows: Figure 2 As shown, the generation of complex product models is first completed within the cross-modal data generation framework: a satellite component image generation instruction (Prompt) is constructed through a text generation module. machine The input is sent to the image generation module to generate an image set (Image). machine), and complete the model reconstruction (3D) within the model remodeling group. machine Next, within the digital twin construction framework, the construction of a digital twin scenario for a complex product manufacturing unit is completed: digital twin model construction based on the complex product model (DT). BUILD Digital twin model assembly and fusion (DT) ASS Digital twin model verification and calibration methods (DT) Testing Specifically, the steps include the following:

[0026] Step 1: Based on the twin requirements and current status of the satellite complex product manufacturing unit scenario, construct the generation instructions (Prompt) required for the image generation module through the text generation module. machine (prompt1-prompt n Then, within the image generation module, the target image group is generated based on the image generation instructions exported from the text generation module. machine (imageData1-imageData n Next, the image group is input into the model remodeling group to generate images from a single viewpoint to multiple viewpoints, and then the 3D model is reconstructed based on the multiple viewpoint images. machine (3Dmodel1-3Dmodel n This allows you to obtain high-quality mesh models and rendering.

[0027] ① Constructing the image generation command Prompt machine (prompt1-prompt n The large language model used in the text generation module within the instance is based on the generative language model ChatGPT4.0. First, the example and test cases are input into In-Context Learning to generate templates for image production instructions within a complex product manufacturing unit scenario. Then, a clearly defined data generation task description is input to obtain the corresponding image generation instructions.

[0028] ② Generate the target image group Image machine (imageData1-imageData n This example demonstrates the generation of a single-view image set for a complex product model based on a stable diffusion model. First, the model detection points, VAE encoder, CLIP encoder, and optimization scheme are selected. Then, the image generation instructions (prompt1-prompt) are input. n The target image set is generated using a sample image and an optional example image.

[0029] ③ 3D model reconstruction machine (3Dmodel1-3Dmodel nThis example demonstrates how to reconstruct a complex product model in 3D using camera viewpoint control and 3D model reconstruction methods within the model remodeling group. First, we analyze the target image group (imageData1-imageData2). n Preprocessing is performed to remove interference and background. Next, seed values ​​and example steps are set within the camera viewpoint control method and 3D model reconstruction method to acquire multi-view image data and generate high-quality mesh models and renderings.

[0030] Step 2: Input the complex product model file and rendering generated in Step 1 into the digital twin scene construction framework. Complete the setting of the geometric model, behavioral model, and rule model based on the digital twin model construction method. Then, using the digital twin model assembly method, assemble the product model-manufacturing unit twin scene in the spatial dimension. Next, complete the fusion of the mechanical / electrical / control models of the manufacturing unit twin scene based on the engineering algorithm library. Finally, verify the accuracy of the model's output compared to the actual scene output.

[0031] In summary, this embodiment can quickly construct digital twin scenarios of complex product manufacturing units, accelerating the construction efficiency of digital twin scenarios. Without large-scale sensor deployment and multimodal data acquisition, the rapid construction method for digital twin scenarios based on cross-modal generation provided by this invention achieves rapid virtual modeling of low-sample application scenarios through a cross-modal data generation framework and a digital twin construction framework.

[0032] Example 2: A twin scenario of a smart street was constructed based on this invention. Taking a proposed smart street scenario as an example, the construction process is as follows: Figure 3 As shown, the main models (streetlights, buildings, bridges, roads, etc.) in the smart street scene are first generated within the cross-modal data generation framework: a smart street image generation instruction (Prompt) is constructed using a text generation module. street The image is input into the image generation module to generate a smart street image. street And within the model remodeling group, the reconstruction of streetlights, buildings, bridges, roads, and other models within the block was completed (3D). street Next, within the digital twin construction framework, the construction of a smart street twin scenario is completed: a digital twin model is constructed based on the models of various elements of the smart street (DT). BUILD Digital twin model assembly and fusion (DT) ASS Digital twin model verification and calibration methods (DT) Testing Specifically, the steps include the following:

[0033] Step 1: Based on the twin requirements and current status of the smart street scene, construct the generation instructions (Prompt) required for generating element images within the street using the text generation module. street(prompt1-prompt n Then, within the image generation module, a set of smart street element images (streetlights, buildings, bridges, roads, etc.) is generated based on the generation instructions. street (imageData1-imageData n Next, the image group is input into the model remodeling group to generate feature images from a single viewpoint to multiple viewpoints, and then a 3D model of the street features is reconstructed based on the multiple viewpoint images. street (3Dmodel1-3Dmodel n This allows you to obtain high-quality mesh models and rendering.

[0034] ① Constructing the Smart Street Image Generation Instruction Prompt street (prompt1-prompt n): The large language model for the text generation module within the instance is based on the generative language model ChatGPT4.0. First, the example and test cases are input into In-Context Learning to generate image production instruction templates for elements within the smart street scene. Then, a clear description of the elements is input to obtain the smart street image generation instructions.

[0035] ② Generate a set of smart street block element images. street (imageData1-imageData n This example generates a single-view image set of smart street elements (streetlights, buildings, bridges, roads, etc.) based on a stable diffusion model. After selecting model detection points, VAE encoder, CLIP encoder, and optimization scheme, input the street image generation command and an example street image (optional) to generate the street element image set.

[0036] ③ 3D reconstruction of street block elements street (3Dmodel1-3Dmodel n This example demonstrates how the core elements of a street block are reconstructed in 3D using camera viewpoint control and 3D model reconstruction methods within the model remodeling group. First, we preprocess the smart street block element images to remove interference and background. Then, we set parameters within the camera viewpoint control and 3D model reconstruction methods to acquire multi-view image data of the street block elements and generate a high-quality 3D model of the street block elements.

[0037] Step 2: Input the 3D model of the street elements generated in Step 1 into the digital twin scene construction framework. Based on the digital twin model construction method, complete the setting of the geometric model, physical model, behavioral model, and rule model. Then, using the digital twin model assembly method, assemble the street elements into a smart street twin scene in the spatial dimension. Next, based on the engineering algorithm library and artificial intelligence algorithms, complete the fusion of multi-domain models and functional models (path planning, etc.) of the smart street twin scene. Finally, verify the real-time operation of the model and the matching degree between the prediction results and the actual state of the scene.

[0038] In summary, this embodiment can quickly construct a digital twin scene of a smart street, accelerating the construction efficiency of the digital twin scene. Using the rapid construction method for digital twin scenes based on cross-modal generation provided by this invention, model data of core elements of a smart street (streetlights, buildings, bridges, roads, etc.) are quickly obtained through a cross-modal data generation framework, and rapid virtual modeling of the smart street twin scene is achieved within the digital twin construction framework.

[0039] It should be noted that the above embodiments are not intended to limit the scope of protection of the present invention. Equivalent transformations or substitutions made based on the above technical solutions all fall within the scope of protection of the claims of the present invention.

Claims

1. A method for rapid construction of digital twin scenes based on cross-modal generation, characterized in that, The method includes the following steps: Step 1: Based on the twin requirements and current status of the project scenario, construct the necessary generation instructions for the image generation module using the text generation module. Then, within the image generation module, generate the target image group based on the image generation instructions exported from the text generation module. Next, input the image group into the model remodeling module to complete the generation of images from a single perspective to multiple perspectives. Based on the multiple perspective images, complete the 3D model reconstruction to obtain a high-quality mesh model and rendering. Step 2: Input the model file and rendering generated in Step 1 into the digital twin scene construction framework. Based on the digital twin model construction method, complete the setting of geometric model, physical model, behavioral model, and rule model. Then, using the digital twin model assembly method, based on the knowledge graph and engineering concept, complete the assembly of the twin model from unit-level model to system-level scene in the spatial dimension. Next, according to the digital twin model fusion method, realize the multi-domain fusion of models involved in the project through multi-domain algorithm library and neural network technology. Finally, after model construction, assembly, and fusion, it is necessary to verify to ensure correctness and effectiveness, verify that the model output is consistent with the actual output, and ensure accuracy. Unit-level models are verified first, and then the assembled or fused models are verified. If they meet the actual project requirements, they can be applied. If they do not meet the requirements, they need to be corrected. The verification and correction are iterated until the project requirements are met. The text generation module, image generation module, and model remodeling module in step 1 are implemented as follows: Step 11: The core of the text generation module is a large language model, which uses a deep learning model pre-trained on large-scale data. The underlying architecture of this model employs a converter neural network, consisting of an encoder and a decoder, and supports self-attention. By setting templates and inputting a clearly defined data generation task description, the text generation module can generate appropriate image generation instructions. The text generation module uses in-context learning methods to assist in executing the instruction generation task, improving the model's learning ability in low-shot and zero-shot environments. Step 12: The core of the image generation module is a stable diffusion model. The model employs a step-by-step and iterative forward and backward diffusion process, exhibiting strong generalization performance. In the forward diffusion stage, the model progressively adds noise to the latent space, transforming the image into a random noise distribution. In the backward diffusion stage, the model estimates the image noise in the latent space using a noise predictor and progressively removes the noise to restore a clear image. Depending on the input content of the specific application scenario, the constructed image generation module can mainly accept three forms of modal data: text input, and combined text and image input. Step 13: Since the image sets output in Step 12 are all image data from a single viewpoint, a mechanism for controlling camera extrinsic parameters other than those captured in the photographs needs to be added to the model remodeling group; In step 13, the core of the model remodeling group consists of two parts: the camera viewpoint control method and the 3D model reconstruction method. Among them, the camera viewpoint control method enables the model to acquire a general mechanism for controlling the camera viewpoint by fine-tuning the pre-trained diffusion model. Based on this method, the model's ability to perform novel view synthesis can be unleashed. The camera viewpoint control method realizes the conversion from single-view image data to multi-view image data. The 3D model reconstruction method references the Score Jacobian Chaining open-source framework, which is the main algorithm used in zero-shot development and 3D generation work. It treats the 3D model as a single point and optimizes the model through stochastic gradient descent. In the 3D model reconstruction method, the classifier-free guidance value is set higher than the normal value, thereby improving the fidelity of the reconstruction.

2. The method for rapid construction of digital twin scenes based on cross-modal generation according to claim 1, characterized in that, The digital twin model construction method, digital twin model assembly method, digital twin model fusion method, and digital twin model verification and calibration method in step 2 are specifically implemented as follows: Step 21: The goal of the digital twin model construction method is to complete the setting of the geometric model, physical model, behavioral model, and rule model based on the mesh data generated and rendered in Step 1. Through the digital twin model construction method, the object's "geometry-physical-behavioral-rule" characteristics are set from multiple dimensions and domains. This can be adjusted according to actual needs to construct the core multi-domain dimensional model. Step 22: The goal of the digital twin model assembly method is to assemble the twin model from a unit-level model to a system-level scene in the spatial dimension. This method includes the following: First, constructing the hierarchical relationships within the twin scene and defining their assembly order; second, adding spatial constraints that conform to physical properties during the model assembly process; and finally, assembling the twin model based on the constructed constraints and assembly order. Step 23: The goal of the digital twin model fusion method is to construct a multi-disciplinary and multi-domain integrated digital twin model based on the subject areas involved in the modeling object. The digital twin model fusion method includes a multi-domain algorithm library covering mechanical, electrical, and hydraulic fields, and neural network technology to achieve multi-domain fusion of the models involved in the project. Step 24: After steps 21, 22, and 23, the constructed digital twin model needs to be verified to ensure its correctness and effectiveness. The verification and calibration standard is to check whether the output of the model is consistent with the output of the physical object according to the project requirements. Model verification and calibration is an iterative process. By selecting model calibration parameters and constructing objective functions, the accuracy of the model is continuously optimized, so that the digital twin model construction method in step 2 can better adapt to different application requirements, conditions, and scenarios until it meets the project requirements.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous take-off and landing cruising method based on digital twinning

    CN113406968A

  • Digital twinning learning scene reconstruction method and system based on context information

    CN117251896A

  • Digital twin modeling method and system based on SAM large model and NeRF

    CN117671138A