An automated construction method for digital twins based on multi-dimensional perception data

Through the automated construction method of multi-dimensional perceptual data, combined with multi-sensor data and DETR network, the automated construction of digital twins in large-scale and special scenarios is realized, solving the problems of cumbersome and high cost in traditional methods and is suitable for multi-environmental conditions.

CN116071559BActive Publication Date: 2025-06-27NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211446984.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-18
Publication Date
2025-06-27
Estimated Expiration
2042-11-18

AI Technical Summary

Technical Problem

The existing digital twin model research is only aimed at a single component, and the construction process is cumbersome and costly, and it cannot adapt to the needs of large-scale digital twin factories or cities. Especially in special scenarios such as extreme cold, the construction of traditional models has failed.

Method used

An automated construction method based on multi-dimensional perceptual data is adopted, and multi-sensor data is obtained using binocular vision sensors and lidar sensors. Combined with multi-scale extraction and depth estimation modules, a fusion depth map and three-dimensional model are built, and target semantic information is obtained through the DETR network, matching the heterogeneous data of the Internet of Things sensors, and the automated construction of digital twins is realized.

Benefits of technology

It realizes the automated construction of digital twins in large-scale scenarios and special scenarios, avoids the tedious processes and high costs in traditional methods, improves the construction efficiency and accuracy, and is suitable for multiple environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116071559B_ABST
    Figure CN116071559B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for automatically constructing a digital twin based on multi-dimensional perception data, comprising the following steps: obtaining a color image and a point cloud depth map of the area to be detected; processing the color image using a multi-scale extraction and depth estimation module to obtain a visual depth map based on pixel information; extracting color-dominated depth information and depth-dominated depth information according to the color image, the point cloud depth map and the visual depth map, constructing a fused depth map, and reconstructing a three-dimensional model of the area to be detected; constructing and using a DETR network to obtain the target semantic information of the visual depth map; obtaining heterogeneous data of Internet of Things sensors based on the area to be detected; matching the three-dimensional model of the area to be detected with the target semantic information, and matching the matching result with the heterogeneous data of Internet of Things sensors to construct a digital twin of the area to be detected. The present invention reduces the difficulty of constructing a digital twin and improves the effectiveness of feature extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital twins, and specifically includes an automated construction method for digital twins based on multi-dimensional perception data. Background Art

[0002] As an emerging technology in the next-generation manufacturing field, digital twins provide key support for the intelligent and digital transformation of the manufacturing industry. With the development of emerging technologies such as the Internet of Things, big data, cloud computing, and artificial intelligence, digital twin technology has gradually matured. The manufacturing industry can achieve data-driven operation status monitoring, trace back data for product iteration, develop innovative products and services, and realize value creation and diverse business models. However, current research on digital twin models only focuses on single components, and the process of constructing digital twins is cumbersome and costly, unable to meet the development needs of current digital twin factories, digital twin cities, etc. At the same time, for some special scenarios such as extremely cold regions, traditional model construction methods fail and cannot meet the requirements for constructing digital twins in specific scenarios. Therefore, the lack of a construction method for large-scale and high-precision digital twin models restricts the development of digital twins, so an automated construction method for digital twins applicable to multiple environments and large scales needs to be solved. Summary of the Invention

[0003] Aiming at the above deficiencies in the prior art, an automated construction method for digital twins based on multi-dimensional perception data provided by the present invention solves the problems of cumbersome construction process, high cost, and difficulty in constructing digital twins in specific scenarios.

[0004] To achieve the above invention objective, the technical solution adopted by the present invention is: an automated construction method for digital twins based on multi-dimensional perception data, including the following steps:

[0005] S1. Use a binocular vision sensor and a lidar sensor to scan the area to be detected, and obtain a color image and a point cloud depth map of the area to be detected;

[0006] S2. Use a multi-scale extraction and depth estimation module to process the color image to obtain a visual depth map based on pixel information;

[0007] S3. Use the binocular vision sensor to extract color-dominated depth information and depth-dominated depth information according to the color image, point cloud depth map, and visual depth map, construct a fused depth map, and reconstruct the three-dimensional model of the area to be detected;

[0008] S4. Construct and use a DETR network to obtain target semantic information of the visual depth map;

[0009] S5. Obtain heterogeneous data of Internet of Things sensors based on the area to be detected;

[0010] S6. Match the 3D model of the area to be detected with the target semantic information, match the matching result with the heterogeneous data of the Internet of Things sensors in the area to be detected, and construct the digital twin of the area to be detected.

[0011] Further, the specific implementation method of step S1 is as follows:

[0012] Use a binocular vision sensor to scan the area to be detected to obtain a color image of the area to be detected; use a lidar sensor to scan the area to be detected to obtain a point cloud depth map of the area to be detected.

[0013] Further, the specific implementation method of step S2 is as follows:

[0014] S2-1. Perform the first convolution on the color image collected by the left-eye sensor of the binocular vision sensor to generate the feature map A1, pass the feature map A1 through a 3*3 convolutional layer, and perform a correlation operation with a 40-pixel displacement on the output of the 3*3 convolutional layer and the feature map A1 to obtain the correlation with a larger range but coarser texture of the left-eye sensor;

[0015] S2-2. Perform the second convolution on the color image collected by the left-eye sensor of the binocular vision sensor to obtain the left-eye feature map; upsample the left-eye feature map to obtain a full-resolution output, and perform a correlation operation with a 20-pixel displacement on the full-resolution output and the output result of the second convolutional layer to obtain the correlation with a smaller range but finer texture of the left-eye sensor;

[0016] S2-3. Perform the same operations as in steps S2-1 to S2-2 on the color image collected by the right-eye sensor of the binocular vision sensor to obtain the right-eye feature map, the correlation with a larger range but coarser texture of the right-eye sensor, and the correlation with a smaller range but finer texture of the right-eye sensor;

[0017] S2-4. Match the left-eye feature map and the right-eye feature map, and concatenate the matching result with the left-eye feature map to obtain the underlying semantic information;

[0018] S2-5. Pass the underlying semantic information through an encoder-decoder structure, and introduce skip connections at each scale of the encoder-decoder to obtain a full-resolution disparity map estimate;

[0019] S2-6. Concatenate the disparity map estimate with the correlation with a larger range but coarser texture of the right-eye sensor, the correlation with a smaller range but finer texture of the right-eye sensor, the correlation with a smaller range but finer texture of the left-eye sensor, and the correlation with a larger range but coarser texture of the left-eye sensor, and obtain the target depth estimate map through a convolution operation.

[0020] Further, the specific implementation method of step S3 is as follows:

[0021] S3-1. Convolve the color image, the point cloud depth map, and the visual depth map to obtain a color image feature map, a point cloud depth feature map, and a visual depth feature map;

[0022] S3-2. Convolve the color image feature map, the point cloud depth feature map, and the visual depth feature map respectively to obtain the unique feature C E of the color image, the unique feature P E of the point cloud depth map, and the unique feature D E of the visual depth map;

[0023] S3-3. Convolve the color image feature map and the point cloud depth feature map to obtain the common feature C S of the color image and the common feature P S of the point cloud depth map;

[0024] Convolve the point cloud depth feature map and the visual depth feature map to obtain the common feature P S of the point cloud depth map and the common feature D S of the visual depth map;

[0025] S3-4. Concatenate the unique feature C E of the color image and the unique feature P E of the point cloud depth map to obtain the unique feature C' E dominated by color; Concatenate the common feature C S of the color image and the common feature P S of the point cloud depth map to obtain the common feature C' S dominated by color;

[0026] Concatenate the unique feature D E of the visual depth map and the unique feature P E of the point cloud depth map to obtain the unique feature D' E dominated by depth; Concatenate the common feature D S of the visual depth map and the common feature P S of the point cloud depth map to obtain the common feature D' S dominated by depth;

[0027] S3-5. According to the formula:

[0028]

[0029] Use the confidence aggregation strategy to obtain the depth image pixel D(u, v) in the fused depth map; where, C' E (u, v) represents the unique feature point dominated by color; C' S (u, v) represents the common feature point dominated by color; D' E(u, (u, v) represents the unique feature points dominated by depth; D' S (u, v) represents the common feature points dominated by depth;

[0030] S3-6. Use the PCL triangulation algorithm to reconstruct the depth map based on the depth image pixels to obtain a three-dimensional model.

[0031] Furthermore, the specific implementation of step S4 is as follows:

[0032] S4-1. Divide the target depth estimation map into N slices to obtain a set of image sequences and corresponding position encodings;

[0033] S4-2. Construct and use the DETR network to process the image sequences and corresponding position encodings to obtain the target semantic information.

[0034] Furthermore, the specific implementation of step S4-2 is as follows:

[0035] S4-2-1. Input the image sequences and corresponding position encodings into the self-attention module of the DETR network for encoding and decoding to obtain an output sequence;

[0036] S4-2-2. Send the output sequence into the target prediction module of the DETR network to obtain the target category;

[0037] S4-2-3. Add a mask head to the target category, calculate the binary mask of each target category, and generate an attention heat map;

[0038] S4-2-4. Send the attention heat map into the target semantic segmentation module of the DETR network to obtain the target semantic information.

[0039] Furthermore, the specific implementation of step S6 is as follows:

[0040] S6-1. Match the three-dimensional model of the area to be detected with the target semantic information, establish a device information database, and store the target semantic information and the storage path of the three-dimensional model of the area to be detected;

[0041] S6-2. Classify and store the heterogeneous data of the Internet of Things sensors according to the target semantic information to complete the acquisition of device information data;

[0042] S6-3. Complete the construction of the digital twin according to the device information data.

[0043] Furthermore, the digital twin includes a model library management module, a storage module, a digital thread module, a virtual-real mapping interface module, an inter-body interaction interface module, a control assembly module, and a service interface module;

[0044] The control assembly module is connected to the model library management module, the storage module, the digital thread module, the assembly model tree reading module, the interaction interface module, and the service interface module.

[0045] The beneficial effects of the present invention are as follows:

[0046] 1. The digital twin construction based on multi-dimensional perception data proposed by the present invention can automatically construct digital twins in large-scale scenarios and special scenarios, eliminating the defects of complex and cumbersome processes for constructing a single digital twin.

[0047] 2. The present invention proposes a fusion neural network model based on binocular vision sensors and lidar sensors. By extracting features from multi-sensor data, the fusion reconstruction of the target model is completed, thereby completing the construction of a high-resolution digital twin geometric model. It avoids the problems that it is difficult for vision sensors to extract effective features in weak texture environments, and is greatly affected by light, and lidar sensors will degenerate in open environments, and will generate excessive noise in extreme weather such as rain, snow, etc.

[0048] 3. The present invention uses a vision sensor to scan for the recognition and semantic segmentation of target objects in the scene, obtains more target semantic information, and matches the geometric model with the multi-source heterogeneous data of the device Internet of Things sensors, which can effectively reduce the workload of manual pairing and automatically complete the construction of the device digital twin. Description of the Drawings

[0049] Figure 1 is the flowchart of the present invention;

[0050] Figure 2 is the structure diagram of the digital twin system;

[0051] Figure 3 is the overall structure schematic diagram of the device of the present invention;

[0052] Figure 4 is the schematic diagram of the visual depth map prediction network;

[0053] Figure 5 is the schematic diagram of the fusion depth network of common features and unique features;

[0054] Figure 6 is the schematic diagram of the visual semantic analysis network based on the self-attention network. Detailed Embodiments

[0055] The specific embodiments of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0056] As Figure 1 shown, a method for automatically constructing a digital twin based on multi-dimensional perception data includes the following steps:

[0057] S1. Use a binocular vision sensor and a lidar sensor to scan the area to be detected, and obtain a color image and a point cloud depth map of the area to be detected;

[0058] S2. Use a multi-scale extraction and depth estimation module to process the color image to obtain a visual depth map based on pixel information;

[0059] S3. Use the binocular vision sensor to extract color-dominated depth information and depth-dominated depth information according to the color image, the point cloud depth map, and the visual depth map, construct a fused depth map, and reconstruct the three-dimensional model of the area to be detected;

[0060] S4. Construct and use a DETR network to obtain target semantic information from the visual depth map;

[0061] S5. Obtain heterogeneous data of Internet of Things sensors based on the area to be detected;

[0062] S6. Match the three-dimensional model of the area to be detected with the target semantic information, and match the matching result with the heterogeneous data of the Internet of Things sensors in the area to be detected to construct the digital twin of the area to be detected.

[0063] The specific implementation manner of step S4 is as follows:

[0064] S4-1. Divide the target depth estimation map into N pieces to obtain a set of image sequences and corresponding position encodings;

[0065] S4-2. Construct and use a DETR network to process the image sequences and the corresponding position encodings to obtain target semantic information.

[0066] The specific implementation manner of step S4-2 is as follows:

[0067] S4-2-1. Input the image sequences and the corresponding position encodings into the self-attention module of the DETR network for encoding and decoding to obtain an output sequence;

[0068] S4-2-2. Feed the output sequence into the target prediction module of the DETR network to obtain the target category;

[0069] S4-2-3. Add a mask head to the target category, calculate the binary mask for each target category, and generate an attention heat map;

[0070] S4-2-4. Feed the attention heat map into the target semantic segmentation module of the DETR network to obtain the target semantic information.

[0071] The specific implementation method of step S6 is as follows:

[0072] S6-1. Match the three-dimensional model of the area to be detected with the target semantic information, establish a device information database, and store the target semantic information and the storage path of the three-dimensional model of the area to be detected;

[0073] S6-2. Classify and store the heterogeneous data of the Internet of Things sensors according to the target semantic information to complete the acquisition of device information data;

[0074] S6-3. Complete the construction of the digital twin according to the device information data.

[0075] As Figure 2 shown, the digital twin includes a model library management module, a storage module, a digital thread module, a virtual-real mapping interface module, an inter-body interaction interface module, a control assembly module, and a service interface module;

[0076] The control assembly module is connected to the model library management module, the storage module, the digital thread module, the virtual-real mapping interface module, the inter-body interaction interface module, and the service interface module.

[0077] As Figure 3 shown, connect the binocular vision sensor, the lidar sensor, and the camera position electrically. The binocular vision sensor is placed at both ends of the camera position, horizontally, and the lidar sensor is placed in the middle of the binocular vision sensor; use the binocular vision sensor to scan the area to be detected to obtain the color image of the area to be detected; use the lidar sensor to scan the area to be detected to obtain the point cloud depth map of the area to be detected.

[0078] As Figure 4 shown, the specific implementation method of step S2 is as follows:

[0079] S2-1. Perform the first convolution on the color image collected by the left-eye sensor of the binocular vision sensor to generate the feature map A1. Pass the feature map A1 through a 3*3 convolutional layer, and perform a correlation operation with a 40-pixel displacement on the output of the 3*3 convolutional layer and the feature map A1 to obtain a relatively large but texture-rough correlation of the left-eye sensor;

[0080] S2-2. Perform a second convolution on the color image collected by the left-eye sensor of the binocular vision sensor to obtain a left-eye feature map; upsample the left-eye feature map to obtain a full-resolution output, and perform a correlation operation with a 20-pixel displacement on the full-resolution output and the output result of the second convolutional layer to obtain a correlation with a smaller range but finer texture in the left-eye sensor;

[0081] S2-3. Perform the same operations as in steps S2-1 to S2-2 on the color image collected by the right-eye sensor of the binocular vision sensor to obtain a right-eye feature map, a correlation with a larger range but coarser texture in the right-eye sensor, and a correlation with a smaller range but finer texture in the right-eye sensor;

[0082] S2-4. Match the left-eye feature map and the right-eye feature map, and concatenate the matching result and the left-eye feature map to obtain the underlying semantic information;

[0083] S2-5. Pass the underlying semantic information through an encoder-decoder structure, and introduce skip connections at each scale of the encoder-decoder to obtain a full-resolution disparity map estimate;

[0084] S2-6. Concatenate the disparity map estimate with the correlation with a larger range but coarser texture in the right-eye sensor, the correlation with a smaller range but finer texture in the right-eye sensor, the correlation with a smaller range but finer texture in the left-eye sensor, and the correlation with a larger range but coarser texture in the left-eye sensor, and obtain the target depth estimate map through a convolution operation.

[0085] As Figure 5 shown, the specific implementation method of step S3 is as follows:

[0086] S3-1. Convolve the color image, the point cloud depth map, and the visual depth map to obtain a color image feature map, a point cloud depth feature map, and a visual depth feature map;

[0087] S3-2. Convolve the color image feature map, the point cloud depth feature map, and the visual depth feature map respectively to obtain the unique feature C E of the color image, the unique feature P E of the point cloud depth map, and the unique feature D E of the visual depth map;

[0088] S3-3. Convolve the color image feature map and the point cloud depth feature map to obtain the common feature C S of the color image and the common feature P S of the point cloud depth map;

[0089] Convolve the point cloud depth feature map and the visual depth feature map to obtain the common feature P S of the point cloud depth map and the common feature D S of the visual depth map;

[0090] S3-4. Concatenate the unique feature C of the color image E and the unique feature P of the point cloud depth map E to obtain the unique feature C' dominated by color E ; Concatenate the common feature C of the color image S and the common feature P of the point cloud depth map S to obtain the common feature C' dominated by color S ;

[0091] Concatenate the unique feature D of the visual depth map E and the unique feature P of the point cloud depth map E to obtain the unique feature D' dominated by depth E ; Concatenate the common feature D of the visual depth map S and the common feature P of the point cloud depth map S to obtain the common feature D' dominated by depth S ;

[0092] S3-5. According to the formula:

[0093]

[0094] Adopt a confidence aggregation strategy to obtain the depth image pixel D(u, v) in the fused depth map; where, C' E (u, v) represents the unique feature point dominated by color; C' S (u, v) represents the common feature point dominated by color; D' E(u ,v) represents the unique feature point dominated by depth; D' S (u, v) represents the common feature point dominated by depth;

[0095] S3-6. Use the PCL triangulation algorithm to reconstruct the depth map based on the depth image pixels to obtain a 3D model.

[0096] Such as Figure 6As shown in the figure, set the number of slices of the target depth estimation map to N, use the image sequence as V, and add the image sequence and the position encoding as K and Q. Input them into the multi-head self-attention module, add the output result to the original image sequence and regularize it, then input it into the FFN network, and perform addition and regularization again to obtain the target sequence features. Send the target sequence into the decoder, where the target sequence features form the new V, and the image target sequence serves as the new K. Use 100 coordinate vectors with initial values of position encoding to form Q for prediction. Through the multi-head self-attention module, parallel decoding with the target sequence features yields the output sequence. Pass the output sequence through the fully-connected prediction feed-forward neural network FFN to obtain the target category, add the mask head to the target category, calculate the binary mask for each target category, generate the attention heatmap, and obtain the semantic segmentation image through the FPN module.

[0097] In one embodiment of the present invention, the point cloud depth map obtained by convolving the color image feature map and the point cloud depth feature map has a common feature P S , and the point cloud depth map obtained by convolving the point cloud depth feature map and the visual depth feature map has a common feature P S can be replaced with each other.

[0098] According to the formula:

[0099]

[0100] the convolution result S(i,j) is obtained; where I represents the original image, K represents the convolution kernel, m and n respectively represent the length and width of the convolution kernel, and (i,j) represents the point on the original image.

[0101] The semantic segmentation image passes through a 3*3 convolution + GN + ReLU layer to obtain the feature map C2; the feature map C2 passes through a convolution with a stride of 4 to obtain the feature map C3; the feature map C3 passes through a convolution with a stride of 8 to obtain the feature map C4; the feature map C4 passes through a convolution with a stride of 16 to obtain the feature map C5; perform a 1*1 convolution on C5 to reduce the number of channels to obtain M5, then perform 2x nearest neighbor upsampling in sequence, and add the results to the results of 1*1 convolutions of C4, C3, and C2 respectively to obtain M4, M3, and M2. After obtaining the added features, use a 3*3 convolution to process the generated M5, M4, M3, and M2 to obtain the final features P2, P3, P4, and P5. Upsample the features of different scales and concatenate them to form the P2, P3, P4, and P5 features, and perform a concatenation operation on them to obtain the final semantic segmentation information for each target category.

[0102] The construction of the model library management module of the digital twin mainly iteratively manages models such as the geometric model, mechanism model, and failure model of the virtual entity generated by the physical entity. The geometric model of the digital twin refers to the digital model that describes the geometric characteristics of the physical entity of the device, which can facilitate users to intuitively understand and recognize the visual appearance characteristics of the physical entity. The construction of the intrinsic principle model of the device mainly analyzes the internal mechanism and information transmission mechanism during the operation of the device itself, and uses the optical principle and basic circuit laws to establish the model of the device operation process. The construction of the failure mechanism model of the device mainly analyzes the relevant state data of the device, establishes the failure mechanism model of the device, and uses algorithms such as data mining and data analysis to identify the obtained key characteristic values to achieve failure mechanism prediction.

[0103] The storage module of the digital twin refers to storing the historical and real-time model data, state data, and monitoring data of the digital twin physical entity and virtual entity into relational databases and non-relational databases according to their respective characteristics; during the operation of the digital twin system, various sensing data of the physical entity and various prediction results of the digital twin need to be stored. Its data has characteristics such as large scale, concurrency, and uninterruptibility, and different storage modes need to be selected according to their different characteristics to achieve fast reading and calling of the digital twin system. For device monitoring data, a relational database is used, and for state data, a bitmap storage of Redis is used. Monitoring data is mainly stored using user tables and device monitoring data tables. State data is mainly stored using bitmap storage, and the fault data of different devices needs to be stored in different bitmaps.

[0104] The operation of the digital twin system needs to interact with the physical entity in real time, and the data transmitted by the physical entity needs to be analyzed and processed in real time. Therefore, there are strict requirements for communication latency. To solve the real-time synchronization between the virtual entity and the physical entity in the digital twin, the concept of digital thread is proposed.

[0105] The virtual-real mapping interface of the device digital twin mainly realizes the real-time data interaction between the digital twin and its corresponding physical entity. For the real-time state mapping of the device digital twin, it is necessary to quickly import different models and replace the historical database parameters according to the changes of the physical entity of the device to achieve the real-time data interaction between the device digital twin and the physical entity. For the process of device display and sensor binding in the virtual space, it is necessary to obtain the assembly relationship of the model in the virtual space and adopt the assembly model tree reading algorithm.

[0106] The inter-body interaction interface mainly realizes the data interaction between multiple digital twins in the digital twin system. In multi-twin applications such as digital twin cities or digital twin factories, it is inevitable to involve the interaction between multiple digital twins. Therefore, an interface for interaction between twins is designed.

[0107] The service interface mainly implements the human-computer interaction mechanism of the digital twin, or the interaction mechanism with other systems. It manages various types of interaction interfaces, including the human-machine interface, software access interface, and service encapsulation function for the digital twin to provide services across domains.

[0108] The control assembly module of the digital twin is the central control center during the operation of the digital twin. It is mainly responsible for coordinating and controlling the orderly operation of various modules such as the geometric model, mechanism model, digital thread, virtual-real mapping interface, inter-body interaction interface, service interface, and the digital twin's own database. According to the unified scheduling and management of the digital twin system interface by the digital twin management platform, an interface management unified scheduling bus is designed and implemented. Through this scheduling bus, modules such as the assembly model tree reading module, interaction interface module, digital thread module, service interface module, and storage module interface can be accessed.

[0109] The present invention can automatically construct digital twins in large-scale scenarios and special scenarios, eliminating the defects of complex and cumbersome processes for constructing a single digital twin; the present invention avoids the problems that it is difficult for vision sensors to extract effective features in weak texture environments and is greatly affected by light, and that laser sensors degrade in open environments and generate excessive noise in extreme weather such as rain and snow. The present invention uses vision sensors to scan for the recognition and semantic segmentation of target objects in the scene, obtains more target semantic information, and matches the geometric model with the multi-source heterogeneous data of device Internet of Things sensors, which can effectively reduce the workload of manual pairing and automatically complete the construction of device digital twins.

Claims

1. A digital twin automated construction method based on multi-dimensional perception data, characterized in that, It includes the following steps: S1. Use a binocular vision sensor and a lidar sensor to scan the area to be detected, and obtain the color image and the point cloud depth map of the area to be detected; S2. Use a multi-scale extraction and depth estimation module to process the color image to obtain a visual depth map based on pixel information; S3. Use the binocular vision sensor to extract color-dominated depth information and depth-dominated depth information according to the color image, the point cloud depth map and the visual depth map, construct a fused depth map, and reconstruct the three-dimensional model of the area to be detected; S4. Construct and use a DETR network to obtain target semantic information from the visual depth map; S5. Obtain heterogeneous data of Internet of Things sensors based on the area to be detected; S6. Match the three-dimensional model of the area to be detected with the target semantic information, match the matching result with the heterogeneous data of the Internet of Things sensors in the area to be detected, and construct a digital twin of the area to be detected.

2. The automated construction method of a digital twin based on multi-dimensional perception data according to claim 1, wherein, The specific implementation method of step S1 is as follows: Use a binocular vision sensor to scan the area to be detected to obtain the color image of the area to be detected; use a lidar sensor to scan the area to be detected to obtain the point cloud depth map of the area to be detected.

3. A digital twin automated construction method based on multi-dimensional perception data according to claim 1, characterized in that The specific implementation method of step S2 is as follows: S2-1. Perform the first convolution on the color image collected by the left-eye sensor of the binocular vision sensor to generate the feature map A1. Pass the feature map A1 through a 3*3 convolutional layer, and perform a correlation operation with a 40-pixel displacement on the output of the 3*3 convolutional layer and the feature map A1 to obtain a relatively large but rough-textured correlation within the range of the left-eye sensor; S2-2. Perform the second convolution on the color image collected by the left-eye sensor of the binocular vision sensor to obtain the left-eye feature map; upsample the left-eye feature map to obtain a full-resolution output, and perform a correlation operation with a 20-pixel displacement on the full-resolution output and the output result of the second convolutional layer to obtain a relatively small but fine-textured correlation within the range of the left-eye sensor; S2-3. Perform the same operations as in steps S2-1 to S2-2 on the color image collected by the right-eye sensor of the binocular vision sensor to obtain the right-eye feature map, a relatively large but rough-textured correlation within the range of the right-eye sensor, and a relatively small but fine-textured correlation within the range of the right-eye sensor; S2-4. Match the left-eye feature map and the right-eye feature map, and concatenate the matching result and the left-eye feature map to obtain the underlying semantic information; S2-5. Pass the underlying semantic information through an encoder-decoder structure, and introduce skip connections at each scale of the encoder-decoder to obtain a full-resolution disparity map estimation; S2-6. Concatenate the disparity map estimation with a relatively large but rough-textured correlation within the range of the right-eye sensor, a relatively small but fine-textured correlation within the range of the right-eye sensor, a relatively small but fine-textured correlation within the range of the left-eye sensor, and a relatively large but rough-textured correlation within the range of the left-eye sensor, and obtain the target depth estimation map through convolution operations.

4. An automated construction method of a digital twin based on multi-dimensional perception data according to claim 3, characterized in that, The specific implementation method of step S3 is as follows: S3-1. Convolve the color image, the point cloud depth map and the visual depth map to obtain the color image feature map, the point cloud depth feature map and the visual depth feature map; S3-2. Convolve the color image feature map, the point cloud depth feature map, and the visual depth feature map respectively to obtain the unique feature C of the color image E , the unique feature P of the point cloud depth map E , and the unique feature D of the visual depth map E ; S3-3. Convolve the color image feature map and the point cloud depth feature map to obtain the common feature C of the color image S and the common feature P of the point cloud depth map S ; Convolve the point cloud depth feature map and the visual depth feature map to obtain the common feature P of the point cloud depth map S and the common feature D of the visual depth map S ; S3-4. Concatenate the unique feature C of the color image E and the unique feature P of the point cloud depth map E to obtain the color-dominated unique feature C'. E Concatenate the common feature C of the color image S and the common feature P of the point cloud depth map S to obtain the color-dominated common feature C'. S ; Concatenate the unique feature D of the visual depth map E and the unique feature P of the point cloud depth map E to obtain the depth-dominated unique feature D' E ; Concatenate the common feature D of the visual depth map S and the common feature P of the point cloud depth map S to obtain the depth-dominated common feature D' S ; S3-5. According to the formula: The depth image pixel D(u, v) in the fused depth map is obtained by using a confidence aggregation strategy; where C' E (u, v) represents the unique feature points dominated by color; C' S (u, v) represents the common feature points dominated by color; D' E(u , v) represents the unique feature points dominated by depth; D' S (u, v) represents the common feature points dominated by depth; S3-6. Use the PCL triangulation algorithm to reconstruct the depth map based on the depth image pixels to obtain a three-dimensional model.

5. A method for automatically constructing a digital twin based on multi-dimensional perception data according to claim 4, characterized in that, The specific implementation method of step S4 is as follows: S4-1. Divide the target depth estimation map into N slices to obtain a set of image sequences and corresponding position encodings. S4-2. Construct and use the DETR network to process the image sequences and corresponding position encodings to obtain the target semantic information.

6. The automated construction method of a digital twin based on multi-dimensional perception data according to claim 5, wherein, The specific implementation method of step S4-2 is as follows: S4-2-1. Input the image sequences and corresponding position encodings into the self-attention module of the DETR network for encoding and decoding to obtain an output sequence. S4-2-2. Send the output sequence into the target prediction module of the DETR network to obtain the target category. S4-2-3. Add a mask head to the target category, calculate the binary mask of each target category, and generate an attention heat map. S4-2-4. Send the attention heat map into the target semantic segmentation module of the DETR network to obtain the target semantic information.

7. A method for automatically constructing a digital twin based on multi-dimensional perception data according to claim 5, characterized in that, The specific implementation method of step S6 is as follows: S6-1. Match the three-dimensional model of the area to be detected with the target semantic information, establish a device information database, and store the target semantic information and the storage path of the three-dimensional model of the area to be detected. S6-2. Classify and store the heterogeneous data of the Internet of Things sensors according to the target semantic information to complete the acquisition of device information data. S6-3. Complete the construction of the digital twin based on the device information data.

8. A digital twin automated construction method based on multi-dimensional perception data according to claim 7, characterized in that The digital twin includes a model library management module, a storage module, a digital thread module, a virtual-real mapping interface module, an inter-body interaction interface module, a control assembly module, and a service interface module. The control assembly module is connected to the model library management module, the storage module, the digital thread module, the virtual-real mapping interface module, the inter-body interaction interface module, and the service interface module.

Citation Information

Patent Citations

  • Virtual reality interaction method based on digital twinning

    CN113485392A

  • Scene flow digital twinning method and system based on dynamic trajectory flow

    CN114970321A