Computer vision-based point cloud and finite element model generation system and method
The system automates the conversion of video data into finite element models using computer vision and multi-view geometric reconstruction, addressing efficiency and precision challenges in civil engineering modeling.
Patent Information
- Application Number
- CN202510386945.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-31
AI Technical Summary
Existing methods for generating high-precision three-dimensional models in civil engineering face challenges in efficiency and cost, requiring expensive equipment and complex operations, especially for complex structures, and struggle to balance precision and dynamic model updates.
A system and method utilizing computer vision to convert video data into finite element models through video data acquisition, multi-view geometric reconstruction, and layer-based segmentation of point clouds to automate the modeling process.
Enables efficient and accurate conversion of video data into high-precision finite element models, reducing manual operations and enhancing model updating capabilities for complex structures.
Smart Images

Figure CN120318456A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of 3D modeling, and in particular, to a point cloud and finite element model generation system and method based on computer vision. Background Art
[0002] Digital Twin, as one of the core technologies of Industry 4.0 and intelligent development, has currently been widely applied in fields such as civil engineering, intelligent manufacturing, and infrastructure management. By establishing a dynamic mapping between a physical entity and its digital virtual model, Digital Twin can achieve real-time monitoring, analysis, and prediction, thereby optimizing structural performance, extending service life, and improving operation and maintenance efficiency.
[0003] In civil engineering, Digital Twin can not only provide the geometric shape of a structure but also integrate environmental data, damage information, and mechanical properties, providing a scientific basis for structural health monitoring, disaster warning, and maintenance decision-making. However, the core challenge in constructing Digital Twin lies in how to quickly and accurately generate a high-precision 3D model and continuously update its state. Traditional modeling methods rely on expensive devices such as 3D laser scanners, with high data acquisition costs and complex operations. Secondly, the modeling process is usually time-consuming and highly dependent on technical experience, especially more obvious when dealing with complex structures or large-scale projects. For example, Chinese Patent CN116611280A proposes a building structure finite element intelligent reverse modeling and analysis system based on 3D computer vision. This system automatically divides the processed structural point cloud data into each component point cloud data using the DBSCAN clustering algorithm, establishes a local coordinate system for each component to convert the component point cloud into the local coordinate system; also uses a slicing-based method to fit each slice of the component, thereby obtaining the geometric parameters of each component of the existing structure for generating a finite element model of the existing structure. This method requires obtaining the point cloud data of the existing structure through 3D scanning, with the problems of high data acquisition costs and cumbersome operations.
[0004] In addition, existing methods are difficult to balance accuracy and efficiency. High-precision modeling is often inefficient, while fast modeling may lead to the loss of geometric details and difficulty in meeting the requirements of model dynamic update. In addition, converting a 3D model into a finite element model requires a complex meshing and attribute definition process. This manual method is not only time-consuming but also prone to errors. Therefore, in order to promote the upgrade of the civil engineering industry and the high-quality development of the civil engineering industry, there is an urgent need for a fast and high-quality 3D modeling method. Summary of the Invention
[0005] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a point cloud and finite element model generation system and method based on computer vision, which can efficiently and accurately convert video data into a finite element model required for structural analysis, avoiding the cumbersome operations of manual modeling.
[0006] The object of the present invention can be achieved by the following technical solutions: A point cloud and finite element model generation system based on computer vision, including a video acquisition and processing module, a three-dimensional point cloud construction module, and a finite element model generation module connected in sequence. The video acquisition and processing module is used to acquire video data of the target structure and process the video data to obtain representative key frames and corresponding spatio-temporal information;
[0007] Based on the perspective differences between multiple key frames and the corresponding spatio-temporal information, the three-dimensional point cloud construction module uses a deep learning-based multi-view geometric reconstruction technology to reconstruct the three-dimensional point cloud of the target object and preprocess the point cloud data;
[0008] The finite element model generation module is used to extract structural spatial geometric information from the preprocessed point cloud data, cut the point cloud into multiple thin layers based on the method of tomographic scanning nodes, establish finite element components layer by layer, and integrate them to obtain a finite element model.
[0009] A point cloud and finite element model generation method based on computer vision includes the following steps:
[0010] S1. Acquire video data of the target structure and process the video data to obtain key pictures and corresponding spatio-temporal information;
[0011] S2. Based on the perspective differences between multiple key pictures and the corresponding spatio-temporal information, use a deep learning-based multi-view geometric reconstruction technology to reconstruct the three-dimensional point cloud of the target object and generate a point cloud model;
[0012] S3. Preprocess the point cloud data of the point cloud model;
[0013] S4. Extract effective structural spatial geometric information from the preprocessed point cloud data, then use the method of tomographic scanning nodes to cut the point cloud into multiple thin layers, establish each finite element component layer by layer, and integrate each finite element component to obtain a finite element model.
[0014] Further, in step S1, the video data of the target structure is specifically acquired by a mobile phone, a camera or a drone.
[0015] Further, the process of processing the video data in step S1 includes:
[0016] Extract picture nodes from the video data, apply Fourier transform to the image to obtain the frequency domain representation F(u, v) = f{f(x, y)} of the image, and calculate the proportion of the high-frequency component of Images with proportion values higher than a preset threshold are regarded as key images, where u and v are the coordinates of the image in the frequency domain, corresponding to the frequency components after Fourier transform; x and y are the coordinates of the image in the spatial domain (time domain), corresponding to the pixel positions of the original image; R is the radius threshold for dividing the high-frequency region, that is, frequency components higher than this radius are considered high-frequency components; E total is the total energy of the image, that is, the sum of the energies of all frequency components in the entire spectrum.
[0017] For key images, background nodes are removed, and the U-Net convolutional neural network structure is used to learn the boundary between the foreground and background of the image to remove the background and obtain the region of interest.
[0018] Further, step S2 specifically calculates the optical flow of each pixel of adjacent key images using the optical flow method to determine the movement of the camera, combines it with the corresponding camera pose information, and performs multi-view three-dimensional point cloud reconstruction in the coordinate system of the first frame image.
[0019] Further, the preprocessing of the point cloud data for the point cloud model in step S3 includes:
[0020] Using the median filtering algorithm and the RANSAC (Random Sample Consensus) algorithm to remove noise points;
[0021] Using the farthest point sampling algorithm for downsampling to reduce the point cloud resolution;
[0022] Using the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm to segment the point cloud to distinguish different structural components.
[0023] Further, step S4 specifically includes the following steps:
[0024] S41. For the segmented point cloud data, use the B-spline surface algorithm for fitting to obtain a fitting surface;
[0025] S42. Adopt the method of tomographic scanning nodes to cut the fitted surface into multiple tomographies, and each layer represents the cross-section of the structure within a certain height range;
[0026] S43. Use the Delaunay method to generate triangular meshes for each cut layer;
[0027] Connect the meshes of each layer to the nodes of the upper layer and the lower layer to form a complete three-dimensional finite element mesh.
[0028] Further, the specific process of step S41 is as follows:
[0029] Through the B-spline surface, a surface equation is fitted for the segmented point cloud data to determine the normal vector and position of the plane. Among them, the B-spline surface is represented by a set of control points, basis functions, and weights, and the B-spline surface is defined by the following formula:
[0030]
[0031] In the formula, u and v are coordinates in the parameter space; S(u, v) is the surface position in the u, v parameter space; P i,j is the control point, N i,p (u) and N j,q (v) are basis functions, which are B-spline basis functions defined in the u and v parameter directions respectively; p and q are the orders of the basis functions;
[0032] The least squares method is used to iteratively calculate the position of the control point P i,j so that the fitted surface is as close as possible to the actual point cloud data, and the optimization objective is
[0033] Further, the basis function is calculated using the Cox-de Boor recurrence formula. For a given parameter u, the basis function calculation formula is:
[0034]
[0035] where u i is an element of the parameter vector, and the basis function is recursively calculated.
[0036] Further, step S42 specifically cuts the object at each fixed z value for the surface S(x, y, z) = 0 to obtain multiple slices S(x, y, z i ) = 0, i = 1, 2, …, N, where z i is the height value of the i-th layer, and N is the total number of layers.
[0037] Compared with the prior art, the present invention has the following advantages:
[0038] The present invention provides a collection and processing video module, a three-dimensional point cloud construction module, and a finite element model generation module that are connected in sequence. The collection and processing video module is used to collect video data of a target structure and process the video data to obtain representative key frames and corresponding spatio-temporal information. The three-dimensional point cloud construction module is used to perform three-dimensional point cloud reconstruction on the target object by using a multi-view geometry reconstruction technology based on deep learning based on the perspective differences between multiple key frames and the corresponding spatio-temporal information, and preprocess the point cloud data. The finite element model generation module is used to extract structural spatial geometric information from the preprocessed point cloud data, and based on the method of tomography nodes, cut the point cloud into multiple thin layers, establish finite element components layer by layer, and integrate them to obtain a finite element model. Thus, by integrating video acquisition, three-dimensional reconstruction, point cloud processing, and finite element modeling technologies, the full process automation from collecting video data to establishing high-precision finite element components is realized, and the video data can be efficiently and accurately converted into a finite element model required for structural analysis, avoiding the cumbersome operations of manual modeling.
[0039] The present invention obtains structural spatial information from the captured video data, extracts clear key frames by means of Fourier transform, and uses a U-Net convolutional neural network to weight the region of interest in the image. Then, through deep learning methods, three-dimensional point cloud reconstruction is performed using the key frames, enabling the present invention to adapt to the modeling requirements of complex structures and irregular shapes, and being particularly suitable for scenarios of rapid modeling.
[0040] The present invention preprocesses the point cloud data, uses a median filtering algorithm and a RANSAC algorithm to remove noise points, uses a farthest point sampling algorithm for downsampling to reduce the point cloud resolution, and also uses a DBSCAN clustering algorithm to segment the point cloud to distinguish different components of the structure, thereby ensuring the accuracy of three-dimensional point cloud reconstruction.
[0041] Considering that the segmented point cloud data has certain geometric features, such as planar and curved surface shapes, the present invention designs a method based on B-spline surface fitting and tomography nodes to generate three-dimensional components from the point cloud data. Through an automated process, the cumbersome operations of manual modeling are avoided, greatly improving the modeling efficiency and meeting the need for rapid and continuous updating of the three-dimensional model state. Description of the Drawings
[0042] Figure 1 It is a schematic diagram of the system structure of the present invention;
[0043] Figure 2 It is a schematic diagram of the method flow of the present invention;
[0044] Figure 3 It is a schematic diagram of the application framework of the embodiment;
[0045] Figure 4Schematic diagram of a one-dimensional B-spline curve of the B-spline surface in the embodiment;
[0046] Figure 5 Schematic diagram of point cloud tomography cutting and layering modeling in the embodiment;
[0047] Explanation of the markings in the figure: 1. Acquisition and processing video module, 2. Three-dimensional point cloud establishment module, 3. Finite element model generation module. Detailed implementation mode
[0048] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0049] Embodiment
[0050] As Figure 1 shown, a point cloud and finite element model generation system based on computer vision includes an acquisition and processing video module 1, a three-dimensional point cloud establishment module 2, and a finite element model generation module 3 that are connected in sequence. Among them, the acquisition and processing video module 1 is used to acquire video data of the target structure and process the video data to obtain representative key frames and corresponding spatio-temporal information;
[0051] Based on the perspective differences between multiple key frames and the corresponding spatio-temporal information, the three-dimensional point cloud establishment module 2 uses a deep learning-based multi-view geometry reconstruction technology to reconstruct the three-dimensional point cloud of the target object and preprocess the point cloud data;
[0052] The finite element model generation module 3 is used to extract structural spatial geometric information from the preprocessed point cloud data, cut the point cloud into multiple thin layers based on the tomographic scanning nodes, establish finite element components through layering, and integrate to obtain a finite element model.
[0053] Based on the above system, a method for generating a point cloud and a finite element model based on computer vision is realized. As Figure 2 shown, it includes the following steps:
[0054] S1. Acquire video data of the target structure and process the video data to obtain key pictures and corresponding spatio-temporal information;
[0055] S2. Based on the perspective differences between multiple key pictures and the corresponding spatio-temporal information, use a deep learning-based multi-view geometry reconstruction technology to reconstruct the three-dimensional point cloud of the target object and generate a point cloud model;
[0056] S3. Preprocess the point cloud data of the point cloud model;
[0057] S4. Extract effective structural space geometric information from the preprocessed point cloud data, and then use the method of tomographic scanning nodes to cut the point cloud into multiple thin layers, establish each finite element component layer by layer, and integrate each finite element component to obtain a finite element model.
[0058] This embodiment applies the above solution to build an application framework as shown in Figure 3 . The main contents include:
[0059] I. Video acquisition and processing module: Collect video data of the target structure through lightweight devices such as mobile phones, cameras, and drones, process the video data to obtain representative key frames and corresponding spatio-temporal information, and further process the key frame images to locate and extract the interested parts;
[0060] Specifically, first use lightweight devices such as mobile phones, cameras, and drones to obtain image information of the structural space, apply the Fourier transform to the image to obtain the frequency domain representation F(u,v) = f{f(x,y)} of the image, and analyze the proportion of high-frequency components The frame with a high proportion is a clear frame. On this basis, use the U-Net convolutional neural network structure to learn the boundary between the foreground and background of the image and remove the background to obtain the region of interest.
[0061] II. 3D point cloud establishment module: Through the perspective differences and corresponding spatio-temporal information between multiple key frames, use the multi-view geometric reconstruction technology based on deep learning to perform 3D point cloud reconstruction of the target object in a unified coordinate system and preprocess the point cloud data.
[0062] Specifically, apply the optical flow method to calculate the optical flow of each pixel in adjacent key frame images, infer the movement of the camera, and combine it with the corresponding camera pose information to perform multi-view 3D point cloud reconstruction in the coordinate system of the first frame image. On this basis, use the median filter algorithm and the RANSAC algorithm to remove noise points; use the farthest point sampling algorithm for downsampling to reduce the point cloud resolution; use the DBSCAN clustering algorithm to segment the point cloud and distinguish different components of the structure.
[0063] III. Finite element model generation module: Extract effective structural space geometric information from the segmented point cloud data, and based on the idea of tomographic scanning, cut the point cloud into multiple thin layers, establish finite element components layer by layer, and finally integrate them into a finite element model.
[0064] Specifically, the segmented point cloud data usually has certain geometric features, such as plane and curved surface shapes. In this solution, the B-spline surface is used to fit the surface equation to determine the normal vector and position of the surface. As shown in
[0065] Figure 4 Figure 4As shown, a B-spline surface is represented by a set of control points, basis functions, and weights. A typical 2D B-spline surface is defined by the following formula u and v are coordinates in the parameter space; S(u, v) is the surface position in the u, v parameter space; P i,j is the control point, N i,p (u) and N j,q (v) are the basis functions, which are the B-spline basis functions defined in the u and v parameter directions respectively; p and q are the orders of the basis functions.
[0066] Among them, the basis functions are calculated using the Cox-de Boor recurrence formula. For a given parameter u, the calculation formula of the basis function is:
[0067]
[0068] where u i is an element of the parameter vector, and the basis function is obtained through recursive calculation.
[0069] Finally, the position of the control point P i,j is iteratively calculated using the least squares method, so that the fitted surface is as close as possible to the actual point cloud data, and the optimization objective is
[0070] As Figure 5 shown, this scheme is based on the idea of tomography. The fitted surface is cut into multiple thin layers (i.e., "tomograms"), and each layer represents the cross-section of the structure within a certain height range. These layers can be uniform or non-uniformly cut according to the shape of the structure. Assuming that the surface of the object is a surface S(x, y, z) = 0, the object can be cut at each fixed z value to obtain multiple slices S(x, y, z i ) = 0, i = 1, 2,..., N, where z i is the height value of the i-th layer and N is the total number of layers.
[0071] Then, the Delaunay method is used to generate a triangular mesh for each cut layer to ensure the reliable quality of the mesh elements and avoid too small angles. For the mesh of each layer, the node coordinates and shape functions of each element are calculated. For a triangular element, its node coordinates are (x1, y1), (x2, y2), (x3, y3) and the shape function is: A is the area of the triangular element. The mesh of each layer is connected to the nodes of the upper and lower layers, thus forming a complete 3D finite element mesh. For each layer z i and z i+1 , the node P(x, y, z i ) will be connected to the node P(x, y, z i+1 ).
[0072] The finally generated finite element mesh is verified and optimized to further ensure the quality, accuracy and computational efficiency of the mesh.
[0073] It should be noted that this solution is not only applicable to the modeling of steel structures, but also can be incorporated and applied to other common structures such as bamboo and wood structures, reinforced concrete, etc.
[0074] This solution first obtains high-quality image data by shooting videos of the structure and extracting frame images from the videos. Then, it constructs a point cloud model of the structure using the extracted images, fits the point cloud data through B-spline surfaces, and then generates a finite element model based on the idea of tomography. It can effectively convert video data into the finite element model required for structural analysis, avoid the cumbersome operations of manual modeling, meet the need for quickly and continuously updating the three-dimensional model state, and provide a new technical path for structural modeling in engineering practice. This solution automatically obtains the spatial geometric information of the structure based on videos, realizes the automatic generation of the end-to-end structural finite element model, simplifies the modeling process, greatly improves the modeling efficiency, and is suitable for long-term dynamic modeling and tracking monitoring of structures.
Claims
1. A point cloud and finite element model generation system based on computer vision, characterized in that, It includes a video acquisition and processing module (1), a 3D point cloud construction module (2), and a finite element model generation module (3) connected in sequence. The video acquisition and processing module (1) is used to acquire video data of the target structure and process the video data to obtain representative key frames and corresponding spatio-temporal information; The 3D point cloud construction module (2) performs 3D point cloud reconstruction on the target object by using a deep learning-based multi-view geometry reconstruction technology based on the perspective differences and corresponding spatio-temporal information between multiple key frames, and preprocesses the point cloud data; The finite element model generation module (3) is used to extract structural spatial geometric information from the preprocessed point cloud data, cut the point cloud into multiple thin layers based on the method of tomography nodes, establish finite element components layer by layer, and integrate them to obtain a finite element model.
2. A method for generating a point cloud and a finite element model based on computer vision, characterized in that, It includes the following steps: S1. Acquire video data of the target structure and process the video data to obtain key pictures and corresponding spatio-temporal information; S2. Based on the perspective differences and corresponding spatio-temporal information between multiple key pictures, perform 3D point cloud reconstruction on the target object by using a deep learning-based multi-view geometry reconstruction technology to generate a point cloud model; S3. Preprocess the point cloud data of the point cloud model; S4. Extract effective structural spatial geometric information from the preprocessed point cloud data, then use the method of tomography nodes to cut the point cloud into multiple thin layers, establish each finite element component layer by layer, and integrate each finite element component to obtain a finite element model.
3. A method for generating a point cloud and a finite element model based on computer vision according to claim 2, characterized in that, The specific operation of step S1 is to acquire video data of the target structure through a mobile phone, a camera or a drone.
4. A method for generating a point cloud and a finite element model based on computer vision according to claim 2, characterized in that, The process of processing the video data in step S1 includes: Extract image nodes from video data, apply Fourier transform to the images to obtain the frequency-domain representation F(u,v) = f{f(x,y)} of the images, and calculate the proportion of high-frequency components of Use the images with proportion values higher than the preset threshold as key images, where u and v are the coordinates of the image in the frequency domain, corresponding to the frequency components after Fourier transform; x and y are the coordinates of the image in the spatial domain, corresponding to the pixel positions of the original image; R is the radius threshold for dividing the high-frequency region, that is, the frequency components higher than this radius are considered high-frequency components; E total is the total energy of the image, that is, the sum of the energies of all frequency components in the entire spectrum; For the key pictures, remove the background nodes, use the U-Net convolutional neural network structure to learn the boundary between the foreground and background of the image, and remove the background to obtain the region of interest.
5. A method for generating a point cloud and a finite element model based on computer vision according to claim 2, wherein, The specific operation of step S2 is to calculate the optical flow of each pixel of adjacent key pictures by using the optical flow method to determine the movement of the camera, and jointly with the corresponding camera pose information, perform multi-view 3D point cloud reconstruction in the coordinate system of the first frame image.
6. A method for generating a point cloud and a finite element model based on computer vision according to claim 2, characterized in that, The preprocessing of the point cloud data of the point cloud model in step S3 includes: Use the median filtering algorithm and the RANSAC algorithm to remove the noise points; Use the farthest point sampling algorithm for downsampling to reduce the point cloud resolution; Use the DBSCAN algorithm to segment the point cloud to distinguish different components of the structure.
7. A method for generating a point cloud and a finite element model based on computer vision according to claim 6, wherein The specific operation of step S4 includes the following steps: S41. For the segmented point cloud data, perform fitting processing by using the B-spline surface algorithm to obtain a fitting surface; S42. Use the method of tomography nodes to cut the fitted surface into multiple tomographies, and each layer represents the cross-section of the structure within a certain height range; S43. Use the Delaunay method to generate a triangular mesh for each cut layer; Connect the meshes of each layer to the nodes of the upper layer and the lower layer to form a complete 3D finite element mesh.
8. A method for generating a point cloud and a finite element model based on computer vision according to claim 7, characterized in that, The specific process of step S41 is: Through the B-spline surface, a surface equation is fitted for the segmented point cloud data to determine the normal vector and position of the plane. Among them, the B-spline surface is represented by a set of control points, basis functions, and weights, and the B-spline surface is defined by the following formula: where u and v are coordinates in the parameter space; S(u, v) is the surface position in the u, v parameter space; P i,j is a control point, N i,p (u) and N j,q (v) are basis functions, which are B-spline basis functions defined in the u and v parameter directions respectively; p and q are the orders of the basis functions; The least squares method is used to iteratively calculate the position of the control point P i,j so that the fitted surface is as close as possible to the actual point cloud data, and the optimization objective is 9. A method for generating a point cloud and a finite element model based on computer vision according to claim 8, wherein, The basis function is calculated using the Cox-de Boor recurrence formula. For a given parameter u, the basis function calculation formula is: where, u i is an element of the parameter vector, and the basis function is calculated recursively.
10. A method for generating a point cloud and a finite element model based on computer vision according to claim 7, characterized in that, Specifically, in step S42, for the surface S(x, y, z) = 0, the object is cut at each fixed z value to obtain multiple slices S(x, y, z i ) = 0, i = 1, 2, …, N, where z i is the height value of the i-th layer and N is the total number of layers.
Citation Information
Patent Citations
Finite element preprocessing method for reconstructing three-dimensional entity model
CN104282040A
Grid curved surface reconstruction method for scene understanding
CN110009743A
Parametric modeling method for planing boat
CN111709086A
Gastrointestinal tract capsule endoscopy video key frame extraction method with self-adaptive threshold value
CN113850299A
Building structure finite element intelligent reverse modeling and analysis system based on three-dimensional computer vision
CN116611280A
Cited By
Pipeline reverse modeling system and method based on point cloud data
CN122115743A