Fast visualization method of airport 3D maps based on deep learning

Through low-cost monocular cameras and deep learning technology, the problem of expensive equipment and difficult reconstruction is solved, and fast and low-cost three-dimensional map reconstruction of airports is realized, suitable for weak texture scenarios.

CN115471623BActive Publication Date: 2025-08-15SOUTHEAST UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211147684.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-19
Publication Date
2025-08-15
Estimated Expiration
2042-09-19

AI Technical Summary

Technical Problem

The existing three-dimensional model establishment methods are expensive, labor-intensive and material-intensive, and are difficult and time-consuming to reconstruct in weak texture scenarios.

Method used

A low-cost monocular camera is used to collect image data, combined with deep learning technology, and three-dimensional visualization is achieved through image feature point extraction, pose solution, depth map estimation and fusion.

Benefits of technology

Rapidly reconstruct dense three-dimensional models in unknown environments, reduce equipment costs, improve reconstruction efficiency, and adapt to weak texture scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115471623B_ABST
    Figure CN115471623B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for rapid visualization of airport three-dimensional maps based on deep learning. In an unknown environment, the carrier uses a monocular camera to collect color images of the surrounding environment during movement, extracts feature points from the image, aligns and calculates its own motion transformation, and then inputs the image and posture into a deep learning network to obtain a depth map of each image. Finally, a dense three-dimensional visualization map is obtained by fusing all the points in the depth map. The present invention provides a rapid three-dimensional visualization method based on deep learning. In order to solve the problem that the existing three-dimensional model establishment method is expensive and requires a lot of manpower and material costs, a low-cost camera is used to replace the expensive equipment. It only needs to collect enough image data to reconstruct the three-dimensional map model. In order to solve the problem that it is difficult for cameras to reconstruct scenes with weak textures and the reconstruction time is long, deep learning is introduced to quickly obtain dense three-dimensional visualization models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of rapid map visualization, and particularly proposes a rapid visualization method for airport three-dimensional maps based on deep learning. Background Art

[0002] Airports are a vital component of large and medium-sized cities. In recent years, the rapid development of the civil aviation industry has put immense pressure on airport and airspace capacity. Furthermore, airport operations are becoming increasingly complex, and passenger service demands are becoming more diverse and personalized. To improve operational efficiency and service quality, airports must fully utilize modern technology and advanced facilities. This has given rise to the concept of smart airports. The Smart Airport Visual Decision Platform comprehensively displays key data from core airport operational systems, supporting daily operational monitoring and management across multiple dimensions, including airport infrastructure, ground services, flight operations, surface supervision, and aircraft stand management, as well as emergency command and dispatch management in the event of an emergency. It provides users with an intelligent operations management platform that integrates operational monitoring, decision support, command and dispatch, and information dissemination, providing data-driven decision-making support for managers to improve airport operational efficiency and management efficiency.

[0003] Building Information Modeling (BIM) is a new tool in architecture, engineering, and civil engineering. BIM is a term used to describe computer-aided design (CAD) systems that utilize 3D graphics and are object-oriented. BIM (Building Information Modeling) technology, first proposed by Autodesk in 2002, has gained widespread industry recognition worldwide. It helps integrate building information, integrating all aspects of a building's design, construction, operation, and lifecycle into a single 3D model database. This allows design teams, construction companies, facility operators, and owners to collaborate based on BIM, effectively improving efficiency, conserving resources, reducing costs, and ultimately achieving sustainable development. BIM-based 3D visualization allows for a more intuitive understanding of the built environment, enhancing the overall realism and experience. 3D visualization technology utilizes computer graphics and image processing techniques to convert data drawings into graphics or images for on-screen display, allowing users to interact with the building as if they were actually there. The rapidly developing VR technology is also based on this visualization technology. However, BIM 3D visualization solutions for large scenes require a lot of money and time. Collecting image data through cameras and using multi-view geometry (MVS) methods for 3D reconstruction can greatly reduce costs and quickly achieve 3D visualization.

[0004] Reconstructing 3D geometry from photographs is a classic computer vision problem that has occupied researchers for over 30 years. Applications range from 3D mapping and navigation to online shopping, 3D printing, computational photography, computer video games, and cultural heritage archiving. However, only recently have these techniques matured sufficiently to move beyond the controlled environment of the laboratory and into the field, providing robustness, accuracy, and scalability at an industrial scale. Multi-view geometry (MVS) is a technique for recovering the dense structure of a scene from multiple overlapping viewpoints. Traditional 3D reconstruction algorithms based on multi-view geometry have long dominated the field. These methods use methods such as photometric consistency to compute dense 3D information. While these methods achieve high depth estimation accuracy in ideal scenarios, they also suffer from common limitations, such as weak textures, high reflectivity, and repetitive textures, which make reconstruction difficult or incomplete. Consequently, traditional 3D reconstruction methods still have significant room for improvement in terms of reconstruction completeness.

[0005] In recent years, artificial intelligence technologies, represented by deep learning, have experienced rapid development in the field of computer vision and have been widely applied to tasks such as object detection, image segmentation, and image classification. Leveraging artificial intelligence to empower classic research problems has become one of the hottest development directions. Recent studies have demonstrated that deep learning-based image depth estimation techniques can effectively improve the quality of depth estimation. The recent success of convolutional neural network (CNN) research has also sparked interest in improving stereo reconstruction. Conceptually, learning-based methods can incorporate global semantic information, such as specular and reflection priors, to achieve more robust matching. Currently, deep learning-based MVS reconstruction methods have surpassed traditional methods in accuracy and completeness, becoming the most advanced method in this field. Furthermore, they consume far less computational resources and require far less time to reconstruct the model than traditional methods.

[0006] After searching the existing technical literature, it was found that:

[0007] Compared to the Chinese patent application (application number: CN201510401528.4), a Beidou ground-based navigation network operation and management system. Technical comparison: The "Beidou ground-based navigation network operation and management system" is used outdoors, such as in mountainous areas and cities where Beidou signals are weak; this article's application context is indoors, such as in airports. Furthermore, the "Beidou ground-based navigation network operation and management system" proposes a signal management system that is connected by pseudo-satellites, monitoring stations, and a ground control center. The operation and management system includes, in sequence, a ground-based network signal receiving antenna, a signal receiving module, a signal processing module, a control analysis and management module, a computer, analysis software, and multiple hardware and software components connected to the control analysis and management module. This optimizes the signal from a macro perspective. This article proposes processing and analyzing the signal source, making judgments through various means such as signal quality monitoring, data quality monitoring, observation quality monitoring, standard deviation and mean monitoring, and message range monitoring, with a greater reliance on algorithms.

[0008] Compared with the Chinese patent (application number: CN202111328313.6) A method, device and related components for monitoring the integrity and health status of pseudo-satellites. Technical comparison: "A method, device and related components for monitoring the integrity and health status of pseudo-satellites" mainly monitors the health status of pseudo-satellites, such as measuring the temperature, voltage and other parameters of pseudo-satellites; this article monitors the signals emitted by pseudo-satellites, so the detection targets of the two are different. In addition, two of the signal parameters mentioned in "A method, device and related components for monitoring the integrity and health status of pseudo-satellites" are signal-to-noise ratio and carrier phase. This article judges five aspects, including signal quality, data quality, observation quality, standard deviation and mean, and message range. Both the monitoring methods and parameter types are different. Summary of the Invention

[0009] This invention aims to provide a method for rapidly visualizing airport 3D maps based on deep learning. This method addresses the challenges of existing 3D modeling methods, which require expensive equipment and significant human and material resources. By replacing this expensive equipment with low-cost cameras, only sufficient image data needs to be collected to reconstruct a 3D map model. Furthermore, deep learning is introduced to rapidly generate dense 3D visualization models, addressing the difficulties cameras face in reconstructing scenes with weak textures and the long reconstruction time.

[0010] To achieve the above object, the technical solution adopted by the present invention is:

[0011] A method for rapid visualization of airport three-dimensional maps based on deep learning is characterized by comprising the following steps:

[0012] S1. Transmit the color image acquired by the monocular camera to the processor to calculate the displacement and posture information of the carrier;

[0013] The step S1 includes the following steps:

[0014] S11. Image feature point extraction;

[0015] S12. Match feature points and calculate position and posture;

[0016] The step S11 includes the following steps: extracting SIFT feature points from the image and matching them to obtain corresponding feature points in different images;

[0017] The step S12 includes the following steps: calculating the basic matrix F, solving the essential matrix E using the basic matrix F and the camera internal parameters, and decomposing the essential matrix E to obtain the position t and posture R of the camera:

[0018] S2. Input the color image and displacement and posture information into the deep learning network to obtain a depth map for each image;

[0019] The step S2 includes the following steps:

[0020] S21. Input multi-view images into the designed end-to-end deep learning architecture;

[0021] S3. Fusing the 3D points in the depth map based on displacement and posture information to obtain a dense 3D visualization map;

[0022] The step S3 includes the following steps:

[0023] S31. Convert the points in the depth map to the world coordinate system according to the position and posture;

[0024] S32. Fuse the converted points to obtain a three-dimensional visualization map.

[0025] As a further improvement of the present invention, the image feature point extraction in step S11 is specifically as follows;

[0026] The SIFT algorithm is used to extract feature points. In the matching process, the SIFT algorithm uses the Kd-tree algorithm, that is, to find the point B closest to the feature point A of the reference image and the point C with the second closest distance in the image to be registered. The formula for determining whether feature points A and B are a pair of matching points is:

[0027]

[0028] Where d(A, B) is the distance between A and the point B closest to A, and d(A, C) is the distance between A and the point C second closest to A. The RANSAC algorithm is then used to eliminate incorrect matching points. At this point, the SIFT feature point extraction and matching work has been completed.

[0029] As a further improvement of the present invention, the feature points are matched in step S21 and the position and posture are calculated as follows;

[0030] Pose solution;

[0031] The coordinate system of the first image is used as the world coordinate system. For each image, the next frame is selected to form an image pair. The basic matrix F between the two images is calculated through the image pair. When there are enough corresponding feature points on the two images, it can be calculated by the following equation:

[0032] p T Fq=0

[0033] Where F is the fundamental matrix to be determined, p is the feature point of the left image, and q is the corresponding feature point of the right image;

[0034] The essential matrix E between the two images is obtained through the basic matrix:

[0035]

[0036] Where K1 is the camera intrinsic parameter of the left image, K2 is the camera intrinsic parameter of the right image, and E is the essential matrix to be determined;

[0037] Finally, the position and attitude of the camera are decomposed from the essential matrix, and the SVD decomposition of the essential matrix is performed:

[0038] SVD(E)=UDV r

[0039] t=U

[0040] R=UWV r

[0041] Where t is the camera position, R is the camera pose,

[0042] As a further improvement of the present invention, step S3 fuses the three-dimensional points in the depth map according to the displacement and posture information to obtain a dense three-dimensional visualization map. The specific steps are as follows:

[0043] 1) Deep feature extraction;

[0044] After the perspective is selected, N paired images are input, namely the reference image and the candidate set. First, an eight-layer two-dimensional convolutional neural network is used to extract the depth feature F of the stereo pair and output a 32-channel feature map.

[0045] 2) Depth estimation;

[0046] The depth estimation of MVSNet is learned directly through neural networks.

[0047] The network training method is to input the cost volume V and the corresponding depth map true value, use SoftMax to regress the probability of each pixel at depth θ, and obtain a probability volume P representing the confidence of each image along the depth direction of the reference image to complete the learning process from cost to depth value. The depth value with the highest confidence is the depth of the current pixel.

[0048] 3) Depth map fusion;

[0049] Outlier Removal

[0050] For each depth point in an image, if it is similar in depth to the corresponding point in other images, the point is considered accurate. If it is dissimilar to all images, it is an outlier and is removed:

[0051] p=KT ii q

[0052] Where q is the coordinate of the point in the current image i, p is the coordinate of the point projected to the image j, K is the camera internal parameter, T ij is the pose between image i and image j;

[0053] Depth map fusion

[0054] Project the points in each image into a point cloud in world coordinates:

[0055] q w =KT i q

[0056] Where T i is the pose of the i-th image to the world coordinate system, q is the point in the current image coordinate system, q w The point converted to the world coordinate system.

[0057] Compared with the prior art, the present invention has the following beneficial effects:

[0058] This invention uses only a low-cost monocular camera to collect data in unknown environments to rapidly reconstruct a 3D model. This addresses the challenges of existing 3D modeling methods, which require expensive equipment and significant human and material resources. By replacing this expensive equipment with a low-cost camera, only sufficient image data is required to reconstruct a 3D map model. Furthermore, to address the difficulties cameras face in reconstructing scenes with weak textures and the long reconstruction time, deep learning is introduced to rapidly acquire dense 3D visualization models. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a schematic diagram of the network architecture of the deep learning model designed by the present invention;

[0060] Figure 2It is a flow chart of the airport rapid visualization method. DETAILED DESCRIPTION

[0061] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0062] As a specific embodiment of the present invention, the present invention proposes a method for rapid visualization of airport three-dimensional maps based on deep learning, wherein the network architecture diagram of the deep learning model designed by the present invention is as follows Figure 1 As shown in the figure, the flow chart of the airport rapid visualization method is as follows Figure 2 The specific steps are as follows:

[0063] (1) Image pose estimation

[0064] 1. Image feature extraction. The scale-invariant feature transform (SIFT) algorithm is used to extract feature points. During the matching process, the SIFT algorithm uses the Kd-tree algorithm, which finds the point B closest to the reference image feature point A and the point C with the second closest distance. The formula for determining whether feature points A and B are a pair of matching points is:

[0065]

[0066] Where d(A, B) is the distance between point A and point B, which is closest to A, and d(A, C) is the distance between point A and point C, which is the second closest to A. We then use the RANSAC algorithm to eliminate incorrect matching points, completing the SIFT feature point extraction and matching process. RANSAC is a random parameter estimation algorithm.

[0067] 2. Pose calculation

[0068] The coordinate system of the first image is used as the world coordinate system. For each image, the next frame is selected to form an image pair. The basic matrix F between the two images is calculated through the image pair. When there are enough corresponding feature points on the two images, it can be calculated by the following equation:

[0069] p T Fq=0

[0070] Where F is the fundamental matrix to be determined, p is the feature point of the left image, and q is the corresponding feature point of the right image.

[0071] The essential matrix E between the two images can be obtained through the basic matrix:

[0072]

[0073] Where K1 is the camera intrinsic parameter of the left image, K2 is the camera intrinsic parameter of the right image, and E is the essential matrix to be determined.

[0074] Finally, the position and attitude of the camera are decomposed from the essential matrix, and the SVD decomposition of the essential matrix is performed:

[0075] SVD(E)=UDV r

[0076] t=U

[0077] R=UWV r

[0078] Where t is the camera position, R is the camera pose,

[0079] (2) Depth Map Estimation Based on Deep Learning

[0080] This patent designs an end-to-end deep learning architecture that takes images and poses as input and outputs a depth map for each image. The network first extracts deep visual image features, then constructs a 3D cost volume through differentiable projective transformations. Next, regularization is performed to output a 3D probability volume, which is then passed through a soft argMin layer to calculate the expected depth along the depth direction, yielding a depth map for the reference image.

[0081] 1. Deep Feature Extraction

[0082] After the perspective is selected, N paired images are input, namely the reference image and the candidate set. First, an eight-layer two-dimensional convolutional neural network is used to extract the depth feature F of the stereo pair and output a 32-channel feature map.

[0083] 2. Depth Estimation

[0084] Depth estimation is learned directly through a neural network. The network training method for MVSNet is to input a cost volume V and the corresponding ground truth depth map. Using SoftMax regression, the probability of each pixel at depth θ is regressed. This yields a probability volume P representing the confidence level of each reference image along the depth direction. This completes the learning process from cost to depth value. The depth value with the highest confidence is the depth of the current pixel.

[0085] (3) Depth map fusion

[0086] 1. Outlier removal

[0087] For each depth point in an image, if it is similar in depth to the corresponding point in other images, the point is considered accurate. If it is dissimilar to all images, it is an outlier and is removed:

[0088] p=KT ij q

[0089] Where q is the coordinate of the point in the current image i, p is the coordinate of the point projected to the image j, K is the camera internal parameter, T ij is the pose between image i and image j.

[0090] 2. Depth Map Fusion

[0091] Project the points in each image into a point cloud in world coordinates:

[0092] q W =KT i q

[0093] Where T i is the pose of the i-th image to the world coordinate system, q is the point in the current image coordinate system, q w The point converted to the world coordinate system.

[0094] The point clouds projected to the world coordinate system are fused into a dense point cloud to obtain a three-dimensional visual map.

[0095] The above description is merely a preferred embodiment of the present invention and does not constitute any other form of limitation to the present invention. Any modification or equivalent variation based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A fast visualization method for airport 3D maps based on deep learning, characterized by: The steps include: S1. Transmits the color image captured by the monocular camera to the processor, which calculates the displacement and posture information of the carrier. The step S1 includes the following steps: S11. Image feature point extraction; S12. Feature point matching and calculation of position and posture; The step S11 includes the following steps: extracting SIFT feature points from the image and matching them to obtain corresponding feature points in different images; The step S12 includes the following steps: calculating the basic matrix F, solving the essential matrix E using the basic matrix F and the camera internal parameters, and decomposing the essential matrix E to obtain the position t and posture R of the camera: S2. Input the color image and displacement and pose information into the deep learning network to obtain a depth map for each image. The step S2 includes the following steps: S21. Input multi-view images into the designed end-to-end deep learning architecture; S3. Fuse the 3D points in the depth map based on displacement and pose information to obtain a dense 3D visualization map. The step S3 includes the following steps: S31. Convert the points in the depth map to the world coordinate system according to the position and posture; S32. Fusing the converted points to obtain a three-dimensional visual map; The step S3 fuses the three-dimensional points in the depth map according to the displacement and posture information to obtain a dense three-dimensional visualization map. The specific steps are as follows: 1) Deep feature extraction; After the perspective is selected, N paired images, namely the reference image and the candidate set, are input. First, an eight-layer two-dimensional convolutional neural network is used to extract the deep features F of the stereo pair and output a 32-channel feature map. 2) Depth estimation; The depth estimation of MVSNet is learned directly through neural networks. The network training method is to input the cost volume V and the corresponding depth map true value, use SoftMax to regress the probability of each pixel at depth θ, and obtain a probability volume P representing the confidence of each image along the depth direction of the reference image to complete the learning process from cost to depth value. The depth value with the highest confidence is the depth of the current pixel; 3) Depth map fusion; Outlier Removal For each depth point in an image, if it is similar in depth to the corresponding point in other images, the point is considered accurate. If it is dissimilar to all images, it is an outlier and is removed: ; Where q is the coordinate of the point in the current image i, and p is the coordinate of the point projected to image j. is the camera internal parameter, is the pose between image i and image j; Depth map fusion Project the points in each image into a point cloud in world coordinates: ; in is the pose of the i-th image to the world coordinate system, is the point in the current image coordinate system, The point converted to the world coordinate system.

2. The method for rapid visualization of airport 3D maps based on deep learning according to claim 1, characterized in that: The image feature point extraction in step S11 is specifically as follows: The SIFT algorithm is used to extract feature points. In the matching process, the SIFT algorithm uses the Kd-tree algorithm, that is, to find the point B closest to the feature point A of the reference image and the point C with the second closest distance in the image to be registered. The formula for determining whether feature points A and B are a pair of matching points is: ; in is the distance between A and the point B closest to A, is the distance between A and the point C that is the second closest to A. Then the RANSAC algorithm is used to eliminate the wrong matching points. At this point, the SIFT feature point extraction and matching work has been completed.

3. The method for rapid visualization of airport 3D maps based on deep learning according to claim 1, characterized in that: In step S21, the feature points are matched and the position and posture are calculated as follows; Pose solution; The coordinate system of the first image is used as the world coordinate system. For each image, the next frame is selected to form an image pair. The basic matrix F between the two images is calculated through the image pair. When there are enough corresponding feature points on the two images, it can be calculated by the following equation: ; in is the fundamental matrix to be determined, are the feature points of the left image, is the feature point corresponding to the right image; The essential matrix E between the two images is obtained through the basic matrix: ; in is the camera intrinsic parameter of the left image, is the camera intrinsic parameter of the right image, is the essential matrix to be determined; Finally, the position and attitude of the camera are decomposed from the essential matrix, and the SVD decomposition of the essential matrix is performed: ; ; ; in is the camera position, is the camera pose, .

Citation Information

Patent Citations

  • Beidou foundation navigation network operation and managing system

    CN105119648A

  • Pseudolite integrity and health state monitoring method and device and related components

    CN114035153A

  • Three-dimensional real-time map construction method and device

    CN111105498A

  • Instant positioning and map construction system and method with semantic perception

    CN111968129A