Method for information extraction and three-dimensional reconstruction based on hyperstereo pair
By employing a method for information extraction and 3D reconstruction of ultra-generalized stereo image pairs, utilizing perspective instance segmentation and observation angle discrimination networks in tilted remote sensing images, and combining information repair strategies, the limitations of existing stereo image pair modeling methods are overcome, achieving high-quality 3D reconstruction and information extraction of buildings.
Patent Information
- Application Number
- CN202211616806.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-15
AI Technical Summary
Existing stereo image pair modeling methods, when not meeting the constraints of standard stereo image pairs, struggle to reconstruct the overall structure of a building with high quality from a small number of local points, making it difficult to extract missing information from the blind spots of orthophoto images.
A method for information extraction and 3D reconstruction based on ultra-generalized stereo image pairs is adopted. Perspective instance segmentation is performed by acquiring tilted remote sensing images of buildings. The main perspective instance segmentation network and the observation angle discrimination network are used for matching. Image inpainting and image outpainting strategies are combined to repair missing information, and finally the 3D reconstruction of buildings is achieved.
It enables high-quality reconstruction of the overall structure of a building using a small number of local points even when the standard stereo image pair constraint is not met. It comprehensively solves the problem of building information extraction and 3D reconstruction under the broad input conditions of ultra-generalized stereo image pairs, and can more easily extract missing information in the blind spots of orthophoto images.
Smart Images

Figure CN116258824B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of trajectory planning, in particular to an information extraction and three-dimensional reconstruction method based on hyper-stereo image pairs. BACKGROUND
[0002] For a long time, building information extraction and three-dimensional reconstruction technology based on visible light remote sensing images has a wide range of application needs in the field of people's livelihood and national defense. In the traditional sense, three-dimensional reconstruction methods based on visible light remote sensing images mainly include three categories according to the different composition types of input data: (1) stereo pair-based method. Generally composed of two or three images that meet certain intersection angles, overlap degrees, and base-height ratios, etc. from spaceborne or airborne; (2) oblique photogrammetry system-based method. Usually composed of five aerial images of fixed observation angles, including one orthographic image and four oblique images of different azimuths; (3) video image-based method. Usually a sequence of aerial images taken continuously along a certain flight line. In order to obtain a relatively ideal three-dimensional reconstruction effect, the input images usually need to meet certain limited conditions, such as limiting the number of images, the observation angle difference and the resolution difference range between multiple images. The combination of two or more visible light images that meet these conditions is collectively referred to as a "standard stereo pair".
[0003] At present, the three-dimensional reconstruction method for the standard stereo pair needs to solve the three-dimensional coordinates of a certain amount of point clouds that can constrain the building surface structure or estimate the scene depth map according to the matching relationship of the same name points between multiple images, and then gradually model to generate a three-dimensional model of the building. However, this three-dimensional information extraction method based on pixel-level processing largely depends on the number and position distribution of correctly matched same name points. Once the input images do not meet the limited conditions of the standard stereo pair, it will be limited by the technical bottleneck of stereo matching, and it is difficult to reconstruct the overall structure of the building with high quality from a small number of local points. Some scholars have proposed the concept of "generalized stereo pair", but the "generalized stereo pair" used in related research is mainly different from the "standard stereo pair" in that the images come from different satellites, and there is a small resolution difference between the images. It does not essentially reduce the requirements for input images, and the core technology used is still the traditional processing method. In addition, whether it is a "standard stereo pair" or a "generalized stereo pair", they generally pay more attention to the accuracy of the three-dimensional coordinates of the overall scene and terrain, lack comprehensive analysis of the three-dimensional structure and surface texture information of each building, and fail to further tap a large amount of missing information in the blind area of the traditional orthographic image perspective. This information is crucial to the accuracy and completeness of building three-dimensional reconstruction. Therefore, the current stereo pair modeling method still has the problem that it is difficult to reconstruct the overall structure of the building with high quality from a small number of local points when it does not meet the limited conditions of the standard stereo pair, which leads to the difficulty in extracting missing information in the blind area of the orthographic image perspective. SUMMARY
[0004] The present application aims to solve the problem that the existing modeling method of stereo pairs cannot meet the standard stereo pair conditions, and it is difficult to reconstruct the overall structure of a building with a small number of local points, resulting in the difficulty of extracting missing information in the perspective blind area of the orthographic image.
[0005] The specific process of the information extraction and three-dimensional reconstruction method based on the super-generalized stereo pair is as follows:
[0006] Step one, obtain building oblique remote sensing images, and perform perspective instance segmentation on the building oblique remote sensing images to obtain perspective instance segmentation results;
[0007] Step two, information repair is performed on the perspective instance segmentation results obtained in step one to obtain single building images;
[0008] Step three, match the same repaired single building images from different building oblique remote sensing images;
[0009] Step four, based on the matching results of step three, use the image of each single building to realize super-generalized stereo pair three-dimensional reconstruction;
[0010] The super-generalized stereo pair is any oblique remote sensing image covering a certain area.
[0011] Further, the step one of obtaining building oblique remote sensing images and performing perspective instance segmentation on the building oblique remote sensing images to obtain perspective instance segmentation results comprises the following steps:
[0012] Step one, based on the virtual knowledge of building oblique remote sensing images, simulate the oblique remote sensing images to obtain oblique remote sensing simulation images;
[0013] Step two, build a main perspective instance segmentation network, and use the oblique remote sensing simulation images and the main perspective instance segmentation network to obtain the perspective instance segmentation results, specifically:
[0014] Step two one, label the oblique remote sensing simulation images;
[0015] The label includes: building overall instance, building top instance, occlusion instance, and background.
[0016] Step two two, build a main perspective instance segmentation network, and use the labeled oblique remote sensing simulation images and building oblique remote sensing images to pre-train the main perspective instance segmentation network to obtain a pre-trained model;
[0017] Step two three, perform transfer learning using the building oblique remote sensing images and the pre-trained model to obtain the perspective instance segmentation results.
[0018] Further, the virtual knowledge of the building tilt remote sensing image is obtained by the following formula:
[0019] K sim ={S Image ,S Label}
[0020] S Image =P(V _sim {Terrain,Object,Texture _Sim})
[0021] S Label =P(V _Label {Terrain,Object,Texture _Label}) (1)
[0022] Wherein, S Image is a building tilt remote sensing simulation image, S Label is the true value corresponding to S Image , Terrain, Object, Texture is the terrain, object and texture of the virtual scene, P is the imaging projection model, V _sim and V _Label are the results of mapping the same simulation three-dimensional scene with simulation texture and label texture respectively, Texture _sim and Texture _Lable are simulation texture and label texture respectively.
[0023] Further, the main perspective instance segmentation network comprises: an FCOS-based target detector layer, a convolution layer, a feature pyramid layer, a feature decoder layer, a mask prediction layer, an occlusion classification layer and an output layer.
[0024] The FCOS-based target detector layer is used to obtain the detection box area of the building whole, top surface and facade in the building tilt remote sensing image.
[0025] The convolution layer is used to obtain the attention map r k of each instance according to the detection box area of the building whole, top surface and facade.
[0026] The attention map of each instance comprises the attention map of the building whole instance, the attention map of the building top surface instance and the attention map of the occlusion instance; the feature pyramid is used to obtain the feature of the building tilt remote sensing image.
[0027] The feature decoder layer is used to obtain the base s k of each instance according to the feature of the building tilt remote sensing image.
[0028] The mask prediction layer is used to predict the base s k and the attention map r k The mask m d is obtained by merging based on the BlendMask mixed mask strategy.
[0029] The occlusion classification layer is used to determine whether there is an occlusion area in the image within the detection frame region of each building as a whole. For buildings with occlusions, the mask m occ of the invisible area is obtained.
[0030] The output layer is used to obtain the comprehensive mask by merging m d and m occ , and output the perspective instance segmentation result.
[0031] Further, the base s k and the attention map r k of each instance are merged to obtain the predicted mask m d based on the BlendMask mixed mask strategy, as follows:
[0032]
[0033] where k is the base number of each instance, K is the set number of bases, represents the element-wise product, d is the detection frame number, and D is the total number of detection frames.
[0034] Further, the perspective instance segmentation result obtained in step one is repaired in step two to obtain a single building image, specifically:
[0035] Step two one, based on the perspective instance segmentation result obtained in step one, the invisible mask area S occ of each building in the building inclined remote sensing image is obtained, and the building overall area mask area S bui is obtained, and the ratio Z s of S occ to S bui is obtained.
[0036] Step two two, set a proportion threshold T, compare T with Z s , if Z s is less than or equal to the threshold T, use the ImageInpainting strategy to repair the local loss to obtain a single building image; if Z s is greater than the threshold T, use the ImageOutpainting strategy to repair the local loss to obtain a single building image.
[0037] Further, the step three of matching the same repaired single building image from different building oblique remote sensing images comprises the following steps:
[0038] Step three one, obtaining the overall, top surface and facade area pixels of each single building respectively;
[0039] Step three two, establishing an observation angle judgment network, inputting the overall, top surface and facade area pixels of each building obtained in step three one into the observation angle judgment network to obtain the observation angle of each building;
[0040] The observation angle judgment network adopts a ResNet-50 network for feature extraction;
[0041] The observation angle of each building comprises a quasi-forward direction, a quasi-lateral direction and an oblique direction;
[0042] Step three three, matching the same single building image from different building oblique remote sensing images based on the observation angle of each building to obtain a matching result.
[0043] Further, the observation angle judgment network loss function is as follows:
[0044] L total =λ w L w +λ t L t +λ s L s (4)
[0045] Wherein, L w , L t and L s are the loss functions of the overall, top surface and facade of the building, λ w , λ t and λ s are regularization coefficients.
[0046] Further, the step three of matching the same repaired single building image from different building oblique remote sensing images comprises the following steps:
[0047] Step three one, obtaining the overall, top surface and facade area pixels of each single building respectively;
[0048] Step three, the matching network of target level is used to classify the matching in the same kind of observation angle data set and the different kind of observation angle data set, the matching result is obtained, that is, all images of the same building, then step three is executed; the shooting time of the single building image without matching object is compared with the record time in the database, the newly added building single image is obtained by screening, the newly added building single image is saved to the database, and step four is directly executed;
[0049] Step three, based on the matching result obtained in step three, the original building tilt remote sensing image where each building is located is restored, then Nm buildings in the neighborhood are selected based on each building as the center, the top surface centroid coordinates of each building are obtained as the center position of the building, the distance and the relative orientation between two centroids are obtained according to the ground resolution and the observation depression angle, that is, the distance and the relative orientation between buildings, the wrong matching in step three is excluded according to the distance and the relative orientation between buildings, the single building image without matching is added to the database to execute step four;
[0050] The wrong matching is: the wrong matching caused by high image feature similarity and one-to-many.
[0051] Further, the shooting time of the single building image without matching object is compared with the record time in the database in step three, and the newly added building single image is obtained by screening.
[0052] The shooting time of the single building image without matching object is compared with the record time in the database, if the shooting time of the single building image with matching object is earlier than the record time in the database, it is a removed building; if the shooting time of the single building image with matching object is later than the record time in the database, it is a newly added building.
[0053] The beneficial effects of the present application are:
[0054] The present application proposes a new building three-dimensional reconstruction technical framework for super-general stereo pairs, discards the three-dimensional information extraction method based on pixel level processing, adopts the strategy based on target level processing, directly reconstructs the three-dimensional structure based on the image features of the building, takes the three-dimensional reconstruction technology based on a single tilt remote sensing image as the core architecture, cooperates with the three-dimensional information fusion technology, copes with the characteristics of uncertain number of images and newly added images at any time, realizes the three-dimensional reconstruction based on super-general stereo pairs, and the present application can reconstruct the overall structure of the building with a small amount of local points in high quality when the image does not meet the limited conditions of the standard stereo pairs, so that the missing information in the perspective blind area of the orthographic image can be extracted more easily. The present application comprehensively solves various problems in building information extraction and three-dimensional reconstruction caused by the wide input conditions of super-general stereo pairs. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 This is a flowchart of the present invention;
[0056] Figure 2 This is a flowchart for information repair.
[0057] Figure 3 Flowchart of target-level matching technology;
[0058] Figure 4 It is a super-generalized stereo image pair. Detailed Implementation
[0059] Specific implementation method one: as follows Figure 1 As shown, the specific process of the information extraction and 3D reconstruction method based on ultra-generalized stereo image pairs in this embodiment is as follows:
[0060] Step 1: Acquire remote sensing images of the building's tilt and perform perspective instance segmentation on these images to obtain the perspective instance segmentation results.
[0061] Step 11: Using virtual knowledge of building tilt remote sensing images, simulate tilt remote sensing images to obtain tilt simulation images:
[0062] The virtual knowledge of the building's tilted remote sensing image is obtained through the following methods:
[0063] K sim ={S Image ,S Label}
[0064] S Image =P(V _sim {Terrain,Object,Texture _Sim})
[0065] S Label =P(V _Label {Terrain,Object,Texture _Label})(1)
[0066] Among them, S Image It is a remote sensing simulation image of a tilted building, S Label It is S Image The corresponding Ground-Truth is: Terrain, Object, and Texture represent the terrain, features, and textures of the virtual scene; P represents the imaging projection model; and V represents the ground truth. _sim and V _Label It is the result of mapping the same simulated 3D scene using simulated textures and labeled textures respectively; the virtual scene is artificially set, Texture _sim and Texture_Lable are simulation texture and label texture respectively;
[0067] As shown in formula (1), the virtual knowledge K sim is composed of two parts of simulation image (S Image ) and corresponding Ground-Truth (S Label ). Both are formed by the same projection P processing on simulation scene V _sim and V _Label , V _sim and V _Label are obtained by mapping the same simulation three-dimensional scene with simulation texture and label texture respectively. Therefore, only the terrain Terrain, object Object and texture Texture for constructing the simulation scene are needed to construct V _sim and V _Label , and then the corresponding S Image and S Label are obtained according to the imaging projection model P to form the virtual knowledge K sim . Programmable related modules of many software platforms such as OPENCV, MATLAB and the like can realize the related functions.
[0068] Step one two, build the main perspective instance segmentation network, and obtain the perspective instance segmentation result by using the oblique remote sensing simulation image and the main perspective instance segmentation network:
[0069] Step one two one, label the oblique remote sensing simulation image, including: building whole instance, building top instance, shielding instance, and other background;
[0070] Considering that the facade features of the building in the remote sensing image are greatly affected by the observation angle and are easily shielded, it is not suitable to be extracted as an instance alone. In order to distinguish the top and facade, the scene is labeled in the following way: building whole instance, building top instance, shielding instance, and other background. Among them, shielding may exist in two cases of background shielding and mutual shielding between buildings, which are both regarded as shielding class, and the building top is not regarded as the shielding of the whole building. Although this overlapping type of class labeling may cause overlapping ambiguity problem, on the one hand, instance segmentation supports the case that the same pixel has two class labels, and most methods have related strategies to handle overlapping, on the other hand, since there is a clear subordinate relationship between the building top and the whole, the ambiguous pixels between the two classes are directly assigned to the top class, which does not affect the extraction of building information as a whole.
[0071] Step 122: Construct a perspective instance segmentation network for the main subject. Pre-train the network using labeled oblique remote sensing simulation images and oblique remote sensing images of buildings to obtain a pre-trained model. Then, use oblique remote sensing images of buildings and the pre-trained model for transfer learning to obtain perspective instance segmentation results.
[0072] The main perspective instance segmentation network includes: FCOS object detector layer, convolutional layer, feature pyramid layer, feature decoder layer, mask prediction layer, occlusion classification layer, and output layer;
[0073] The perspective instance segmentation results are obtained using a subject perspective instance segmentation network, specifically as follows:
[0074] ① First, the entire building and its facade are processed separately using an FCOS-based object detector module to obtain the detection bounding boxes for the entire building, its top surface, and its facade. Then, an attention map r for each instance is obtained through a convolutional layer. k ;
[0075] ② Simultaneously, low-level features of the building tilt remote sensing image are obtained through the feature pyramid network at the front end, and input into the feature decoder to add the bases of the overall image to obtain the base of each instance.
[0076] ③ Based on the BlendMask blending mask strategy, for each instance base s k With attention map r k The merging process achieves the final mask prediction m. d As shown in equation (2), where K is the set number of bases. This indicates element-wise multiplication, where D is the total number of predicted detection boxes.
[0077]
[0078] ④ Using the image within the detection box of each building as a whole, determine whether there are occlusion areas. For buildings with occlusion, predict the mask m for their invisible areas. occ ;
[0079] ⑤ For each building, combine the mask md and occlusion mask m of the overall building facade predicted in ③ and ④. occ A comprehensive mask is applied to generate the final perspective instance segmentation result.
[0080] Step Two, as follows Figure 2 As shown, information restoration is performed on the perspective instance segmentation results obtained in step one to obtain an image of a single building:
[0081] First, based on the perspective instance segmentation results, obtain the mask area S of the invisible region for each building. occand the building overall area mask area S bui , judging the proportion Z of the occlusion area relative to the building overall s , setting a proportion threshold T, if the occlusion proportion Z s is less than or equal to the threshold T, it indicates that the building overall has more visible information and can provide sufficient features, so an Image Inpainting strategy based on edge deduction can be used to repair the local missing; if the occlusion proportion Z s is greater than the threshold T, it indicates that the building overall has a large occluded area and it is difficult to extract effective texture edges. As much as possible, select the part with more information in the unoccluded area as the basis, and use an Image Outpainting strategy based on contour constraint for expansion prediction;
[0082]
[0083] Combined with the original image and the various mask results of the perspective instance segmentation, the following data of each independent building can be obtained through simple processing, which can be used as the input conditions required by different strategies respectively: the original image area I RGB , the comprehensive mask graph M of the building top, facade, occlusion information and background area A , the original image I with background area and occlusion area mask Z_RGB , its gray image I Z_gray , the occlusion area mask graph M Z , the mask graph Mb of the building overall and the background based on the predicted outer contour Co.
[0084] Step three, as shown in Figure 3 , matching the same single building from different building inclined remote sensing images, including the following steps:
[0085] Step three one, respectively obtaining the overall, top and facade area pixels of each single building;
[0086] Step three two, establishing an observation angle discrimination network, inputting the overall, top and facade area pixels of each building obtained in step three one into the observation angle discrimination network to obtain the observation angle of each building;
[0087] The observation angle discrimination network comprises a feature extraction layer and an observation angle discrimination layer.
[0088] Among them, the overall, top and facade area pixels of each building in the image enter different channels of the observation angle discrimination network respectively, and the three channels share parameters in the network.
[0089] The feature extraction layer is ResNet-50, and for the building in the image whose facade is not visible, the facade features are directly set to zero.
[0090] The observation angles include: quasi-forward, quasi-lateral, oblique;
[0091] The observation angle discriminant network loss function is as follows:
[0092] L total = λ w L w + λ t L t + λ s L s (4)
[0093] Wherein, L w , L t and L s are the loss functions of the whole building, the top surface and the facade, and λw, λt and λs are regularization coefficients.
[0094] Step three, based on the observation angle of each building, the same single building image from different building oblique remote sensing images is matched:
[0095] Step three, based on the observation angle obtained in step three, the same building in different building oblique remote sensing images is divided into same observation angle (S-view) and different observation angle (D-view);
[0096] Step three, the target level matching network is adopted to match the same building in the same observation angle and different observation angle data sets respectively; for the single building image without matching object, the shooting time is obtained, which is earlier than the database record, and is regarded as a demolished building, and is later than the database record, and is regarded as a new building, and is added into the database, and subsequent three-dimensional reconstruction based on single image is directly carried out;
[0097] The target level matching network of the building is established, the buildings are divided into same observation angle (S-view) and different observation angle (D-view) according to the observation angle, two independent depth measurements of the samples are learned from S-view and D-view respectively, the spatial constraint and cross-space constraint are strengthened, and thus the first stage matching of the buildings with observation angle difference is better completed. For the buildings without matching object, the shooting time is processed in sequence. Earlier than the database record, it is regarded as a demolished building, and later than the database record, it is regarded as a new building, and is added into the database, and subsequent three-dimensional reconstruction based on single image is directly carried out.
[0098] Step 333: Based on the matching results obtained in Step 332, the original tilted remote sensing images of each building are restored. Nm neighboring buildings are selected with each building as the center. The centroid coordinates of the top surface of each building in the image are extracted as the building's center position. The distance and relative azimuth between the two centroids are calculated according to the ground resolution and observation tilt angle. Although the estimated distance between buildings is not precise, considering the high image resolution, these errors are relatively small compared to the actual distances between different buildings. The relative azimuth and distance features are used to perform a second screening of the first-stage quasi-matching results. Mismatches caused by high image feature similarity and one-to-many cases are excluded. Similarly, mismatched buildings are added to the database, and subsequent 3D reconstruction is performed directly based on a single image.
[0099] Step 4: Based on the matching results of Step 3, use the images of each individual building to achieve 3D reconstruction of ultra-generalized stereo pairs.
[0100] "Ultra-generalized stereo image pair"—an arbitrary-width oblique remote sensing image covering a certain area (oblique remote sensing images refer to visible light remote sensing images acquired at an observation angle deviating from the vertical direction by more than 3°), such as... Figure 4 As shown. Specifically, the composition of ultra-generalized stereo image pairs is free; they can be single images, multiple images, or new images can be added at any time based on available image resources. They can be acquired from different payload platforms at different times, and multiple images have non-fixed differences in observation angles and ground resolution. The intersection angle, base-to-height ratio, and overlap between pairs of images may not necessarily meet the constraints of standard stereo image pairs.
[0101] Super-generalized stereo image pairs are simple to construct, abundant in resources, and can contain a large amount of building facade information. With the rapid development of the aerospace industry, my country has achieved remarkable capabilities in acquiring high-resolution visible light remote sensing images. If 3D reconstruction of buildings meeting certain accuracy requirements can be achieved through super-generalized stereo image pairs, the application potential of 3D building reconstruction technology based on visible light remote sensing images can be significantly improved, possessing significant theoretical and practical value. However, given the characteristics of super-generalized stereo image pairs, existing 3D reconstruction technology frameworks struggle to handle their wide range of input conditions. In most cases, dense stereo matching of every building in the scene is impossible, let alone obtaining the 3D coordinates of a large number of points. Furthermore, super-generalized stereo image pairs consist of oblique remote sensing images, and the processing and analysis of oblique remote sensing images face a variety of scientific problems arising from the diversity of observation angles, ground resolution, building shapes, and occlusion conditions. This invention solves this problem.
Claims
1. A method for information extraction and 3D reconstruction based on hyperstereo pairs, characterized in that The method specifically comprises the following steps: Step one, obtaining building oblique remote sensing images, and performing perspective instance segmentation on the building oblique remote sensing images to obtain perspective instance segmentation results; Step two, performing information repair on the perspective instance segmentation results obtained in step one to obtain single building images; Step three, matching the same single building images from different building oblique remote sensing images, comprising the following steps: Step three one, obtaining the overall, top surface and facade area pixels of each single building respectively; Step three two, establishing an observation angle discrimination network, inputting the overall, top surface and facade area pixels of each building obtained in step three one into the observation angle discrimination network to obtain the observation angle of each building; The observation angle discrimination network adopts a ResNet-50 network for feature extraction; The observation angle of each building comprises a quasi-forward direction, a quasi-lateral direction and an oblique direction; Step three three, based on the observation angle of each building, matching the same single building images from different building oblique remote sensing images to obtain matching results, comprising the following steps: Step three three one, based on the observation angle obtained in step three two, dividing the same building in different building oblique remote sensing images into same-class observation angle data sets and different-class observation angle data sets; Step three three two, using a target-level matching network to classify and match in the same-class observation angle data sets and the different-class observation angle data sets to obtain matching results, i.e. all images of the same building, and then performing step three three three; comparing the shooting time of the single building image that does not exist a matching object with the recorded time in the database to filter and obtain new single building images, and saving the new single building images to the database to directly perform step four; Step three three three, based on the matching results obtained in step three three two, restoring the original building oblique remote sensing images where each building is located, then selecting Nm buildings in the neighborhood centering on each building to obtain the top surface centroid coordinates of each building as the center position of the building, obtaining the distance and relative orientation between two centroids according to the ground resolution and the observation depression angle, i.e. the distance and relative orientation between buildings, and excluding the false matching in step three three two according to the distance and relative orientation between buildings, adding the unmatched single building images to the database to perform step four; The false matching is the false matching and one-to-many situation caused by high image feature similarity; Step four, based on the matching results of step three, realizing hyper-generalized stereo image pair three-dimensional reconstruction by using each single building image; The hyper-generalized stereo image pair is any oblique remote sensing image covering a certain area.
2. The hyperstereo pair based information extraction and 3D reconstruction method according to claim 1, characterized in that: In step one, the building oblique remote sensing images are obtained, and perspective instance segmentation is performed on the building oblique remote sensing images to obtain perspective instance segmentation results, comprising the following steps: Step one one, simulating the building oblique remote sensing images based on virtual knowledge of the building oblique remote sensing images to obtain oblique remote sensing simulation images; Step one two, building a main perspective instance segmentation network, and obtaining perspective instance segmentation results by using the oblique remote sensing simulation images and the main perspective instance segmentation network, specifically as follows: Step one two one, label the inclined remote sensing simulation image; The label includes: building whole instance, building top surface instance, occlusion instance, background; Step one two two, build a main perspective instance segmentation network, use the labeled inclined remote sensing simulation image and building inclined remote sensing image to pretrain the main perspective instance segmentation network to obtain a pre-trained model; Step one two three, use the building inclined remote sensing image and the pre-trained model for transfer learning to obtain the perspective instance segmentation result.
3. The method of claim 2, wherein: The virtual knowledge of the building inclined remote sensing image is obtained by the following formula: (1) wherein, is a building tilt remote sensing simulation image, is a corresponding ground truth, Terrain, Object, Texture are the terrain, object, texture of the virtual scene, P is an imaging projection model, and are the results of mapping the same simulation three-dimensional scene with simulation texture and label texture respectively, and are simulation texture and label texture respectively.
4. The method of claim 3, wherein: The main perspective instance segmentation network includes: FCOS-based target detector layer, convolution layer, feature pyramid layer, feature decoder layer, mask prediction layer, occlusion classification layer and output layer; The FCOS-based target detector layer is used to obtain the detection box region of the building whole, top surface and facade in the building inclined remote sensing image; The convolutional layer is used to obtain the attention map r of each instance according to the detection frame region of the building as a whole, the top surface and the facade k ; The attention map of each instance includes: the attention map of the building whole instance, the attention map of the building top surface instance and the attention map of the occlusion instance; The feature pyramid is used to obtain the features of the building inclined remote sensing image; The feature decoder layer is configured to obtain the base of each instance according to the features of the building tilt remote sensing image ; The mask prediction layer is configured to generate a base mask for each instance and attention map merge, obtain a predicted mask based on a BlendMask blending mask strategy ; The occlusion classification layer is used to determine whether there is an occlusion region in the image within the detection frame region of each building as a whole, and for the building with occlusion, the mask m of the invisible region is obtained occ ; The output layer is used to output m d and m occ The integrated mask is obtained by merging, and the perspective instance segmentation result is output.
5. The method of claim 4, wherein: The base of each instance is described and attention map r k Merging, the predicted mask m is obtained based on the BlendMask blending mask strategy d As follows: (2) where k is the base number of each instance, K is the number of bases set, represents the element product, d is the detection frame number, and D is the total number of detection frames.
6. The method of claim 5, wherein: The perspective instance segmentation result obtained in step one is repaired in step two to obtain a single building image, specifically: Step two, based on the perspective instance segmentation result obtained in step one, obtaining the invisible mask area S of each building in the building tilt remote sensing image occ and the building overall area mask area S bui , obtaining the ratio Z s of S occ and S bui ; Step two, set a ratio threshold T, compare T with Z s If Z s is less than or equal to threshold T, adopt Image Inpainting strategy to repair local missing to obtain single building image; if Z s is greater than threshold T, adopt Image Outpainting strategy to repair local missing to obtain single building image.
7. The hyperstereo pair based information extraction and 3D reconstruction method according to claim 6, characterized in that: The observation angle judgment network loss function is as follows: (4) where, , and are the loss functions for the building as a whole, the roof and the facade, respectively, , and are regularization coefficients.
8. The method of claim 7, wherein: In step three two, the shooting time of the single building image without matching object is compared with the record time in the database to obtain the newly added building single image, specifically: The shooting time of the single building image without matching object is compared with the record time in the database, if the shooting time of the single building image with matching object is earlier than the record time in the database, it is a demolished building; if the shooting time of the single building image with matching object is later than the record time in the database, it is a newly added building.
Citation Information
Patent Citations
Three-dimensional architecture rapid modelling approach based on stereopair
CN101286241A
Restoration method for building shielding information in inclined remote sensing image
CN113298808A