Method, System, Electronic Device, and Storage Medium for Obtaining Feature Information

By integrating the representation information of satellite remote sensing images and nDSM information step by step, the problem of low acquisition efficiency of high-precision real-life three-dimensional model in the prior art is solved, and automated and precise landform profile extraction and segmentation are realized.

CN116524372BActive Publication Date: 2025-07-25BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310506294.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-06
Publication Date
2025-07-25
Estimated Expiration
2043-05-06

AI Technical Summary

Technical Problem

The prior art has low efficiency, long cycles and high labor costs when acquiring high-precision real-life three-dimensional models. Traditional surveying and mapping solutions rely on a large amount of labor, and the scanning efficiency of airborne lidar is also limited.

Method used

Using step-by-step fusion processing, the cross-modal representation information of satellite remote sensing images, including image information and normalized digital surface model nDSM, is used to fuse the two representation information through weight parameters, and generate the geographic outline information step by step to realize automatic geographic information extraction.

Benefits of technology

It improves the accuracy and efficiency of land object profile segmentation, realizes automatic land object information acquisition without manual intervention, and improves the accuracy and processing efficiency of land object profiles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116524372B_ABST
    Figure CN116524372B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, a system, an electronic device, and a storage medium for obtaining ground object information, which relates to the field of artificial intelligence, and particularly relates to fields such as cloud computing and big data. The specific implementation solution is as follows: The method includes: obtaining two types of cross-modal representation information corresponding to a satellite remote sensing image; based on the two types of representation information, performing fusion processing in a step-by-step fusion processing manner to obtain a target fusion result, and according to the target fusion result, obtaining the contour information of each ground object contour; The step of performing fusion processing in a step-by-step fusion processing manner to obtain a target fusion result includes: obtaining the weight parameters corresponding to each type of representation information at the current level; based on the weight parameters, fusing the two types of representation information to obtain an intermediate fusion result; generating new representation information at the next level based on the intermediate fusion result, and re-executing the step of obtaining the weight parameters corresponding to each type of representation information at the current level until the result of the last level is obtained and used as the target fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and particularly to fields such as cloud computing and big data. Background Art

[0002] As a key information that truly and stereoscopically reflects the human production, living and ecological space, a high-precision real-scene three-dimensional model can realize the real-time correlation and interconnection between the digital space and the real space, and has very wide applications, including urban digital government, urban planning and design, electronic map navigation, etc.

[0003] Currently, the following two main schemes are mainly adopted to obtain the real-scene three-dimensional model: (1) The traditional surveying and mapping scheme relies on a large amount of manual work. Surveying and mapping personnel use professional surveying and mapping instruments to design the acquisition route and complete the elevation measurement of each building one by one to obtain the real-scene three-dimensional model; (2) Use airborne lidar scanning to obtain the point cloud information of the scanning area, so as to obtain a high-precision real-scene 3D (three-dimensional) model. However, the above schemes generally have problems such as low efficiency, long cycle, and high labor costs. Summary of the Invention

[0004] The present disclosure provides a method, a system, an electronic device and a storage medium for obtaining ground object information.

[0005] According to one aspect of the present disclosure, a method for obtaining ground object information is provided. The method includes:

[0006] Obtain two cross-modal representation information corresponding to the satellite remote sensing image;

[0007] Based on the two representation information, perform fusion processing in a hierarchical fusion processing manner to obtain a target fusion result, and according to the target fusion result, obtain the contour information of the contour of each ground object in the satellite remote sensing image;

[0008] Wherein, the step of performing fusion processing in a hierarchical fusion processing manner to obtain a target fusion result includes:

[0009] Obtain the weight parameter corresponding to each representation information of the current level;

[0010] Based on their respective weight parameters, fuse the two representation information to obtain an intermediate fusion result;

[0011] Generate new representation information corresponding to the next level based on the intermediate fusion result, and re-execute the step of obtaining the weight parameter corresponding to each representation information of the current level until the intermediate fusion result corresponding to the last level is obtained and used as the target fusion result.

[0012] According to another aspect of the present disclosure, a device for obtaining ground object information is provided. The device includes:

[0013] A characterization information acquisition module for acquiring two types of cross-modal characterization information corresponding to a satellite remote sensing image;

[0014] A fusion processing module for performing fusion processing in a hierarchical fusion processing manner based on the two types of characterization information to obtain a target fusion result;

[0015] A contour information acquisition module for acquiring contour information of the contour of each ground object in the satellite remote sensing image according to the target fusion result;

[0016] Wherein, the fusion processing module includes:

[0017] A weight parameter acquisition unit for acquiring weight parameters corresponding to each type of the characterization information at the current level;

[0018] A fusion processing unit for fusing the two types of characterization information based on their respective weight parameters to obtain an intermediate fusion result;

[0019] A characterization information generation unit for generating new characterization information corresponding to the next level based on the intermediate fusion result, and calling the weight parameter acquisition unit until the fusion processing unit obtains the intermediate fusion result corresponding to the last level and uses it as the target fusion result.

[0020] According to another aspect of the present disclosure, there is provided an electronic device, including:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.

[0024] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above method.

[0025] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the above method when executed by a processor.

[0026] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0027] The accompanying drawings are used to better understand the present solution and do not constitute a limitation to the present disclosure. Among them:

[0028] Figure 1 is the first flowchart of the method for obtaining ground object information according to the first embodiment of the present disclosure;

[0029] Figure 2 is the second flowchart of the method for obtaining ground object information according to the first embodiment of the present disclosure;

[0030] Figure 3 is the first schematic diagram of the fusion processing process according to the first embodiment of the present disclosure;

[0031] Figure 4 is the second schematic diagram of the fusion processing process according to the first embodiment of the present disclosure;

[0032] Figure 5 is the third flowchart of the method for obtaining ground object information according to the first embodiment of the present disclosure;

[0033] Figure 6 is the first schematic diagram of the boundary where the contour is located according to the first embodiment of the present disclosure;

[0034] Figure 7 is the second schematic diagram of the boundary where the contour is located according to the first embodiment of the present disclosure;

[0035] Figure 8 is the third schematic diagram of the boundary where the contour is located according to the first embodiment of the present disclosure;

[0036] Figure 9 is the fourth schematic diagram of the boundary where the contour is located according to the first embodiment of the present disclosure;

[0037] Figure 10 is the fourth flowchart of the method for obtaining ground object information according to the first embodiment of the present disclosure;

[0038] Figure 11 is the schematic diagram of the roof building slope of the building according to the first embodiment of the present disclosure;

[0039] Figure 12 is the fifth flowchart of the method for obtaining ground object information according to the first embodiment of the present disclosure;

[0040] Figure 13 is the schematic diagram of a single original satellite image according to the first embodiment of the present disclosure;

[0041] Figure 14 is the schematic diagram of the cut map sheet block according to the first embodiment of the present disclosure;

[0042] Figure 15It is a schematic diagram of the matching of characteristic pixel points according to the first embodiment of the present disclosure;

[0043] Figure 16 It is a schematic diagram of the three-dimensional point cloud file according to the first embodiment of the present disclosure;

[0044] Figure 17 It is a schematic diagram of the digital surface model DSM according to the first embodiment of the present disclosure;

[0045] Figure 18 It is a schematic diagram of feature extraction according to the first embodiment of the present disclosure;

[0046] Figure 19 It is a schematic diagram of the module of the ground object information acquisition device according to the second embodiment of the present disclosure;

[0047] Figure 20 It is a block diagram of an electronic device for implementing the method for acquiring ground object information according to the first embodiment of the present disclosure. Detailed implementation manners

[0048] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0049] Embodiment 1

[0050] As Figure 1 shown, the method for acquiring ground object information in this embodiment includes:

[0051] S101. Obtain two cross-modal representation information corresponding to the satellite remote sensing image;

[0052] Among them, the two cross-modal representation information belongs to two completely different dimensional information, and both can illustrate the ground object situation in the satellite remote sensing image from a certain angle; the two cross-modal representation information is mutually registered, that is, the same pixel points correspond one by one.

[0053] For example, the two cross-modal representation information includes the first image information of the satellite remote sensing image, and the first nDSM information in the nDSM (normalized digital surface model) that matches the satellite remote sensing image, etc.

[0054] S102. Based on the two representation information, perform fusion processing in a hierarchical fusion processing manner to obtain a target fusion result;

[0055] Among them, based on the advantages of different representation information, two types of cross-modal representation information are fused, enabling the network to fully extract the representation information of the two modalities to ensure the accuracy of ground object contour segmentation.

[0056] S103. Obtain the contour information of the contour of each ground object in the satellite remote sensing image according to the target fusion result; where the ground objects include but are not limited to buildings.

[0057] Specifically, as Figure 2 shown, the steps of performing fusion processing in a hierarchical fusion processing manner to obtain the target fusion result include:

[0058] S201. Obtain the weight parameter corresponding to each type of representation information at the current level;

[0059] This weight parameter is used to represent the influence degree of each type of representation information on the position where the contour of the ground object is located.

[0060] S202. Based on their respective weight parameters, fuse the two types of representation information to obtain an intermediate fusion result;

[0061] S203. Generate new representation information corresponding to the next level based on the intermediate fusion result, and re-execute step S201 until the intermediate fusion result corresponding to the last level is obtained and used as the target fusion result.

[0062] In this solution, the influence degree of each type of representation information on the position where the contour of the ground object is located is obtained, and it is combined with the two types of representation information for hierarchical fusion calculation to obtain the final target fusion result, making full use of the advantages of different cross-modal representation information, enabling the network to fully extract the representation information of the two modalities, effectively improving the determination effect of the contour of the ground object in the satellite remote sensing image, and realizing automated extraction of ground object information without manual intervention, improving the accuracy and efficiency of ground object contour segmentation.

[0063] In an implementable solution, the two types of cross-modal representation information include the first image information of the satellite remote sensing image and the first nDSM information in the normalized digital surface model nDSM that matches the satellite remote sensing image.

[0064] In this solution, selecting the first image information and the first nDSM information as the two types of cross-modal representation information can effectively ensure the accuracy and efficiency of ground object contour segmentation in the satellite remote sensing image compared with other representation information; a multi-modal fusion-based ground object semantic segmentation solution is proposed, taking the satellite remote sensing image and its corresponding nDSM information as input data at the same time, and using the respective advantages of these two cross-modal information to ensure the reliability of building semantic segmentation.

[0065] In an implementable solution, step S201 includes:

[0066] S2011. Obtain the first global information of the first image information and the second global information of the first nDSM information respectively by using global average pooling;

[0067] S2012. Obtain the first attention vector corresponding to the first global information and the second attention vector corresponding to the second global information.

[0068] In this solution, first, the global average pooling method is used to obtain the global information corresponding to the two types of information respectively, and then the global information is processed by MLP (Multi-Layer Perceptron) respectively to obtain the attention vectors corresponding to the two types of information. The attention vector of each type of information is the corresponding weight parameter, and it is combined with the two types of representation information for hierarchical fusion calculation, so that the network can fully extract the representation information of the two modalities and ensure the accuracy of subsequent ground object contour acquisition.

[0069] Among them, in order to improve the processing efficiency, the first image information and the first nDSM information can also be spliced, and then global average pooling is used for synchronous processing to obtain the corresponding first global information and second global information respectively.

[0070] In an implementable solution, step S202 includes:

[0071] S2021. Multiply the first attention vector by the first image information to obtain the second image information;

[0072] S2022. Fuse the second image information with the first nDSM information to obtain the first fusion feature;

[0073] Among them, the second image information and the first nDSM information at the same pixel position are added, and all pixel points are traversed to complete the fusion between the second image information and the first nDSM information to obtain the first fusion feature;

[0074] S2023. Multiply the second attention vector by the first nDSM information to obtain the second nDSM information;

[0075] S2024. Fuse the second nDSM information with the first image information to obtain the second fusion feature;

[0076] Among them, the second nDSM information and the first image information at the same pixel position are added, and all pixel points are traversed to complete the fusion between the second nDSM information and the first image information to obtain the second fusion feature;

[0077] S2025. Splice the first fusion feature and the second fusion feature to obtain the intermediate fusion result.

[0078] Among them, the first fusion feature and the second fusion feature are concatenated to retain all information and avoid the loss of important information, etc., thereby ensuring the accuracy of the final target fusion result, and then ensuring the reliability of the ground object contour acquisition.

[0079] In this solution, for any fusion processing stage, the first attention vector is multiplied by the first image information, and the second attention vector is multiplied by the first nDSM information to respectively obtain new image information and new nDSM information; the new image information is fused with the original first nDSM information, and at the same time the new nDSM information is fused with the original first image information to respectively obtain their respective fusion results, and finally the two fusion results are concatenated to obtain the intermediate fusion result of the current fusion processing stage; that is, the attention vector is used as an offset and combined into the original image information and nDSM information, and then the combined new image information and the original nDSM information, as well as the new nDSM information and the original image information are respectively fused, so that both the original cross-modal information is retained and the information of the attention vector is introduced for fusion coding, realizing the fusion of cross-modal information, improving the determination quality of the position information of the ground object contour in the satellite remote sensing image, and finally ensuring the accuracy of the ground object contour acquisition.

[0080] In an implementable solution, after step S2012, it further includes:

[0081] Perform noise filtering processing on the first attention vector and the second attention vector.

[0082] In this solution, performing noise filtering processing on the first attention vector and the second attention vector can suppress cross-modal noise, ensure the effect of more robust feature expression, and ensure the accuracy of the fusion result of each level of fusion processing.

[0083] In an implementable solution, the steps of performing noise filtering processing on the first attention vector and the second attention vector specifically include:

[0084] Obtain the weight value corresponding to each pixel channel in the satellite remote sensing image;

[0085] Multiply the first value corresponding to each pixel channel in the first attention vector by the corresponding weight value to perform noise filtering processing on the first attention vector;

[0086] Multiply the second value corresponding to each channel in the second attention vector by the corresponding weight value to perform noise filtering processing on the second attention vector.

[0087] In this solution, by performing multiplication processing on each channel of the attention vector, the feature representation of the attention vector after filtering out noise is obtained, ensuring a more robust effect of the feature expression and the accuracy of the fusion result of each level of fusion processing.

[0088] The schematic diagram corresponding to the fusion process of each level is realized by constructing a two-stream information fusion encoder, as Figure 3 shown. Here, A represents the input image information, and B represents the input nDSM information. The input image information and nDSM information are concatenated (Concat), and the concatenated information is processed by global average pooling (Pooling) to obtain the first global information T1 and the second global information T2 corresponding to the two types of information respectively. Then, the first global information T1 and the second global information T2 are processed by MLP respectively to obtain the corresponding first attention vector K1 and second attention vector K2. The first attention vector K1 is multiplied by the image information A to obtain the image information A1, and the second attention vector K2 is multiplied by the nDSM information B to obtain the nDSM information B1. Then, the image information A and the nDSM information B1 are fused to obtain the image information A2, and the nDSM information B1 and the nDSM information B are fused to obtain the nDSM information B2. Finally, the image information A2 and the image information B2 are concatenated to obtain the intermediate fusion result O generated by the current fusion process.

[0089] As Figure 4 shown, P1 is the satellite remote sensing image, P2 is the normalized digital surface model nDSM corresponding to the satellite remote sensing image, and DF-Encoder (two-stream information fusion encoder) corresponds to Figure 9 the intermediate fusion result O. Conv corresponds to performing convolution processing on the intermediate fusion result O to obtain the image information and nDSM information for use in the next-level fusion process. The fusion processing logic of each level is the same. Until the number of fusion levels of the hierarchical fusion processing (for example, 5), the fusion processing is stopped, and the intermediate fusion result O obtained after the 5th fusion processing is used as the target fusion result R1, and the target fusion result R is decoded (Decoder) to obtain the contour segmentation result R of the building.

[0090] In an implementable solution, generating the new representation information corresponding to the next level based on the intermediate fusion result in step S203 includes:

[0091] Performing convolution processing on the intermediate fusion result to obtain two new representation information of cross-modalities;

[0092] That is, the new first image information and the new first nDSM information for the next-level fusion processing are obtained at this time;

[0093] Among them, the characterization information at each level is different and the corresponding information sizes are the same.

[0094] In this solution, the step-by-step fusion process corresponds to N (N is a positive integer) fusion levels. The input information for the next-level fusion process is obtained based on the fusion result of the previous level. Specifically, the intermediate fusion result of the previous level is convolved to obtain new image data and new nDSM information. Moreover, the new image information obtained at each level is consistent with the information size of the original image information, and the new nDSM information is consistent with the information size of the original nDSM information, thereby ensuring the feasibility and processing efficiency of the fusion process.

[0095] In addition, of course, according to actual needs, in addition to convolution processing, other processing methods can also be combined as long as new image data and new nDSM information of two cross-modal types that meet the requirements of the fusion process are obtained.

[0096] In an implementable solution, the method further includes:

[0097] During the step-by-step fusion process, it is judged whether the current intermediate fusion result has reached the first preset processing condition. If not, the total number of fusion levels is increased to the first fusion level;

[0098] If it has reached and has not reached the preset number of fusion levels, the total number of fusion levels is decreased to the second fusion level.

[0099] By presetting the preset number of fusion levels, it is considered that the corresponding number of fusion processes in different actual scenarios may be different. For example, the preset number of fusion levels is 5. When the fusion process reaches the 3rd level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result representation can obtain a ground object contour effect that meets the requirements, the fusion process can be stopped, and at this time, the number of fusion levels is automatically updated to 3; conversely, when the fusion process reaches the 5th level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result still cannot obtain a ground object contour effect that meets the requirements, the fusion process can be increased, and at this time, the number of fusion levels is automatically updated to 7; when the fusion process reaches the 5th level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result representation can obtain a ground object contour effect that meets the requirements, there is no need to adjust the number of fusion levels at this time.

[0100] In this solution, by dynamically adjusting the number of fusion levels according to the actual situation, the flexibility and rationality of the overall processing are effectively improved, which not only ensures the effective and reliable acquisition of ground object information but also avoids unnecessary waste of computing resources.

[0101] In an implementable solution, as Figure 5 shown, after step S103, it further includes:

[0102] S501. Based on the contour information, obtain a number of boundary points corresponding to the contour of the ground object and the position information corresponding to the boundary points;

[0103] Among them, the Alpha-Shape (algorithm for extracting boundary points) method is used to determine the boundary points on the contour of the ground object; of course, other feasible methods can also be used, which will not be elaborated here.

[0104] S502. Update the contour information according to the position information.

[0105] In this solution, based on the contour information obtained from the above process, the boundary points on these contours and their positions are obtained, and then the positions of the contours of each ground object are optimized, improving the accuracy of determining the contour information of the ground object.

[0106] In an implementable solution, step S502 includes:

[0107] S5021. Obtain a number of segments according to the position information of the number of boundary points;

[0108] Among them, according to the position information, the boundary points are classified geometrically. The boundary points generally in one direction correspond to a line segment, that is, the distance difference between two boundary points on the same line segment is less than a preset value.

[0109] S5022. Perform fitting processing on the position information of multiple boundary points on each segment to obtain the corresponding straight line;

[0110] S5023. Based on a number of connected straight lines, obtain the boundary line corresponding to the contour of the ground object, and use the boundary information corresponding to the boundary line as the contour information.

[0111] In this solution, the initial contour information segmented based on two cross-modal representation information is improved. Each straight line intersects to form the boundary line of the ground object contour as the final contour, more finely and accurately segmenting the contour of the ground object to meet higher requirements of usage scenarios.

[0112] In an implementable solution, the method further includes:

[0113] Determine the main direction of the ground object according to the position information of the number of boundary points;

[0114] Among them, before step S5023, it further includes:

[0115] Regularize the straight lines according to the main direction to update and obtain the straight lines that meet the second preset processing condition.

[0116] In this solution, through the main direction of the ground object, the drawn straight lines are regularized. For those with a large deviation from the main direction, adjustment is required to ensure the accuracy of the boundary line corresponding to the contour. For example, the corresponding straight line is adjusted to be parallel or perpendicular to the main direction. Of course, for buildings with special shapes, synchronous adjustment can also be made in combination with their special shapes, which will not be elaborated here.

[0117] Specifically, as Figure 6 shown, to identify several boundary points of the building contour determined based on the representation information of two cross-modalities; as Figure 7 shown, corresponding segments are obtained based on several boundary points; as Figure 8 shown, the straight line corresponding to each segment and the intersection points between different straight line connections are obtained; as Figure 9 shown, the boundary line of the contour formed by the connections of different straight lines.

[0118] In an implementable solution, when the ground object is a building, as Figure 10 shown, the method further includes:

[0119] S401. Obtain the style information of the roof inside the contour of the building;

[0120] S402. Based on the style information, adopt a matching preset height determination strategy to obtain the height information of the building.

[0121] In this solution, for the height information of different buildings, different height determination strategies will be adopted according to the style of the roof, rather than a unified determination scheme, thus ensuring the accuracy of determining the height information of different ground objects.

[0122] In an implementable solution, step S402 includes:

[0123] When the style information represents that the roof is flat, obtain the first mode among the height values corresponding to each pixel point inside the contour, and the frequency of occurrence of the first mode;

[0124] When the frequency is greater than or equal to the first set value, calculate the first mean value of all height values whose difference from the mode is less than the first distance, and take the first mean value as the height information of the building;

[0125] When the frequency is less than the first set value, calculate the second mean value of all height values whose difference from the first mode is less than the second distance, and take the second mean value as the height information of the building;

[0126] Wherein, the second distance is greater than the first distance.

[0127] For example: Based on the schematic diagram of the roof surface of a building, select buildings with flat roofs, and then obtain the mode and frequency of the height (unit: meter). If the frequency ≥ 0.7, the height value of the building is the mean value within the range of the mode ± 1 meter (unit: centimeter); if the frequency < 0.7, the height value of the building is the mean value within the range of the mode ± 3 meters (unit: centimeter).

[0128] In this solution, for flat-topped buildings, based on the mode and frequency of the height values of all pixel points within the contour, determine the corresponding height determination scheme, thereby ensuring the accuracy and reliability of obtaining the height information of flat-topped buildings.

[0129] In an implementable solution, step S402 includes:

[0130] When the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is greater than or equal to the second set value, obtain the third mean value of the maximum height value and the minimum height value, and use the third mean value as the height information of the building.

[0131] When the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is less than the second set value, obtain the first mode of the height values corresponding to each pixel point within the contour, eliminate the height values that differ from the first mode by more than the third distance, obtain the fourth mean value of the maximum height value and the minimum height value among the remaining all height values, and use the fourth mean value as the height information of the building.

[0132] For example: As Figure 11 shown, based on the schematic diagram of the roof surface of a building, select buildings with pitched roofs, domes, etc., and then obtain whether the difference in height values corresponding to the pixel points within the calibration surface coordinate string is within 6m. If so, the height value of the building is the mean value of the maximum height value and the minimum height value;

[0133] If not, obtain the mode based on the height values (unit: meter), eliminate the pixel points outside the range of the mode ± 3 meters, and use the mean value of the maximum value and the minimum value among the remaining all height values as the height value of the building (unit: centimeter).

[0134] In this solution, for buildings with non-flat roofs (irregular raised structures such as pitched roofs and domes), based on the mode and frequency of the height values of all pixel points within the contour, determine the corresponding height determination scheme, thereby ensuring the accuracy and reliability of obtaining the height information of flat-topped buildings.

[0135] In an implementable solution, the steps of obtaining the first nDSM information include:

[0136] Obtain the digital surface model DSM corresponding to the satellite remote sensing image;

[0137] Obtain the corresponding normalized digital surface model nDSM based on the digital surface model DSM;

[0138] Among them, obtaining the normalized digital surface model nDSM from the digital surface model DSM belongs to the mature technology in this field and will not be elaborated here.

[0139] Obtain the first nDSM information according to the normalized digital surface model nDSM.

[0140] In this solution, the normalized digital surface model nDSM is obtained from the digital surface model DSM to obtain the first nDSM information, so as to ensure the timeliness and reliability of obtaining the nDSM information.

[0141] In an implementable solution, as Figure 12 shown, the steps of obtaining the digital surface model DSM corresponding to the satellite remote sensing image specifically include:

[0142] S1201. Obtain the original satellite image stereo pair;

[0143] Among them, the original satellite image stereo pair is a set of satellite remote sensing images taken by a remote sensing satellite from different angles of the same area;

[0144] S1202. Segment the original satellite image stereo pair to obtain a number of sub-satellite image stereo pairs with a preset resolution;

[0145] Among them, since the data file of the original satellite image stereo pair is large, with a resolution of over 100 million pixels; for satellite image data with a resolution of 10000*10000, the coverage range can reach dozens of square kilometers, and it is difficult and inefficient to use it for feature matching. Therefore, it is necessary to divide the satellite image stereo pair into tiles. The tile division with a resolution of 1024*1024 can be adopted to divide the original satellite image stereo pair into dozens or even hundreds of small stereo pairs, improving the efficiency of subsequent feature extraction and feature matching. Specifically, see Figure 13 is a single original satellite image, see Figure 14 is from Figure 1 a tile block cut from the upper left corner, that is, the sub-satellite image stereo pair.

[0146] S1203. Use a preset feature extraction method to process the sub-satellite image stereo pair to obtain the target feature vector;

[0147] S1204. Based on the target feature vector, extract the candidate feature pixel points;

[0148] S1205. According to the similarity between the candidate feature pixel points, obtain all the matching pixel points corresponding to the sub-satellite image stereo pair;

[0149] S1206. Calculate the elevation information corresponding to each matching pixel point;

[0150] S1207. Obtain the digital surface model DSM based on the elevation information.

[0151] Feature extraction is respectively performed on the sub-satellite image stereo pair to obtain candidate feature pixel points and generate feature descriptors. The feature descriptor composed of multi-dimensional vectors contains the scale where the descriptor is located, the position of the descriptor, and the feature information of the descriptor.

[0152] In the feature similarity matching stage, the Euclidean distance of feature vectors is used to calculate the similarity between feature points in two images, and the KD-tree (a tree-shaped data structure for storing instance points in k-dimensional space for rapid retrieval) algorithm is used to improve the detection speed; the sub-satellite image stereo pair includes Image 1 and Image 2. Take a certain feature point A in Image 1, and find two feature points, the point B with the closest feature vector distance and the second-closest point C in Image 2. Use d(A,B) to represent the similarity between A and B. If the ratio between d(A,B) and d(A,C) is less than a certain set threshold, then determine this pair of feature points A and B with the closest feature vector distance as matching points; and so on, to obtain multiple pairs of mutually matching feature points in the sub-satellite image stereo pair.

[0153] The fundamental matrix F and the homography matrix H are calculated using epipolar geometry, and the matching points of all points in the image are calculated, as Figure 15 shown. Specifically, O1 and O2 are the optical centers of the cameras at two positions respectively, P is a three-dimensional point in space, and p1 and p2 are the pixel points corresponding to point P on different imaging planes. The three points P, O1, and O2 are coplanar, which is called the epipolar plane. The intersection lines l1 and l2 of the epipolar plane and the image planes are the epipolar lines of the two image planes; from the epipolar geometry constraint, p2 T Fp1 = 0, where F is the fundamental matrix. The epipolar constraint contains translation and rotation, and gives the spatial position relationship between two matching points.

[0154] In addition, for the homography matrix H describing the mapping relationship between planes, there is At least 4 pairs of mutually matching feature points are used and substituted into the transformation formula to obtain a set of linear equations, and then the fundamental matrix F and the homography matrix H are calculated; for the matrices F and H, the matrix with a small reprojection error is selected as the final estimated matrix; the estimated matrix is used to calculate the pixel points in Image 2 corresponding to all pixel points in Image 1 to determine the one-to-one matching of the corresponding pixel points and obtain all the matching pixel points.

[0155] Among them, the elevation information corresponding to each matching pixel point is obtained by using the RPC model. Of course, other implementable methods can also be used, which will not be elaborated here.

[0156] Linear interpolation encryption is performed to generate a point cloud file. Specifically, after obtaining the elevation information h of each matching pixel point, according to the actual accuracy requirements, linear interpolation encryption is performed on the pixel points in the sub-satellite image stereo pair, converting the original planar pixel points into three-dimensional point pixel points in space, thereby generating a three-dimensional point cloud file, as Figure 16 shown; sampling the three-dimensional point cloud file at a certain interval and using its height value as the gray value of the pixel point to obtain the final tif (Tagged Image File Format) file, as Figure 17 shown, which is the required Digital Surface Model DSM.

[0157] In this solution, a deep learning model is used to extract feature points and perform image feature matching. Without ground control points, three-dimensional point clouds are automatically reconstructed from the sub-satellite image stereo pair. Specifically, a deep learning model based on convolution extracts the features of the satellite image, and then an image matching algorithm is used to find the mutually matching feature points in the sub-satellite image stereo pair; then the epipolar geometry is applied to obtain the mapping relationship of the sub-satellite image stereo pair, and all the image points in the satellite image are mapped, and then used to calculate the elevation information of all points in the image, generating a dense three-dimensional point cloud, thus ensuring the accuracy and reliability of obtaining the Digital Surface Model DSM.

[0158] In an implementable solution, step S1207 includes:

[0159] Perform linear interpolation encryption processing on each matching pixel point to convert each two-dimensional matching pixel point into a three-dimensional matching pixel point;

[0160] Generate a three-dimensional point cloud file based on the three-dimensional matching pixel points;

[0161] Use the elevation information corresponding to each three-dimensional matching pixel point as the gray value, and perform sampling processing on the three-dimensional point cloud file to obtain the Digital Surface Model DSM.

[0162] In this solution, based on the high-resolution satellite image stereo pair, the satellite image stereo pair is divided into tiles; then through key steps such as feature matching, stereo rectification, and triangulation on the divided satellite image stereo pair, the method for photogrammetric three-dimensional reconstruction is applied to large-scale satellite images, and the elevation information of each pair of matching points is obtained from the RPC model, automatically generating a real-scene three-dimensional point cloud model, and thus calculating the Digital Surface Model DSM, thereby ensuring the accuracy and reliability of obtaining the Digital Surface Model DSM.

[0163] In an implementable solution, the preset feature extraction method includes downsampling processing, continuous four-convolution pooling processing, and feature compression processing performed in sequence.

[0164] Specifically, as Figure 18 shown, an image feature extraction model is constructed. The input is the sub-satellite image stereo pair of the tiled satellite image with a resolution of 256*256 after downsampling. After passing through 4 convolutional pooling modules, features of 16*16*512 are output. Then, the PCA (Principal Component Analysis Technique) method is used to compress the features output by the model to retain the preset key information, and finally a 128-dimensional feature vector is obtained.

[0165] In this solution, the above-mentioned feature extraction is adopted to ensure the accuracy of feature extraction, making the Digital Surface Model (DSM) more accurate and reliable, and then ensuring the accuracy of obtaining ground object information.

[0166] Embodiment 2

[0167] As Figure 19 shown, the device for obtaining ground object information in this embodiment includes:

[0168] A characterization information acquisition module 191, configured to acquire two cross-modal characterization information corresponding to the satellite remote sensing image;

[0169] Among them, the two cross-modal characterization information belongs to two completely different-dimensional information, and both can illustrate the ground object situation in the satellite remote sensing image from a certain angle; the two cross-modal characterization information is registered with each other, that is, the same pixel points correspond one by one.

[0170] For example, the two cross-modal characterization information includes the first image information of the satellite remote sensing image, and the first nDSM information in the normalized Digital Surface Model (nDSM) that matches the satellite remote sensing image, etc.

[0171] A fusion processing module 192, configured to perform fusion processing based on the two characterization information by using a hierarchical fusion processing method to obtain a target fusion result;

[0172] Among them, based on the advantages of different characterization information, the two cross-modal characterization information is fused, so that the network fully extracts the characterization information of the two modalities to ensure the accuracy of ground object contour segmentation.

[0173] A contour information acquisition module 193, configured to acquire the contour information of the contour of each ground object in the satellite remote sensing image according to the target fusion result; among them, the ground objects include but are not limited to buildings.

[0174] Among them, the fusion processing module 192 includes:

[0175] A weight parameter acquisition unit 194, configured to acquire the weight parameter corresponding to each characterization information at the current level;

[0176] The weight parameter is used to characterize the influence degree of each characterization information on the position where the contour of the ground object is located.

[0177] The fusion processing unit 195 is configured to fuse the two types of characterization information based on their respective weight parameters to obtain an intermediate fusion result;

[0178] The characterization information generation unit 196 is configured to generate new corresponding characterization information at the next level based on the intermediate fusion result, and call the characterization information generation unit 196 until the fusion processing unit 195 obtains the intermediate fusion result corresponding to the last level and uses it as the target fusion result.

[0179] In this solution, the influence degree of each type of characterization information on the position of the contour of the ground object is obtained, and it is combined with the two types of characterization information for hierarchical fusion calculation to obtain the final target fusion result. The advantages of different cross-modal characterization information are fully utilized, enabling the network to fully extract the characterization information of the two modalities, effectively improving the effect of determining the contour of the ground object in the satellite remote sensing image, and realizing automatic extraction of ground object information without manual intervention, thus improving the accuracy and efficiency of the contour segmentation of the ground object.

[0180] In an implementable solution, the two cross-modal characterization information includes the first image information of the satellite remote sensing image and the first nDSM information in the normalized digital surface model nDSM that matches the satellite remote sensing image.

[0181] In this solution, the first image information and the first nDSM information are selected as the two cross-modal characterization information. Compared with other characterization information, it can effectively ensure the accuracy and efficiency of the contour segmentation of the ground object in the satellite remote sensing image; a multi-modal fusion-based ground object semantic segmentation solution is proposed, taking the satellite remote sensing image and its corresponding nDSM information as input data at the same time, and using the respective advantages of these two cross-modal information to ensure the reliability of the semantic segmentation of the building.

[0182] In an implementable solution, the characterization information generation unit 196 includes:

[0183] The first information acquisition subunit is configured to acquire the first global information of the first image information by using global average pooling;

[0184] The second information acquisition subunit is configured to acquire the second global information of the first nDSM information by using global average pooling;

[0185] The first vector acquisition subunit is configured to acquire the first attention vector corresponding to the first global information;

[0186] The second vector acquisition subunit is configured to acquire the second attention vector corresponding to the second global information.

[0187] In this scheme, the global average pooling method is first used to obtain the global information corresponding to the two types of information, and then the global information is processed by MLP (multi-layer perceptron) to obtain the attention vectors corresponding to the two types of information respectively. The attention vector of each type of information is the corresponding weight parameter, and it is combined with the two representation information for step-by-step fusion calculation, so that the network can fully extract the representation information of the two modalities, ensuring the accuracy of the subsequent acquisition of the contour of the object.

[0188] In order to improve processing efficiency, the first image information and the first nDSM information may be concatenated, and then synchronously processed using global average pooling to obtain the corresponding first global information and second global information, respectively.

[0189] In one feasible solution, the fusion processing unit 195 includes:

[0190] A first multiplication processing subunit, used for multiplying the first attention vector and the first image information to obtain second image information;

[0191] A first feature acquisition subunit is used to fuse the second image information with the first nDSM information to obtain a first fusion feature;

[0192] The second image information at the same pixel position is added to the first nDSM information, and all pixels are traversed to complete the fusion between the second image information and the first nDSM information to obtain the first fusion feature;

[0193] A second multiplication processing subunit is used to multiply the second attention vector by the first nDSM information to obtain second nDSM information;

[0194] A second feature acquisition subunit is used to fuse the second nDSM information with the first image information to obtain a second fusion feature;

[0195] The second nDSM information at the same pixel point position is added to the first image information, and all pixels are traversed to complete the fusion between the second nDSM information and the first image information to obtain a second fusion feature;

[0196] The fusion processing subunit is used to splice the first fusion feature and the second fusion feature to obtain an intermediate fusion result.

[0197] Among them, the first fusion feature and the second fusion feature are spliced to retain all information and avoid the loss of important information, thereby ensuring the accuracy of the final target fusion result and then ensuring the reliability of obtaining the contour of the object.

[0198] In this solution, for any fusion processing stage, the first attention vector is multiplied by the first image information, and the second attention vector is multiplied by the first nDSM information to respectively obtain new image information and new nDSM information; the new image information is fused with the original first nDSM information, and at the same time the new nDSM information is fused with the original first image information to respectively obtain their respective fusion results, and finally the two fusion results are spliced to obtain the intermediate fusion result of the current fusion processing stage; that is, the attention vector is used as an offset and combined into the original image information and nDSM information, and then the combined new image information and the original nDSM information, as well as the new nDSM information and the original image information are respectively fused, so that both the original cross-modal information is retained and the information of the attention vector is introduced for fusion coding, realizing the fusion of cross-modal information, improving the determination quality of the position information of the object contour in the satellite remote sensing image, and finally ensuring the accuracy of the object contour acquisition.

[0199] In an implementable solution, the fusion processing module 192 further includes: a noise processing unit for filtering noise from the first attention vector and the second attention vector.

[0200] In this solution, filtering noise from the first attention vector and the second attention vector can suppress cross-modal noise, ensure the effect of more robust feature expression, and ensure the accuracy of the fusion result of each level of fusion processing.

[0201] In an implementable solution, the noise processing unit includes:

[0202] a weight value acquisition sub-unit for acquiring the weight value corresponding to each pixel channel in the satellite remote sensing image;

[0203] a first noise processing sub-unit for multiplying the first value corresponding to each pixel channel in the first attention vector by the corresponding weight value to filter noise from the first attention vector;

[0204] a second noise processing sub-unit for multiplying the second value corresponding to each channel in the second attention vector by the corresponding weight value to filter noise from the second attention vector.

[0205] In this solution, through the multiplication processing of each channel in the attention vector, the feature representation after filtering noise from the attention vector is obtained, ensuring the effect of more robust feature expression and the accuracy of the fusion result of each level of fusion processing.

[0206] The schematic diagram corresponding to the fusion process of each level above is realized by constructing a two-stream information fusion encoder, as Figure 3As shown in the figure, where A represents the input image information and B represents the input nDSM information. The input image information and nDSM information are concatenated (Concat), and the concatenated information is processed by global average pooling (Pooling) to obtain the first global information T1 and the second global information T2 corresponding to the two types of information respectively. Then, the first global information T1 and the second global information T2 are processed by MLP respectively to obtain the corresponding first attention vector K1 and second attention vector K2. The first attention vector K1 is multiplied by the image information A to obtain the image information A1, and the second attention vector K2 is multiplied by the nDSM information B to obtain the nDSM information B1. Then, the image information A and the nDSM information B1 are fused to obtain the image information A2, and the nDSM information B1 and the nDSM information B are fused to obtain the nDSM information B2. Finally, the image information A2 and the image information B2 are concatenated to obtain the intermediate fusion result O generated by the current fusion process.

[0207] As Figure 4 shown, P1 is a satellite remote sensing image, P2 is the normalized digital surface model nDSM corresponding to the satellite remote sensing image, and DF-Encoder (dual-stream information fusion encoder) corresponds to Figure 9 the intermediate fusion result O. Conv corresponds to performing convolution processing on the intermediate fusion result O to obtain the image information and nDSM information for use in the next-level fusion process. The fusion processing logic for each level is the same. When the number of fusion levels for hierarchical fusion processing (e.g., 5) is reached, the fusion processing is stopped, and the intermediate fusion result O obtained after the 5th fusion processing is used as the target fusion result R1, and the target fusion result R is decoded (Decoder) to obtain the contour segmentation result R of the building.

[0208] In an implementable solution, the feature information generation unit 196 is used to perform convolution processing on the intermediate fusion result to obtain two new cross-modal feature information.

[0209] That is, at this time, new first image information and new first nDSM information for the next-level fusion process are obtained.

[0210] Among them, the feature information for each level is different and the corresponding information sizes are the same.

[0211] In this solution, the hierarchical fusion process corresponds to N (where N is a positive integer) fusion levels. The input information for the next-level fusion process is obtained based on the fusion result of the previous level. Specifically, the intermediate fusion result of the previous level is convolved to obtain new image data and new nDSM information. Moreover, the new image information obtained at each level has the same information size as the original image information, and the new nDSM information has the same information size as the original nDSM information, thereby ensuring the feasibility and processing efficiency of the fusion process.

[0212] In addition, of course, according to actual requirements, in addition to convolution processing, other processing methods can also be combined, as long as new image data and new nDSM information of two cross-modal types that meet the requirements of the fusion process are obtained.

[0213] In an implementable solution, the device further includes:

[0214] A fusion level update module, which is used to determine whether the current intermediate fusion result has reached a first preset processing condition during the hierarchical fusion process. If not, the total number of fusion levels is increased to a first fusion level; if it has reached and has not reached the preset number of fusion levels, the total number of fusion levels is decreased to a second fusion level.

[0215] By presetting the preset number of fusion levels, it is considered that the number of fusion processes corresponding to different actual scenarios can be different. For example, the preset number of fusion levels is 5. When the fusion process reaches the 3rd level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result indicates that a satisfactory ground object contour effect can be obtained, the fusion process can be stopped, and at this time, the number of fusion levels is automatically and dynamically updated to 3; conversely, when the fusion process reaches the 5th level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result still indicates that a satisfactory ground object contour effect cannot be obtained, the fusion process can be increased, and at this time, the number of fusion levels is automatically and dynamically updated to 7; when the fusion process reaches the 5th level, the corresponding intermediate fusion result is automatically searched. If the current intermediate fusion result indicates that a satisfactory ground object contour effect can be obtained, there is no need to adjust the number of fusion levels at this time.

[0216] In this solution, by dynamically adjusting the number of fusion levels according to the actual situation, the flexibility and rationality of the overall processing are effectively improved, which not only ensures the effective and reliable acquisition of ground object information but also avoids the waste of unnecessary computing resources.

[0217] In an implementable solution, the device further includes:

[0218] A boundary information acquisition module, which is used to obtain a number of boundary points corresponding to the contour of the ground object and the position information corresponding to the boundary points based on the contour information.

[0219] Among them, the Alpha-Shape method is used to determine the boundary points on the contour of the ground object; of course, other implementable methods can also be used, which will not be elaborated here.

[0220] A contour information update module, configured to update the contour information according to the position information.

[0221] In this solution, based on the contour information obtained through the above process, the boundary points on these contours and their positions are obtained, and then the positions where the contours of each ground object are located are optimized, improving the accuracy of determining the contour information of the ground object.

[0222] In an implementable solution, the contour information update module includes:

[0223] A segmented acquisition unit, configured to acquire a plurality of segments according to the position information of a plurality of boundary points;

[0224] Among them, the boundary points are classified according to the geometric relationship of the boundary points based on the position information. The boundary points generally in one direction correspond to a line segment, that is, the distance difference between two boundary points on the same line segment is less than a preset value.

[0225] A line fitting unit, configured to perform fitting processing on the position information of multiple boundary points on each segment to obtain the corresponding line;

[0226] A contour information update unit, configured to obtain the boundary line corresponding to the contour of the ground object based on a plurality of connected lines, and use the boundary information corresponding to the boundary line as the contour information.

[0227] In this solution, the initial contour information segmented based on two cross-modal representation information is improved. Each line intersects to form the boundary line of the ground object contour as the final contour, more finely and accurately segmenting the contour of the ground object to meet the usage scenarios with higher requirements.

[0228] In an implementable solution, the device further includes:

[0229] A main direction determination module, configured to determine the main direction of the ground object according to the position information of a plurality of boundary points;

[0230] A regularization processing module, configured to perform regularization processing on the line according to the main direction to update and obtain a line that meets the second preset processing condition.

[0231] In this solution, through the main direction of the ground object, the drawn lines are regularized. For those with too large a deviation from the main direction, adjustment is required to ensure the accuracy of the determined boundary line of the contour. For example, the corresponding line is adjusted to be parallel or perpendicular to the main direction; of course, for buildings with special shapes, synchronous adjustment can also be performed in combination with their special shapes, which will not be elaborated here.

[0232] Specifically, as Figure 6 shown, to identify several boundary points of the building outline determined based on the representation information of two cross-modalities; as Figure 7 shown, corresponding segments are obtained based on several boundary points; as Figure 8 shown, the straight line corresponding to each segment and the intersection points between different straight line connections are obtained; as Figure 9 shown, the boundary line of the outline formed by the connections of different straight lines.

[0233] In an implementable solution, when the ground object is a building, the device further includes:

[0234] A style information acquisition module, configured to acquire the style information of the inner roof of the building outline;

[0235] A height information acquisition module, configured to acquire the height information of the building based on the style information by using a matching preset height determination strategy.

[0236] In this solution, for the height information of different buildings, different height determination strategies are adopted according to the style of the roof, rather than a unified determination scheme, thus ensuring the accuracy of determining the height information of different ground objects.

[0237] In an implementable solution, the height information acquisition module is configured to, when the style information represents that the roof is a flat roof, acquire the first mode among the height values corresponding to each pixel point within the outline, and the frequency of occurrence of the first mode;

[0238] When the frequency is greater than or equal to the first set value, calculate the first mean value of all height values whose difference from the mode is less than the first distance, and use the first mean value as the height information of the building;

[0239] When the frequency is less than the first set value, calculate the second mean value of all height values whose difference from the first mode is less than the second distance, and use the second mean value as the height information of the building;

[0240] Wherein, the second distance is greater than the first distance.

[0241] For example: Based on the schematic diagram of the roof slope of the building, select buildings with flat roofs, and then take the mode and frequency according to the height (unit: meter). If the frequency ≥ 0.7, the height value of the building is taken as the mean value within the range of the mode ± 1 meter (unit: centimeter); if the frequency < 0.7, the height value of the building is taken as the mean value within the range of the mode ± 3 meters (unit: centimeter).

[0242] In this solution, for flat-roofed buildings, based on the mode and frequency among the height values of all pixel points within the outline, the corresponding height determination scheme is determined, thus ensuring the accuracy and reliability of obtaining the height information of flat-roofed buildings.

[0243] In an implementable solution, the height information acquisition module is configured to obtain a third mean value of the maximum height value and the minimum height value when the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is greater than or equal to a second set value, and use the third mean value as the height information of the building.

[0244] When the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is less than the second set value, obtain the first mode of the height values corresponding to each pixel point within the contour, eliminate the height values that differ from the first mode by more than a third distance, obtain the fourth mean value of the maximum height value and the minimum height value among the remaining all height values, and use the fourth mean value as the height information of the building.

[0245] For example: as Figure 11 shown, based on the schematic diagram of the roof surface building slope, select buildings with pitched roofs, domes, etc. for the roof, and then obtain whether the difference in height values corresponding to the pixel points within the calibration surface coordinate string is within 6m. If so, the height value of the building is taken as the mean value of the maximum height value and the minimum height value;

[0246] If not, take the mode according to the height values (unit: meter), eliminate the pixel points outside the range of ±3 meters from the mode, and take the mean value of the maximum and minimum values among all the height values after elimination as the height value of the building (unit: centimeter).

[0247] In this solution, for buildings with non-flat roofs (irregular convex structures such as pitched roofs and domes), based on the mode and frequency of the height values of all pixel points within the contour, determine the corresponding height determination scheme, thereby ensuring the accuracy and reliability of the height information acquisition of flat-roofed buildings.

[0248] In an implementable solution, the characterization information acquisition module 191 includes:

[0249] The first module acquisition unit is configured to obtain the digital surface model DSM corresponding to the satellite remote sensing image;

[0250] The second module acquisition unit is configured to obtain the corresponding normalized digital surface model nDSM based on the digital surface model DSM;

[0251] Among them, obtaining the normalized digital surface model nDSM from the digital surface model DSM belongs to the mature technology in this field and will not be elaborated here.

[0252] The first nDSM information acquisition unit is configured to obtain the first nDSM information according to the normalized digital surface model nDSM.

[0253] In this solution, a normalized digital surface model nDSM is obtained from the digital surface model DSM to obtain the first nDSM information, so as to ensure the timeliness and reliability of nDSM information acquisition.

[0254] In an implementable solution, the first module acquisition unit includes:

[0255] A stereo pair acquisition subunit, configured to acquire an original satellite image stereo pair;

[0256] Wherein, the original satellite image stereo pair is a set of satellite remote sensing images taken by a remote sensing satellite of the same area from different angles;

[0257] A segmentation processing subunit, configured to segment the original satellite image stereo pair to obtain a plurality of sub-satellite image stereo pairs with a preset resolution;

[0258] Wherein, since the data file of the original satellite image stereo pair is large, with a resolution of over 100 million pixels; for satellite image data with a resolution of 10000*10000, the coverage range can reach dozens of square kilometers, and it is difficult and inefficient to perform feature matching using it. Therefore, it is necessary to segment the satellite image stereo pair. It is possible to use a map segmentation with a resolution of 1024*1024 to segment the original satellite image stereo pair into dozens or even hundreds of small stereo pairs, improving the efficiency of subsequent feature extraction and feature matching. Specifically, see Figure 13 For a single original satellite image, see Figure 14 It is Figure 1 A map block cut out from the upper left corner, that is, a sub-satellite image stereo pair.

[0259] A feature extraction subunit, configured to process the sub-satellite image stereo pair by using a preset feature extraction method to obtain a target feature vector;

[0260] A candidate pixel point extraction subunit, configured to extract candidate feature pixel points based on the target feature vector;

[0261] A matching pixel point acquisition subunit, configured to obtain all matching pixel points corresponding to the sub-satellite image stereo pair according to the similarity between the candidate feature pixel points;

[0262] An elevation information calculation subunit, configured to calculate the elevation information corresponding to each matching pixel point;

[0263] The first module acquisition subunit, configured to obtain the digital surface model DSM according to the elevation information.

[0264] Feature extraction is performed on the sub-satellite image stereo pair respectively to obtain candidate feature pixel points and generate feature descriptors. The feature descriptors composed of multi-dimensional vectors contain the scale where the descriptor is located, the position of the descriptor, and the feature information of the descriptor.

[0265] In the feature similarity matching stage, the Euclidean distance of the feature vectors is used to calculate the similarity between the feature points in two images, and the KD-tree algorithm is used to improve the detection speed; the sub-satellite image stereo pair includes Image 1 and Image 2. Take a certain feature point A in Image 1, and find two feature points, the point B with the closest feature vector distance and the second-closest point C in Image 2. Use d(A,B) to represent the similarity between A and B. If the ratio between d(A,B) and d(A,C) is less than a certain set threshold, then determine this pair of feature points A and B with the closest feature vector distance as matching points; and so on, to obtain multiple pairs of mutually matching feature points in the sub-satellite image stereo pair.

[0266] The fundamental matrix F and the homography matrix H are calculated using epipolar geometry, and the matching points of all points in the image are calculated, as Figure 15 shown. Specifically, O1 and O2 are the optical centers of the cameras in two positions respectively, P is a three-dimensional point in space, and p1 and p2 are the pixel points corresponding to point P on different imaging planes. The three points P, O1, and O2 are coplanar, which is called the epipolar plane. The intersection lines l1 and l2 of the epipolar plane and the image planes are the epipolar lines of the two image planes; according to the epipolar geometry constraint, p2 T Fp1 = 0, where F is the fundamental matrix. The epipolar constraint contains translation and rotation, and gives the spatial position relationship of two matching points.

[0267] In addition, for the homography matrix H that describes the mapping relationship between planes, there is At least 4 pairs of mutually matching feature points are used and substituted into the transformation formula to obtain a set of linear equations, and then the fundamental matrix F and the homography matrix H are calculated; for the matrices F and H, the matrix with a small reprojection error is selected as the final estimated matrix; the estimated matrix is used to calculate the pixel points in Image 2 corresponding to all pixel points in Image 1 to determine the one-to-one matching of the corresponding pixel points, and all matching pixel points are obtained.

[0268] Among them, the elevation information corresponding to each matching pixel point is obtained by using the RPC model. Of course, other implementable methods can also be used, which will not be elaborated here.

[0269] Linear interpolation encryption is performed to generate a point cloud file; specifically, after obtaining the elevation information h of each matching pixel point, according to the actual accuracy requirements, linear interpolation encryption is performed on the pixel points in the sub-satellite image stereo pair, and the original planar pixel points are converted into three-dimensional point pixel points in space, thereby generating a three-dimensional point cloud file, as Figure 16As shown; sample the three-dimensional point cloud file at a certain interval, and use its height value as the grayscale value of the pixel points to obtain the final tif file, such as Figure 17 shown, which is the required digital surface model DSM.

[0270] In this solution, a deep learning model is used to extract feature points and perform image feature matching. Without ground control points, three-dimensional point clouds are automatically reconstructed from satellite image stereo pairs; specifically, a convolutional-based deep learning model is used to extract satellite image features, and then an image matching algorithm is used to find mutually matching feature points in the satellite image stereo pair; then, the epipolar geometry is applied to obtain the mapping relationship of the satellite image stereo pair, and all image points in the satellite image are mapped, and then used to calculate the elevation information of all points in the image, generating a dense three-dimensional point cloud, thereby ensuring the accuracy and reliability of obtaining the digital surface model DSM.

[0271] In an implementable solution, the first module acquisition subunit is used to perform linear interpolation encryption processing on each matching pixel point to convert each two-dimensional matching pixel point into a three-dimensional matching pixel point, generate a three-dimensional point cloud file based on the three-dimensional matching pixel points, use the elevation information corresponding to each three-dimensional matching pixel point as the grayscale value, and perform sampling processing on the three-dimensional point cloud file to obtain the digital surface model DSM.

[0272] In this solution, based on high-resolution satellite image stereo pairs, the satellite image stereo pairs are divided into tiles; then, through key steps such as feature matching, stereo rectification, and triangulation on the divided satellite image stereo pairs, the method for photogrammetric three-dimensional reconstruction is applied to large-scale satellite images, and the elevation information of each pair of matching points is obtained from the RPC model, automatically generating a real-scene three-dimensional point cloud model, and thus calculating the digital surface model DSM, thereby ensuring the accuracy and reliability of obtaining the digital surface model DSM.

[0273] In an implementable solution, the preset feature extraction method includes downsampling processing, consecutive four convolutional pooling processing, and feature compression processing performed in sequence.

[0274] Specifically, as Figure 18 shown, an image feature extraction model is constructed, and the sub-satellite image stereo pair of the tiled satellite image with a resolution of 256*256 after downsampling is input, and then passed through 4 convolutional pooling modules to output features of 16*16*512; then the PCA method is used to compress the features output by the model to retain preset key information, and finally a 128-dimensional feature vector is obtained.

[0275] In this solution, the above-mentioned feature extraction is adopted, which ensures the accuracy of feature extraction, makes the digital surface model (DSM) more accurate and reliable, and then also ensures the accuracy of obtaining ground object information.

[0276] Embodiment 3

[0277] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0278] Figure 20 The schematic block diagram of an example electronic device 2000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0279] As Figure 20 shown, the device 2000 includes a computing unit 2001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 2002 or the computer program loaded from the storage unit 2008 into the random access memory (RAM) 2003. In the RAM 2003, various programs and data required for the operation of the device 2000 can also be stored. The computing unit 2001, the ROM 2002, and the RAM 2003 are connected to each other through a bus 2004. The input / output (I / O) interface 2005 is also connected to the bus 2004.

[0280] A plurality of components in the device 2000 are connected to the I / O interface 2005, including: an input unit 2006, such as a keyboard, a mouse, etc.; an output unit 2007, such as various types of displays, speakers, etc.; a storage unit 2008, such as a magnetic disk, an optical disc, etc.; and a communication unit 2009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 2009 allows the device 2000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0281] The computing unit 2001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 2001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 2001 executes the various methods and processes described above, such as the above-mentioned method. For example, in some embodiments, the above-mentioned method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 2008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 2000 via the ROM 2002 and / or the communication unit 2009. When the computer program is loaded into the RAM 2003 and executed by the computing unit 2001, one or more steps of the above-mentioned method described above can be executed. Alternatively, in other embodiments, the computing unit 2001 can be configured to execute the above-mentioned method in any other suitable manner (e.g., by means of firmware).

[0282] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0283] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0284] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0285] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0286] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0287] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0288] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0289] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A method for obtaining ground object information, the method comprising: Obtaining two cross-modal representation information corresponding to a satellite remote sensing image; Based on the two representation information, performing fusion processing in a step-by-step fusion processing manner to obtain a target fusion result, and according to the target fusion result, obtaining contour information of the contour of each ground object in the satellite remote sensing image; Wherein, the step of performing fusion processing in a step-by-step fusion processing manner to obtain a target fusion result includes: Obtaining a weight parameter corresponding to each of the representation information at the current level; The weight parameter is used to characterize the influence degree of each of the representation information on the position where the contour of the ground object is located; Fusing the two representation information based on their respective weight parameters to obtain an intermediate fusion result; Generating new representation information corresponding to the next level based on the intermediate fusion result, and re-executing the step of obtaining the weight parameter corresponding to each of the representation information at the current level until the intermediate fusion result corresponding to the last level is obtained and used as the target fusion result; The method further includes: During the step-by-step fusion processing, determining whether the current intermediate fusion result has reached a first preset processing condition. If not, increasing the total number of fusion levels to a first fusion level; If it has reached and has not reached the preset fusion level, reducing the total number of fusion levels to a second fusion level.

2. The method according to claim 1, wherein the two cross-modal representation information includes first image information of the satellite remote sensing image and first nDSM information in the normalized digital surface model nDSM that matches the satellite remote sensing image.

3. The method according to claim 2, wherein The step of obtaining a weight parameter corresponding to each of the representation information at the current level includes: Using global average pooling to respectively obtain first global information of the first image information and second global information of the first nDSM information; Obtaining a first attention vector corresponding to the first global information and a second attention vector corresponding to the second global information.

4. The method according to claim 3, wherein, The step of fusing the two representation information based on their respective weight parameters to obtain an intermediate fusion result includes: Multiplying the first attention vector by the first image information to obtain second image information; Fusing the second image information with the first nDSM information to obtain a first fusion feature; Multiplying the second attention vector by the first nDSM information to obtain second nDSM information; Fusing the second nDSM information with the first image information to obtain a second fusion feature; Concatenating the first fusion feature and the second fusion feature to obtain the intermediate fusion result.

5. The method according to claim 3, wherein After the step of obtaining a first attention vector corresponding to the first global information and a second attention vector corresponding to the second global information, it further includes: Performing noise filtering processing on the first attention vector and the second attention vector.

6. The method according to claim 5, wherein, The step of performing noise filtering processing on the first attention vector and the second attention vector includes: Obtaining a weight value corresponding to each pixel channel in the satellite remote sensing image; Multiply each first value corresponding to a pixel channel in the first attention vector by the corresponding weight value to filter noise from the first attention vector; Multiply each second value corresponding to a pixel channel in the second attention vector by the corresponding weight value to filter noise from the second attention vector.

7. The method according to claim 1, wherein The step of generating new representation information corresponding to the next level based on the intermediate fusion result includes: Performing a convolution process on the intermediate fusion result to obtain two new types of cross-modal representation information; Among them, the representation information at each level is different and the corresponding information sizes are the same.

8. The method according to any one of claims 1-7, after the step of obtaining the contour information of each feature in the satellite remote sensing image, further includes: Based on the contour information, obtaining a plurality of boundary points corresponding to the contour of the feature and the position information corresponding to the boundary points; Updating the contour information according to the position information.

9. The method according to claim 8, wherein, The step of updating the contour information according to the position information includes: Obtaining a plurality of segments according to the position information of a plurality of the boundary points; Performing a fitting process on the position information of a plurality of boundary points on each segment to obtain a corresponding straight line; Based on a plurality of connected straight lines, obtaining a boundary line corresponding to the contour of the feature, and using the boundary information corresponding to the boundary line as the contour information.

10. The method according to claim 9, the method further includes: Determining the main direction of the feature according to the position information of a plurality of the boundary points; Among them, before the step of based on a plurality of connected straight lines, further includes: Regularizing the straight lines according to the main direction to update and obtain the straight lines that meet the second preset processing condition.

11. The method according to claim 2, when the feature is a building, the method further includes: Obtaining the style information of the roof inside the contour of the building; Based on the style information, adopting a matching preset height determination strategy to obtain the height information of the building.

12. The method according to claim 11, wherein, The step of adopting a matching preset height determination strategy to obtain the height information of the building based on the style information includes: When the style information represents that the roof is a flat roof, obtaining the first mode of the height values corresponding to each pixel point inside the contour and the frequency of occurrence of the first mode; When the frequency is greater than or equal to a first set value, calculating a first mean value of all height values whose difference from the first mode is less than a first distance, and using the first mean value as the height information of the building; When the frequency is less than the first set value, calculating a second mean value of all height values whose difference from the first mode is less than a second distance, and using the second mean value as the height information of the building; Wherein, the second distance is greater than the first distance.

13. The method according to claim 11, wherein, The step of adopting a matching preset height determination strategy to obtain the height information of the building based on the style information includes: When the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is greater than or equal to a second set value, obtain a third mean value of the maximum height value and the minimum height value, and use the third mean value as the height information of the building; When the style information represents that the roof is not flat and the difference in height values at different pixel points within the contour is less than the second set value, obtain a first mode value among the height values corresponding to each pixel point within the contour, eliminate the height values that differ from the first mode value by more than a third distance, obtain a fourth mean value of the maximum height value and the minimum height value among the remaining all height values, and use the fourth mean value as the height information of the building.

14. The method according to claim 2, wherein the step of obtaining the first nDSM information comprises: Obtain a digital surface model DSM corresponding to the satellite remote sensing image; Obtain a corresponding normalized digital surface model nDSM based on the digital surface model DSM; Obtain the first nDSM information according to the normalized digital surface model nDSM.

15. The method according to claim 14, wherein, The step of obtaining the digital surface model DSM corresponding to the satellite remote sensing image comprises: Obtain a stereo pair of original satellite images; Wherein, the stereo pair of original satellite images is a set of the satellite remote sensing images taken by a remote sensing satellite from different angles of the same area; Segment the stereo pair of original satellite images to obtain a plurality of stereo pairs of sub-satellite images with a preset resolution; Process the stereo pairs of sub-satellite images by using a preset feature extraction method to obtain target feature vectors; Extract candidate feature pixel points based on the target feature vectors; Obtain all matching pixel points corresponding to the stereo pairs of sub-satellite images according to the similarity between the candidate feature pixel points; Calculate the elevation information corresponding to each of the matching pixel points; Obtain the digital surface model DSM according to the elevation information.

16. The method according to claim 15, wherein, The step of obtaining the digital surface model DSM according to the elevation information comprises: Perform linear difference encryption processing on each of the matching pixel points to convert each two-dimensional matching pixel point into a three-dimensional matching pixel point; Generate a three-dimensional point cloud file based on the three-dimensional matching pixel points; Use the elevation information corresponding to each of the three-dimensional matching pixel points as a gray value, and perform sampling processing on the three-dimensional point cloud file to obtain the digital surface model DSM.

17. The method according to claim 15, wherein the preset feature extraction method comprises downsampling processing, four consecutive convolutional pooling processes, and feature compression processing performed in sequence.

18. An apparatus for obtaining ground object information, the apparatus comprising: A characterization information acquisition module, configured to acquire two cross-modal characterization information corresponding to a satellite remote sensing image; A fusion processing module, configured to perform fusion processing in a hierarchical fusion processing manner based on the two characterization information to obtain a target fusion result; A contour information acquisition module, configured to acquire contour information of the contour of each ground object in the satellite remote sensing image according to the target fusion result; Among them, the fusion processing module includes: A weight parameter acquisition unit, configured to acquire weight parameters corresponding to each type of the characterization information at the current level; The weight parameters are used to characterize the influence degree of each type of the characterization information on the position where the contour of the ground object is located; A fusion processing unit, configured to fuse the two types of the characterization information based on their respective weight parameters to obtain an intermediate fusion result; A characterization information generation unit, configured to generate new characterization information corresponding to the next level based on the intermediate fusion result, and call the weight parameter acquisition unit until the fusion processing unit obtains the intermediate fusion result corresponding to the last level and use it as the target fusion result; The device further includes: A fusion level update module, configured to determine whether the current intermediate fusion result has reached a first preset processing condition during the process of hierarchical fusion processing. If not, increase the total fusion level to a first fusion level; if it has reached and has not reached the preset fusion level, decrease the total fusion level to a second fusion level.

19. The device according to claim 18, wherein the two cross-modal characterization information includes first image information of the satellite remote sensing image and first nDSM information in the normalized digital surface model nDSM that matches the satellite remote sensing image.

20. The device according to claim 19, wherein the weight parameter acquisition unit includes: A first information acquisition subunit, configured to acquire first global information of the first image information by using global average pooling; A second information acquisition subunit, configured to acquire second global information of the first nDSM information by using global average pooling; A first vector acquisition subunit, configured to acquire a first attention vector corresponding to the first global information; A second vector acquisition subunit, configured to acquire a second attention vector corresponding to the second global information.

21. The device according to claim 20, wherein the fusion processing unit includes A first multiplication processing subunit, configured to multiply the first attention vector by the first image information to obtain second image information; A first feature acquisition subunit, configured to fuse the second image information with the first nDSM information to obtain a first fusion feature; A second multiplication processing subunit, configured to multiply the second attention vector by the first nDSM information to obtain second nDSM information; A second feature acquisition subunit, configured to fuse the second nDSM information with the first image information to obtain a second fusion feature; A fusion processing subunit, configured to splice the first fusion feature and the second fusion feature to obtain the intermediate fusion result.

22. The apparatus according to claim 20, wherein the fusion processing module further comprises: A noise processing unit, configured to perform noise filtering processing on the first attention vector and the second attention vector.

23. The device according to claim 22, wherein the noise processing unit includes: A weight value acquisition subunit, configured to acquire a weight value corresponding to each pixel channel in the satellite remote sensing image; The first noise processing sub-unit is configured to multiply each first value corresponding to a pixel channel in the first attention vector by the corresponding weight value to perform noise filtering processing on the first attention vector; The second noise processing sub-unit is configured to multiply each second value corresponding to a pixel channel in the second attention vector by the corresponding weight value to perform noise filtering processing on the second attention vector.

24. The apparatus according to claim 18, wherein the characterization information generation unit is configured to perform a convolution process on the intermediate fusion result to obtain two new pieces of cross-modal characterization information; Among them, The characterization information at each level is different and the corresponding information sizes are the same.

25. The apparatus according to any one of claims 18-24, wherein the apparatus further comprises: A boundary information acquisition module, configured to acquire a plurality of boundary points corresponding to the contour of the ground object and position information corresponding to the boundary points based on the contour information; A contour information update module, configured to update the contour information according to the position information.

26. The apparatus according to claim 25, wherein the contour information update module comprises: A segmentation acquisition unit, configured to acquire a plurality of segments according to the position information of the plurality of boundary points; A straight line fitting unit, configured to perform a fitting process on the position information of a plurality of boundary points on each segment to obtain a corresponding straight line; A contour information update unit, configured to obtain a boundary line corresponding to the contour of the ground object based on a plurality of connected straight lines, and use the boundary information corresponding to the boundary line as the contour information.

27. The apparatus according to claim 26, wherein the apparatus further comprises: A main direction determination module, configured to determine the main direction of the ground object according to the position information of the plurality of boundary points; A regularization processing module, configured to perform a regularization process on the straight line according to the main direction to update and obtain the straight line that satisfies the second preset processing condition.

28. The apparatus according to claim 19, when the ground object is a building, the apparatus further comprises: A style information acquisition module, configured to acquire the style information of the roof within the contour of the building; A height information acquisition module, configured to acquire the height information of the building by adopting a matching preset height determination strategy based on the style information.

29. The apparatus according to claim 28, wherein the height information acquisition module is configured to, when the style information represents that the roof is a flat roof, acquire the first mode among the height values corresponding to each pixel point within the contour and the frequency of occurrence of the first mode; When the frequency is greater than or equal to a first set value, calculate a first mean value of all height values that differ from the first mode by less than a first distance, and use the first mean value as the height information of the building; When the frequency is less than the first set value, calculate a second mean value of all height values that differ from the first mode by less than a second distance, and use the second mean value as the height information of the building; Among them, The second distance is greater than the first distance.

30. The device of claim 28, wherein the height information acquisition module is used to acquire a third mean of a maximum height value and a minimum height value when the style information indicates that the roof is non-flat and the difference in height values at different pixel points within the outline is greater than or equal to a second set value, and use the third mean as the height information of the building; When the style information indicates that the roof is non-flat, and the difference in height values at different pixel points within the outline is less than the second set value, a first mode among the height values corresponding to each pixel point in the outline is obtained, and height values that differ from the first mode by more than a third distance are eliminated, and a fourth mean of the maximum height value and the minimum height value among all the remaining height values is obtained, and the fourth mean is used as the height information of the building.

31. The apparatus according to claim 19, wherein the characterization information acquisition module comprises: The first module acquisition unit is used to acquire a digital surface model DSM corresponding to the satellite remote sensing image; A second module acquisition unit, configured to acquire a corresponding normalized digital surface model nDSM based on the digital surface model DSM; The first nDSM information acquiring unit is configured to acquire the first nDSM information according to the normalized digital surface model nDSM.

32. The apparatus according to claim 31, wherein the first module acquisition unit comprises: A stereo image pair acquisition subunit, used for acquiring a stereo image pair of original satellite images; The original satellite image stereo pair is a set of satellite remote sensing images taken by a remote sensing satellite from different angles on the same area; A segmentation processing subunit, used for segmenting the original satellite image stereo pair to obtain a plurality of sub-satellite image stereo pairs with a preset resolution; A feature extraction subunit, used to process the sub-satellite image stereo pair using a preset feature extraction method to obtain a target feature vector; A candidate pixel point extraction subunit is used to extract candidate feature pixel points based on the target feature vector; A matching pixel point acquisition subunit is used to acquire all matching pixel points corresponding to the stereo image pair of the sub-satellite images according to the similarity between the candidate feature pixel points; The elevation information calculation subunit is used to calculate the elevation information corresponding to each matching pixel point; The first module is an acquisition subunit, which is used to acquire the digital surface model DSM according to the elevation information.

33. In the device as described in claim 32, the first module acquisition subunit is used to perform linear difference encryption processing on each of the matching pixel points to convert each two-dimensional matching pixel point into a three-dimensional matching pixel point, generate a three-dimensional point cloud file based on the three-dimensional matching pixel points, use the elevation information corresponding to each of the three-dimensional matching pixel points as a grayscale value, and sample the three-dimensional point cloud file to obtain the digital surface model DSM.

34. The device as described in claim 32, wherein the preset feature extraction method includes a downsampling process, four consecutive convolution pooling processes and a feature compression process performed in sequence.

35. An electronic device, comprising: at least one processor; and a memory communicatively coupled to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, enable the at least one processor to perform the method according to any one of claims 1-17.

36. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are for causing the computer to perform the method according to any one of claims 1-17.

37. A computer program product, comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-17.

Citation Information

Patent Citations

  • Remote sensing data ground feature element refined classification method and device

    CN111582102A

  • Remote sensing image semantic segmentation method and device

    CN111582104A