Wetland water body remote sensing identification system based on large model analysis
By combining feature region analysis and weighted fusion of visible light and near-infrared remote sensing images, the problem of misidentification of wetland water bodies in complex backgrounds was solved, achieving higher accuracy and robustness in wetland water body identification.
Patent Information
- Application Number
- CN202511299614.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing wetland water body remote sensing identification technologies based on large models are prone to misidentifying vegetation, soil, and other water bodies as water bodies in complex backgrounds, leading to reduced identification accuracy.
A method combining visible light and near-infrared remote sensing images is used to obtain confidence level and relative reflectance through feature region segmentation, grayscale distribution, texture features and contour morphology feature analysis, weight analysis and image fusion, and wetland water body identification is performed using convolutional neural networks.
It improves the accuracy and robustness of wetland water body identification, reduces the amount of computation, enhances the real-time performance and computational efficiency of the system, generates more accurate fusion results, and enhances the accuracy of wetland water body identification.
Smart Images

Figure CN120808179B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image analysis technology, specifically to a wetland water body remote sensing identification system based on large model analysis. Background Technology
[0002] Wetlands, as an important part of the Earth's ecosystem, play multiple functions, including regulating climate, purifying water quality, providing habitats for organisms, and protecting soil and water. Remote sensing imagery can quickly acquire information on the distribution, area, and changes of wetland water bodies. To better process this large-scale, high-dimensional remote sensing data, large-scale modeling technology has gradually become one of the key technologies for wetland water body remote sensing identification. Large-scale models can efficiently extract spatial and spectral features from images, improving the accuracy and robustness of wetland water body identification.
[0003] Currently, when identifying water bodies in remote sensing images based on large models, "attention" is evenly distributed across all areas of the image. However, since wetland water bodies usually coexist with surrounding vegetation, soil, and other water bodies, the boundaries of water bodies in remote sensing images are not always clear and are often affected by complex backgrounds. Therefore, traditional large model identification may result in complex backgrounds (such as weeds, buildings, and roads) being misidentified as water bodies, thereby reducing the accuracy of water body identification. Summary of the Invention
[0004] This invention provides a wetland water body remote sensing identification system based on large model analysis to solve existing problems.
[0005] The wetland water body remote sensing identification system based on large model analysis of the present invention adopts the following technical solution:
[0006] In a first aspect, one embodiment of the present invention provides a wetland water body remote sensing identification system based on large model analysis, the system comprising the following modules:
[0007] The remote sensing acquisition module is used to acquire remote sensing images and ground reflectance, wherein the remote sensing images include visible light remote sensing images and near-infrared remote sensing images;
[0008] The feature region module is used to divide the remote sensing image into several feature regions. The feature regions are divided into visible light feature regions and near-infrared feature regions. The confidence level of the feature regions is obtained by using the gray-scale distribution, texture features, and contour morphology features of the feature regions. The relative reflectance of the near-infrared feature regions is obtained by using the ground reflectance within the near-infrared feature regions.
[0009] The weighting analysis module is used to match visible light feature regions and near-infrared feature regions at the same time to form matching groups, and to perform weighting analysis based on the confidence of visible light feature regions and the relative reflectance of near-infrared feature regions in the matching groups to determine the fusion weights.
[0010] The remote sensing recognition module is used to perform image fusion on the feature regions in the matching group using fusion weights, and uses the fusion result as the input of the neural network to output the wetland water body recognition result.
[0011] Optionally, the specific method for dividing the remote sensing image region into several feature regions includes:
[0012] For visible light remote sensing images and near-infrared remote sensing images, the ISODATA algorithm is used to cluster the pixels in the images. Several clusters are obtained in the visible light remote sensing images and near-infrared remote sensing images. The pixels contained in each cluster form a region in the image, which is called the feature region. The feature region in the visible light remote sensing image is called the visible light feature region, and the feature region in the near-infrared remote sensing image is called the near-infrared feature region.
[0013] Optionally, the method for obtaining the confidence level of a feature region by utilizing its grayscale distribution, texture features, and contour morphology features includes:
[0014] For any visible light feature region in a visible light remote sensing image, the gray-level co-occurrence matrix of the visible light feature region is obtained, and the entropy value of all elements in the gray-level co-occurrence matrix is used as the texture information entropy of the visible light feature region. Based on the difference between the average gray value of the visible light feature region and the average gray value of all visible light feature regions, the texture information entropy of the visible light feature region, and the fractal dimension of the visible light feature region, the confidence level of the visible light feature region in the visible light remote sensing image is obtained. The difference is positively correlated with the confidence level, and the texture information entropy and the fractal dimension are both negatively correlated with the confidence level. The confidence level of any near-infrared feature region in a near-infrared remote sensing image is obtained using the same method as that for the visible light feature region.
[0015] Optionally, the method for obtaining the relative reflectance of the near-infrared feature region by utilizing the ground reflectance within the near-infrared feature region includes:
[0016] A preset reflectance threshold is set, and the ground reflectance of any near-infrared feature region in the near-infrared remote sensing image is obtained. The ratio of the reflectance threshold to the ground reflectance of the near-infrared feature region in the near-infrared remote sensing image is normalized to obtain the relative reflectance of the near-infrared feature region.
[0017] Optionally, the specific method for matching visible light feature regions and near-infrared feature regions at the same time to form a matching group includes:
[0018] By utilizing the spatial distance and overlap between visible light feature regions in visible light remote sensing images and near-infrared feature regions in near-infrared remote sensing images at the same time, the matching degree between visible light feature regions and near-infrared feature regions is obtained. Several matching groups are then determined using the matching degree, and each matching group consists of a visible light feature region and a near-infrared feature region.
[0019] Optionally, the specific method for obtaining the matching degree is as follows:
[0020] For visible light remote sensing images and near-infrared remote sensing images at the same time, the centroids of the visible light feature regions in the visible light remote sensing image and the near-infrared feature regions in the near-infrared remote sensing image are obtained respectively, and the Euclidean distance between the corresponding coordinates of the centroids is obtained. The sets formed by the coordinates of all pixels contained in any visible light feature region and any near-infrared feature region are obtained respectively, denoted as the first region set and the second region set. The intersection of the first region set and the second region set is obtained, and the number of elements in the intersection is used as the overlap degree of the visible light feature regions and near-infrared feature regions corresponding to the first region set and the second region set respectively. Combining the Euclidean distance between the corresponding coordinates of the centroids of any visible light feature region and any near-infrared feature region and the overlap degree, the matching degree of the visible light feature region and the near-infrared feature region is calculated. The Euclidean distance is negatively correlated with the matching degree, and the overlap degree is positively correlated with the matching degree.
[0021] Optionally, the method for determining the fusion weight based on the confidence level of the visible light feature region and the relative reflectance of the near-infrared feature region in the matching group through weight analysis includes:
[0022] The confidence level of the visible light feature region is used as the spatial attention weight of the corresponding visible light feature region.
[0023] The relative reflectance of the near-infrared feature region is used as the spatial attention weight of the corresponding near-infrared feature region.
[0024] The confidence level is used to divide the water body region into water body region and non-water body region in the remote sensing image. The discrimination coefficient of the corresponding remote sensing image is calculated based on the difference in spatial attention weight level of all water body regions and non-water body regions in the remote sensing image.
[0025] For any matching group, the comprehensive spatial attention weight of the matching group is calculated based on the spatial attention weight of the visible light feature region in the matching group, the discrimination coefficient of the visible light remote sensing image to which the visible light feature region belongs, the spatial attention weight of the near-infrared feature region, and the discrimination coefficient of the near-infrared remote sensing image to which the near-infrared feature region belongs.
[0026] Obtain the comprehensive spatial attention weights of the same matching group at different times, and form a corresponding sequence in chronological order, denoted as the attention weight sequence. Calculate the weight stability of each matching group by using the differences in the comprehensive spatial attention weights of the matching groups in the attention weight sequence.
[0027] Obtain the maximum comprehensive spatial attention weight and the mean of all comprehensive spatial attention weights in the attention weight sequence. Obtain the difference between the comprehensive spatial attention weight of any matching group in the attention weight sequence and the maximum comprehensive spatial attention weight and the mean. Calculate the weight stability of the matching group. For any attention weight sequence, obtain the comprehensive spatial attention weight of the matching group with the maximum weight stability as the fusion weight.
[0028] Optionally, the specific method for obtaining the discrimination coefficient of the remote sensing image is as follows:
[0029] For any remote sensing image, feature regions with a confidence level greater than or equal to a preset confidence threshold are designated as water bodies, and feature regions outside the water bodies are designated as non-water bodies. The water and non-water body regions in the remote sensing image are obtained. The average spatial attention weight of all water bodies in the remote sensing image is obtained as the water body attention weight of the remote sensing image, and the average spatial attention weight of all non-water body regions in the remote sensing image is obtained as the non-water body attention weight of the remote sensing image. Based on the difference between the water body attention weight and the non-water body attention weight, the discrimination coefficient of the remote sensing image is obtained.
[0030] Optionally, the specific method for obtaining the weight stability of the matching group is as follows:
[0031] For any two time points, the first set of visible light feature regions and the second set of near-infrared feature regions contained in all matching groups are used as the total set of the matching groups. The intersection-union ratio (IUR) of the total sets of any two matching groups at different time points is obtained. When the IUR is the largest, the two matching groups are the same matching groups at different time points. The sequence of comprehensive spatial attention weights corresponding to the same matching groups at all consecutive time points is obtained as the attention weight sequence.
[0032] Obtain the maximum comprehensive spatial attention weight and the mean of all comprehensive spatial attention weights in the attention weight sequence. Obtain the difference between the comprehensive spatial attention weight of any matching group in the attention weight sequence and the maximum comprehensive spatial attention weight and the mean. Calculate the weight stability of the matching group.
[0033] Optionally, the method of using fusion weights to perform image fusion on the feature regions in the matching group, using the fusion result as the input to the neural network, and outputting the wetland water body identification result includes the following specific methods:
[0034] For any sequence of attention weights, the matching group corresponding to the elements contained in the sequence of attention weights with the highest weight stability is obtained as the target matching group. The comprehensive spatial attention weight of the target matching group is used as the fusion weight in the alpha fusion algorithm, thereby fusing the visible light feature region and the near-infrared feature region in the target matching group to obtain the fused region. All fused regions are obtained and used as the input of the trained convolutional neural network. The convolutional neural network is used to identify wetland water bodies, and the identification result is output by the convolutional neural network.
[0035] The beneficial effects of the technical solution of this invention are as follows: Visible light images and near-infrared images represent different dimensions of information. Visible light images reflect the visible area that the human eye can perceive, while near-infrared images can reveal the reflective properties of different materials on the ground, and are particularly sensitive to the identification of wetland water bodies. By fusing the two, not only can the shortcomings of a single image source be compensated, but the accuracy and robustness of wetland water body identification can also be effectively improved. In addition, during the image recognition process, dividing the image into feature regions can effectively reduce the amount of computation, and processing the data of different feature regions separately avoids redundant calculations on the entire image. This approach enables more efficient extraction and analysis of key features with limited computing resources, improving system real-time performance and computational efficiency. By combining different image features (such as grayscale distribution, texture, and contour morphology), more discriminative information can be introduced into the analysis of feature regions, improving accuracy. Furthermore, confidence assessment allows the system to adjust its judgment strategy based on the complexity of image features, making it more flexible and precise. Further, by matching visible light feature regions with near-infrared feature regions and determining fusion weights through weight analysis, the scheme can weight the image based on the confidence and reflectivity of different regions, thereby generating a more accurate fusion result. This effectively enhances the accuracy of wetland water body identification in the image fusion results. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a structural block diagram of the wetland water body remote sensing identification system based on large model analysis according to the present invention. Detailed Implementation
[0038] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of the wetland water body remote sensing identification system based on large model analysis proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0040] The following description, in conjunction with the accompanying drawings, details the specific scheme of the wetland water body remote sensing identification system based on large model analysis provided by this invention.
[0041] Please see Figure 1 The diagram illustrates a structural block diagram of a wetland water body remote sensing identification system based on large model analysis, according to an embodiment of the present invention. The system includes the following modules:
[0042] The remote sensing acquisition module 101 is used to acquire remote sensing images and ground reflectance.
[0043] It should be noted that Convolutional Neural Networks (CNNs) are one of the most common large-scale deep learning models for processing remote sensing images. They can automatically extract spatial features from images to identify the boundaries, morphology, and other features of wetland water bodies. CNNs typically use labeled remote sensing images for supervised learning to achieve accurate water body extraction. Since wetland water bodies often coexist with surrounding vegetation, soil, and other water bodies, the boundaries of water bodies in remote sensing images are affected by complex backgrounds. The uniform attention mechanism used in current large-scale models can lead to identification errors in water bodies against these complex backgrounds. Therefore, this embodiment of the invention introduces multispectral remote sensing images for accurate identification. Furthermore, since most water bodies (especially still water bodies or wetland water bodies) have near-zero or extremely low reflectance in the near-infrared band, while vegetation or bare land has very high reflectance in the near-infrared band, they can be distinguished more significantly.
[0044] To implement the wetland water body remote sensing identification system based on large model analysis proposed in this embodiment, it is first necessary to acquire remote sensing images of the wetland. The specific process is as follows:
[0045] First, visible light remote sensing images and near-infrared remote sensing images of any wetland area are acquired using a preset fixed sampling interval. The visible light remote sensing images and near-infrared remote sensing images are collectively referred to as remote sensing images, and ground reflectance data are obtained.
[0046] It should be noted that, in the embodiments of the present invention, the resolution and position of the corresponding image are preset and guaranteed to be the same for each sampling based on experience.
[0047] Then, the visible light remote sensing images and near-infrared remote sensing images are filtered.
[0048] Thus, visible light remote sensing images and near-infrared remote sensing images were obtained using the methods described above.
[0049] The feature region module 102 is used to divide the remote sensing image region into several feature regions. The feature regions are divided into visible light feature regions and near-infrared feature regions. The confidence level of the feature regions is obtained by using the gray-scale distribution, texture features, and contour morphology features of the feature regions. The relative reflectance of the near-infrared feature regions is obtained by using the ground reflectance within the near-infrared feature regions.
[0050] It should be noted that the uniform attention mechanism used in current large-scale model recognition can lead to recognition errors for water bodies in the complex background described above. Therefore, this method analyzes the differences in water body characteristics across different regions in the multispectral image of this area to determine their spatial attention weights. Regions in the multispectral image are matched, and the contribution of different matched regions to the water body is determined, resulting in spatiotemporal fusion of attention weights. The unbalanced attention mechanism described in this embodiment assigns different weights to different regions of the image, allowing the large network model to focus on important regions while ignoring less important ones, thus improving the accuracy of water body recognition. Therefore, this step first determines the regional spatial attention weights for the multispectral image of this area. Furthermore, for multispectral images of wetland water bodies, since water areas differ from vegetation and bare land areas in terms of grayscale and reflectance, it is desirable to assign a larger attention weight to water areas compared to other areas. This difference can be used to determine the spatial attention weights for different regions of the image.
[0051] Specifically, firstly, a clustering algorithm is used to segment the visible light remote sensing image and the near-infrared remote sensing image to obtain several visible light feature regions in the visible light remote sensing image and several near-infrared feature regions in the near-infrared remote sensing image.
[0052] As a preferred embodiment, a clustering algorithm is used to segment the visible light remote sensing image and the near-infrared remote sensing image into regions. The specific method is as follows: For the visible light remote sensing image and the near-infrared remote sensing image, the ISODATA algorithm is used to cluster the pixels in the image. Several clusters are obtained in the visible light remote sensing image and the near-infrared remote sensing image. The pixels contained in each cluster form a region in the image, which is denoted as a feature region. The feature region in the visible light remote sensing image is denoted as the visible light feature region, and the feature region in the near-infrared remote sensing image is denoted as the near-infrared feature region.
[0053] It should be noted that in general wetlands, water bodies and terrestrial vegetation are interspersed. In visible light remote sensing images, the characteristic areas corresponding to water bodies are darker in color and have a more uniform color texture compared to bare vegetation areas. At the same time, due to the continuity of water areas, their boundaries are usually smoother and more regular, while the boundaries of vegetation areas are usually more irregular.
[0054] Then, by utilizing the grayscale distribution, texture features, and contour morphology features in the visible light remote sensing image, the confidence level of the visible light feature region in the visible light remote sensing image is obtained.
[0055] As a preferred embodiment, the confidence level is obtained as follows: For any visible light feature region in a visible light remote sensing image, the gray-level co-occurrence matrix of the visible light feature region is obtained, and the entropy value of all elements in the gray-level co-occurrence matrix is used as the texture information entropy of the visible light feature region. Based on the difference between the average gray value of the visible light feature region and the average gray value of all visible light feature regions, the texture information entropy of the visible light feature region, and the fractal dimension of the visible light feature region, the confidence level of the visible light feature region in the visible light remote sensing image is obtained. The difference is positively correlated with the confidence level, and the texture information entropy and the fractal dimension are both negatively correlated with the confidence level. The confidence level of any near-infrared feature region in a near-infrared remote sensing image is obtained using the same method as that for the visible light feature region.
[0056] As an optional embodiment, the specific method for calculating the confidence level of the visible light feature region is as follows:
[0057] ;
[0058] In the formula, A a represents the confidence level of the a-th feature region in the visible light remote sensing image; Let be the average gray value of the a-th visible light feature region. H represents the maximum average gray value across all visible light characteristic regions. a D represents the texture information entropy of the a-th visible light feature region; a Let f be the fractal dimension of the a-th visible light feature region; norm() represents the linear normalization function.
[0059] It should be noted that, therefore This reflects the gray-level distribution characteristics of the feature region; the larger the value, the darker and more uniform the gray level of the corresponding region, and the more likely it is to be a water body; fractal dimension D aThis reflects the morphological complexity of the corresponding feature region. The larger the fractal dimension, the higher the morphological complexity. Since the outline of wetland water bodies usually does not exhibit excessive complexity, the smaller the fractal dimension, the lower the complexity of the edge outline of the corresponding feature region, i.e., the smoother the corresponding edge, and the greater the confidence that it is a water body region. In addition, since the image of the wetland obtained under near-infrared remote sensing scanning conditions, i.e., the near-infrared remote sensing image, can also reflect the possibility of each area in the wetland being a water body region to a certain extent, this embodiment of the invention also obtains the confidence of the near-infrared feature region in the near-infrared remote sensing image, so as to use it to divide the water body region and non-water body region in the near-infrared remote sensing image.
[0060] In addition, to facilitate the subsequent differentiation between water and non-water areas in remote sensing images, confidence level is used as a weight of attention in this embodiment of the invention. This allows sufficient attention to be given to the water body when identifying wetland water bodies, while less attention is paid to other non-water areas such as vegetation areas. Furthermore, for visible light remote sensing images, grayscale texture and fractal features can effectively describe the relevant characteristics of wetland water bodies in remote sensing images obtained through visible light.
[0061] In addition, the confidence level of the visible light feature region is used as the spatial attention weight of the corresponding visible light feature region.
[0062] It should be noted that for near-infrared remote sensing images, due to the large difference in reflectance between water bodies and vegetation (the reflectance of vegetation is usually between 40% and 60%, while the near-infrared reflectance of water bodies is usually less than 10%), and because the reflectance of different feature regions has the same trend as the spatial attention weight, the relative reflectance of any feature image in the near-infrared remote sensing image is obtained.
[0063] Finally, a reflectance threshold is preset, and the ground reflectance of any near-infrared feature region in the near-infrared remote sensing image is obtained. The ratio of the reflectance threshold to the ground reflectance of the near-infrared feature region in the near-infrared remote sensing image is normalized to obtain the relative reflectance of the near-infrared feature region.
[0064] As an optional embodiment, the specific method for calculating the relative reflectance of the feature region is as follows:
[0065] ;
[0066] In the formula, B b Rb is the relative reflectance of the b-th near-infrared feature region in the near-infrared remote sensing image; R0 is a preset reflectance threshold, Rb... bLet be the ground reflectance of the b-th near-infrared feature region in the near-infrared remote sensing image; norm() represents the linear normalization function.
[0067] It should be noted that in the embodiments of the present invention, the preset reflectivity threshold is 10% based on experience, which can be adjusted according to the actual situation. The embodiments of the present invention do not impose specific limitations.
[0068] It should be noted that because water absorbs most of the near-infrared light, the corresponding ground reflectivity will be lower. A higher value indicates a higher relative reflectance in the corresponding near-infrared feature region, thus making the near-infrared feature region more likely to be a water body. In addition, since the relative reflectance of different near-infrared feature regions in this area represents their probability of being a water body, using relative reflectance as the focus weight for the corresponding near-infrared feature regions in near-infrared remote sensing images can better reflect their characteristics in the near-infrared environment, thereby enabling better identification of wetland water bodies through near-infrared remote sensing images.
[0069] In addition, the relative reflectance of the near-infrared feature region is used as the spatial attention weight of the corresponding near-infrared feature region.
[0070] Thus, the confidence level of visible light feature regions in visible light remote sensing images and the relative reflectance of near-infrared feature regions in near-infrared remote sensing images are obtained through the above methods.
[0071] The weight analysis module 103 is used to match the visible light feature regions and near-infrared feature regions at the same time to form a matching group, and to perform weight analysis based on the confidence of the visible light feature regions and the relative reflectance of the near-infrared feature regions in the matching group to determine the fusion weight.
[0072] It should be noted that the spatial attention weights for different feature regions were determined using the above method for the multispectral images of this water body area. Since remote sensing images under different spectra capture different levels of information about the feature regions (water body, vegetation) of this area, fusing the attention weights from the multispectral images can make the differences between vegetation and water more apparent, improving the accuracy of subsequent target recognition based on a large model. Therefore, this step performs image fusion based on the spatial attention weights of the visible light and near-infrared light remote sensing images of this area.
[0073] It should be noted that, since the attention weights for the two images are acquired independently and simultaneously in the above steps, spatial fusion requires corresponding matching (lateral matching) of different feature regions in the two images at the same sampling time point. Therefore, for any two feature regions in the two images, if they represent the same region, the contours and positions of the corresponding feature regions will be more similar.
[0074] Specifically, firstly, by utilizing the spatial distance and overlap between the visible light feature regions in the visible light remote sensing image and the near-infrared feature regions in the near-infrared remote sensing image at the same time, the matching degree between the visible light feature regions and the near-infrared feature regions is obtained. Then, several matching groups are determined using the matching degree, and each matching group consists of a visible light feature region and a near-infrared feature region.
[0075] As a preferred embodiment, the specific method for obtaining the matching degree is as follows: For a visible light remote sensing image and a near-infrared remote sensing image at the same time, the centroids of the visible light feature regions in the visible light remote sensing image and the near-infrared feature regions in the near-infrared remote sensing image are obtained respectively, and the Euclidean distance between the corresponding coordinates of the centroids is obtained; the sets formed by the coordinates of all pixels contained in any visible light feature region and any near-infrared feature region are obtained respectively, denoted as the first region set and the second region set; the intersection of the first region set and the second region set is obtained, and the number of elements in the intersection is taken as the overlap degree of the visible light feature regions and near-infrared feature regions corresponding to the first region set and the second region set respectively; combining the Euclidean distance between the corresponding coordinates of the centroids of any visible light feature region and any near-infrared feature region and the overlap degree, the matching degree of the visible light feature region and the near-infrared feature region is calculated, wherein the Euclidean distance is negatively correlated with the matching degree, and the overlap degree is positively correlated with the matching degree.
[0076] When the matching degree is greater than or equal to the preset matching degree threshold, the visible light feature region and the near-infrared feature region corresponding to the matching degree form a matching group.
[0077] It should be noted that the preset matching threshold is preset to 0.8 based on experience in this embodiment of the invention, and can be adjusted according to the actual situation. It is not specifically limited in this embodiment of the invention.
[0078] As an optional embodiment, the specific method for calculating the matching degree is as follows:
[0079] ;
[0080] In the formula, Z a,b R represents the matching degree between the a-th visible light feature region in the visible light remote sensing image and the b-th near-infrared feature region in the near-infrared remote sensing image; a R represents the first set of regions representing the a-th visible light feature region in a visible light remote sensing image; b C represents the second set of regions in a near-infrared remote sensing image, specifically the b-th near-infrared feature region. a C represents the centroid coordinates of the a-th visible light feature region in a visible light remote sensing image; brepresents the centroid coordinates of the b-th near-infrared feature region in the near-infrared remote sensing image; L() represents the Euclidean distance function; norm[] represents the linear normalization function.
[0081] It should be noted that in the method for obtaining the matching degree, This reflects the degree of overlap between the contours of two feature regions (visible light feature region and near-infrared feature region) in different remote sensing images. The overlap is determined by superimposing the contours of the two feature regions through intersection. The more elements in the intersection, the greater the degree of overlap, and the more likely the two feature regions in different remote sensing images represent the same location. Furthermore, L(C a C b The value represents the degree of proximity between two feature regions. The smaller the value, the closer the two feature regions are in the remote sensing image, and the greater the matching degree.
[0082] It should be noted that after matching the visible light feature regions in the visible light remote sensing image and the near-infrared feature regions in the near-infrared remote sensing image at the same time using the matching degree, for any matching group at any time point, the spatial attention weights corresponding to the two feature regions (i.e., visible light feature regions and near-infrared feature regions) are different. When fusion of weights, it is necessary to weight them based on the contribution of the corresponding remote sensing image to the accuracy of water body identification. If the contribution of the remote sensing image to wetland water body identification is greater, then the corresponding spatial attention weight will account for a larger proportion during fusion, and vice versa.
[0083] It should be noted that since the contribution of a remote sensing image represents its ability to accurately distinguish between water and non-water bodies, a stronger distinguishing ability corresponds to a larger contribution. Therefore, during spatial weighted fusion, the focus weight corresponding to this image will be more biased. Thus, it is necessary to evaluate the distinguishing ability of different images. For two images, visible light remote sensing images represent the difference between water and non-water bodies through grayscale differences, while near-infrared remote sensing images represent this through reflectance differences mapped to grayscale images. That is, the grayscale features of both images can represent their distinguishing ability. Therefore, the greater the difference in focus weight between water and non-water body regions in an image, the more obvious the distinction between water and non-water bodies in that image, and the stronger its distinguishing ability.
[0084] The confidence level is used to divide the remote sensing image into water body regions and non-water body regions. Based on the difference in spatial attention weight levels among all water body regions and non-water body regions in the remote sensing image, the discrimination coefficient of the corresponding remote sensing image is calculated.
[0085] As a preferred embodiment, the method for obtaining the discrimination coefficient of the remote sensing image is as follows: For any remote sensing image, feature regions in the remote sensing image with a confidence level greater than or equal to a preset confidence threshold are designated as water bodies, and feature regions outside the water bodies are designated as non-water bodies. The water bodies and non-water bodies in the remote sensing image are obtained, and the average spatial attention weight of all water bodies in the remote sensing image is obtained as the water body attention weight of the remote sensing image. The average spatial attention weight of all non-water bodies in the remote sensing image is obtained as the non-water body attention weight of the remote sensing image. The discrimination coefficient of the remote sensing image is obtained based on the difference between the water body attention weight and the non-water body attention weight.
[0086] In a specific embodiment of the present invention, the discrimination coefficients of the remote sensing images are used to obtain the discrimination coefficients of the visible light remote sensing images and the near-infrared remote sensing images, respectively.
[0087] As an optional embodiment, for any remote sensing image, the specific method for calculating the discrimination coefficient of the remote sensing image is as follows:
[0088] ;
[0089] In the formula, Y is the discrimination coefficient of the remote sensing image; Water body attention weights in remote sensing images The non-water body focus weights in the remote sensing image; u represents the preset first parameter; || represents the absolute value function.
[0090] It should be noted that, The larger the value, the greater the difference between the two images, the more obvious the water area discrimination ability, and the stronger the discrimination ability. The discrimination ability coefficient of the two images at this time point is calculated using the above method. Therefore, the image with stronger discrimination ability contributes more to spatial fusion, and the attention weight of this image plays a more significant role in the fusion weight, and vice versa.
[0091] It should be noted that, in this embodiment of the invention, the confidence threshold is preset to 0.8 based on experience. In order to avoid the first parameter being preset to 0.01 to avoid the denominator being 0, the confidence threshold and the first parameter can be adjusted according to the actual situation. This embodiment of the invention does not make specific requirements.
[0092] For any matching group, the comprehensive spatial attention weight of the matching group is calculated based on the spatial attention weight of the visible light feature region in the matching group, the discrimination coefficient of the visible light remote sensing image to which the visible light feature region belongs, the spatial attention weight of the near-infrared feature region, and the discrimination coefficient of the near-infrared remote sensing image to which the near-infrared feature region belongs.
[0093] As an optional embodiment, for any matching group, the specific calculation method for the comprehensive spatial attention weight of the matching group is as follows:
[0094] ;
[0095] In the formula, γk represents the spatial attention weight of the matching group; Yk represents the spatial attention weight of the visible light feature region in the matching group; γh represents the discrimination coefficient of the visible light remote sensing image to which the visible light feature region in the matching group belongs; Yh represents the spatial attention weight of the near-infrared feature region in the matching group; and Yh represents the discrimination coefficient of the near-infrared remote sensing image to which the near-infrared feature region in the matching group belongs.
[0096] It should be noted that, and These represent the relative discrimination coefficients of the visible light remote sensing image and the near-infrared remote sensing image compared to the two remote sensing images. They reflect the contribution of the two remote sensing images in accurately distinguishing between water and non-water areas. The comprehensive attention weight is obtained by weighting the attention weights under the corresponding images with these coefficients and then annotating the corresponding areas of the visible light remote sensing image.
[0097] It should be noted that, since the feature significance of different feature regions varies under different sampling environments (for example, the feature distinction of water bodies relative to vegetation is more obvious when the sunlight is strong than when the weather is cloudy or rainy), in order to prevent the spatial attention weight of the environment to this area from affecting the error at a single sampling time, it is necessary to further analyze the matching groups at different times.
[0098] It should be noted that, under normal circumstances, the changes in water bodies and vegetation land in wetlands are relatively stable and do not change drastically within the sampling period. Therefore, it is desirable that the spatial attention weight of the feature region after fusing visible light remote sensing images and near-infrared remote sensing images can reflect the stable characteristics of this feature region. Therefore, this embodiment of the invention further utilizes remote sensing images collected at multiple times for analysis.
[0099] Obtain the comprehensive spatial attention weights of the same matching group at different times, and form a corresponding sequence in chronological order, denoted as the attention weight sequence. Calculate the weight stability of each matching group by using the differences in the comprehensive spatial attention weights of the matching groups in the attention weight sequence. For any attention weight sequence, obtain the comprehensive spatial attention weight of the matching group with the highest weight stability, and use it as the fusion weight.
[0100] As a preferred embodiment, the method for obtaining the weight stability is as follows:
[0101] First, for any two time points, the first set of visible light feature regions and the second set of near-infrared feature regions contained in all matching groups are further combined into a set of the first and second region sets contained in any matching group as the total set of the matching group. The intersection-union ratio (IUR) of the total sets of any two matching groups at different time points is obtained. When the IUR is the largest, the two matching groups are the same matching group at different time points. The sequence of comprehensive spatial attention weights corresponding to the same matching group at all consecutive time points is obtained as the attention weight sequence.
[0102] Then, obtain the maximum comprehensive spatial attention weight and the mean of all comprehensive spatial attention weights in the attention weight sequence, obtain the difference between the comprehensive spatial attention weight of any matching group in the attention weight sequence and the maximum comprehensive spatial attention weight and the mean, and calculate the weight stability of the matching group.
[0103] As an optional embodiment, for any sequence of attention weights, the specific method for calculating the weight stability of the matching group corresponding to the element in the sequence of attention weights is as follows:
[0104] ;
[0105] In the formula, τ t This indicates that we are concerned with the stability of the weights of the matching group corresponding to the t-th element in the weight sequence. This represents the comprehensive spatial focus weight of the matching group corresponding to the t-th element in the focus weight sequence; E represents the mean of all spatially relevant weights in the weight sequence. γ represents the maximum comprehensive space of the focus weights in the focus weight sequence; || represents the absolute value function.
[0106] It should be noted that, This represents the difference between the weight of the t-th element in the weight sequence and the mean of all weights in the overall space. The smaller the value, the closer the weight is to the mean weight, and the higher its stability. To address the difference between the integrated spatial attention weight corresponding to the t-th element in the attention weight sequence and the maximum integrated spatial attention weight, a larger value corresponds to a greater distance from the extreme value and higher stability. Since the maximum integrated spatial attention weight in the attention weight sequence represents its most unstable integrated spatial attention weight, in order to effectively fuse visible light remote sensing images and near-infrared remote sensing images to accurately identify wetland water bodies, this embodiment of the invention selects a final integrated spatial attention weight that is far removed from the maximum integrated spatial attention weight in the attention weight sequence. Furthermore, the mean of all integrated spatial attention weights in the attention weight sequence reflects the overall level of this attention weight sequence; therefore, this embodiment of the invention selects a final integrated spatial attention weight that is as close as possible to the mean of all integrated spatial attention weights.
[0107] Thus, the weight stability of the matching group is obtained through the above method.
[0108] The remote sensing recognition module 104 is used to perform weighted image fusion on the feature regions in the matching group based on the stability of the weights, obtain the fused region, use the fused region as the input of the neural network, and output the wetland water body recognition result.
[0109] Specifically, firstly, for any sequence of attention weights, the matching group corresponding to the elements contained in the sequence of attention weights with the highest weight stability is obtained as the target matching group. The comprehensive spatial attention weight of the target matching group is used as the fusion weight in the alpha fusion algorithm, thereby fusing the visible light feature region and the near-infrared feature region in the target matching group to obtain the fused region.
[0110] Then, all fused regions are acquired and used as input to a trained convolutional neural network (CNN). The CNN is then used to identify wetland water bodies, and the CNN outputs the identification results.
[0111] As an optional embodiment, the training process of the convolutional neural network is as follows: 1) Acquire remote sensing images from several different regions and under different acquisition environments, and obtain several fusion regions through the feature region module and weight analysis module, and assign corresponding artificial labels to each fusion region, including wetland water bodies and non-wetland water bodies; 2) Take any fusion region with an artificial label as a sample to obtain a dataset formed by all samples, and divide the dataset into a training set, a validation set and a test set according to a 6:2:2 relationship; 3) Use the dataset as the input of the convolutional neural network, and use the artificial labels of the samples in the dataset as the output of the convolutional neural network to train the convolutional neural network, and select the cross-entropy loss function as the loss function of the convolutional neural network.
[0112] This concludes the embodiment.
[0113] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A wetland water body remote sensing identification system based on large model analysis, characterized in that, The system comprises the following modules: a remote sensing acquisition module, configured to acquire remote sensing images and ground reflectivity, the remote sensing images comprising visible light remote sensing images and near-infrared remote sensing images; a feature region module, configured to divide the remote sensing image region to obtain a plurality of feature regions, the feature regions being divided into visible light feature regions and near-infrared feature regions, the confidence of the feature region being acquired by using the gray scale distribution and texture features and contour shape features of the feature region, and the relative reflectivity of the near-infrared feature region being acquired by using the ground reflectivity in the near-infrared feature region; a weight analysis module, configured to match the visible light feature regions and the near-infrared feature regions at the same time to form a matching group, and perform weight analysis based on the confidence of the visible light feature region and the relative reflectivity of the near-infrared feature region in the matching group to determine the fusion weight; a remote sensing recognition module, configured to perform image fusion on the feature regions in the matching group by using the fusion weight, take the fusion result as the input of a neural network, and output a wetland water body recognition result; the method for acquiring the relative reflectivity comprises the following steps: presetting a reflectivity threshold, acquiring the ground reflectivity of any near-infrared feature region in the near-infrared remote sensing image, performing normalization processing on the ratio of the reflectivity threshold to the ground reflectivity of the near-infrared feature region in the near-infrared remote sensing image, and obtaining the relative reflectivity of the near-infrared feature region; the method for acquiring the fusion weight comprises the following steps: taking the confidence of the visible light feature region as the spatial attention weight of the corresponding visible light feature region; taking the relative reflectivity of the near-infrared feature region as the spatial attention weight of the corresponding near-infrared feature region; dividing the water body region and the non-water body region in the remote sensing image by using the size of the confidence, calculating the discrimination coefficient of the corresponding remote sensing image according to the spatial attention weight level difference of all the water body regions and the non-water body regions in the remote sensing image; for any matching group, calculating the comprehensive spatial attention weight of the matching group according to the spatial attention weight of the visible light feature region, the discrimination coefficient of the visible light remote sensing image to which the visible light feature region belongs, the spatial attention weight of the near-infrared feature region, and the discrimination coefficient of the near-infrared remote sensing image to which the near-infrared feature region belongs; 2. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, acquiring the comprehensive spatial attention weight of the same matching group at different time points, forming a corresponding sequence in time sequence, denoted as an attention weight sequence, calculating the weight stability degree of each matching group by using the comprehensive spatial attention weight difference of the matching groups in the attention weight sequence; acquiring the maximum comprehensive spatial attention weight in the attention weight sequence and the mean value of all the comprehensive spatial attention weights, acquiring the difference between the comprehensive spatial attention weight of any matching group in the attention weight sequence and the maximum comprehensive spatial attention weight and the mean value, and calculating the weight stability degree of the matching group; for any attention weight sequence, the comprehensive spatial attention weight of the matching group corresponding to the maximum weight stability degree is taken as the fusion weight. the method for dividing the remote sensing image region to obtain a plurality of feature regions comprises the following steps: For the visible light remote sensing image and the near-infrared remote sensing image, the ISODATA algorithm is used to cluster the pixel points in the images, and a plurality of clustering clusters are obtained in the visible light remote sensing image and the near-infrared remote sensing image, each clustering cluster contains pixel points forming a region in the image, which is recorded as a feature region, the feature region in the visible light remote sensing image is recorded as a visible light feature region, and the feature region in the near-infrared remote sensing image is recorded as a near-infrared feature region.
3. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, The confidence of the feature region is obtained by using the gray distribution and the texture feature and the contour shape feature of the feature region, and the specific method comprises: For any visible light feature region in the visible light remote sensing image, the gray level co-occurrence matrix of the visible light feature region is obtained, and the entropy value of all elements in the gray level co-occurrence matrix is taken as the texture information entropy of the visible light feature region, the confidence of the visible light feature region in the visible light remote sensing image is obtained according to the difference between the average gray value of the visible light feature region and the average gray value of all visible light feature regions, the texture information entropy of the visible light feature region and the fractal dimension of the visible light feature region, the difference is positively correlated with the confidence, and the texture information entropy and the fractal dimension are negatively correlated with the confidence; the confidence of any near-infrared feature region in the near-infrared remote sensing image is obtained, and the confidence of the near-infrared feature region is obtained in the same way as the confidence of the visible light feature region.
4. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, The visible light feature region and the near-infrared feature region at the same time are matched to form a matching group, and the specific method comprises: The matching degree of the visible light feature region and the near-infrared feature region is obtained by using the distance and the coincidence of the visible light feature region in the visible light remote sensing image and the near-infrared feature region in the near-infrared remote sensing image at the same time, and the matching degree is used to determine a plurality of matching groups, and each matching group is composed of a visible light feature region and a near-infrared feature region.
5. The wetland water body remote sensing identification system based on large model analysis according to claim 4, characterized in that, The specific method for obtaining the matching degree is: For the visible light remote sensing image and the near-infrared remote sensing image at the same time, the centroids of the visible light feature region in the visible light remote sensing image and the near-infrared feature region in the near-infrared remote sensing image are obtained, and the Euclidean distance between the coordinates corresponding to the centroids is obtained; the set formed by the coordinates of all pixel points in any visible light feature region and any near-infrared feature region is obtained, which is recorded as a first region set and a second region set, the intersection of the first region set and the second region set is obtained, and the number of elements in the intersection is taken as the coincidence degree of the visible light feature region and the near-infrared feature region corresponding to the first region set and the second region set respectively, the Euclidean distance between the coordinates corresponding to the centroids of any visible light feature region and any near-infrared feature region and the coincidence degree are combined to calculate the matching degree of the visible light feature region and the near-infrared feature region, the Euclidean distance is negatively correlated with the matching degree, and the coincidence degree is positively correlated with the matching degree.
6. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, The specific method for obtaining the discrimination coefficient of the remote sensing image is: For any remote sensing image, the feature region with a confidence greater than or equal to a preset confidence threshold in the remote sensing image is taken as a water body region, and the feature region other than the water body region is taken as a non-water body region, the water body region and the non-water body region in the remote sensing image are obtained, the average spatial attention weight of all water body regions in the remote sensing image is taken as a water body attention weight of the remote sensing image, the average spatial attention weight of all non-water body regions in the remote sensing image is taken as a non-water body attention weight of the remote sensing image, and a discrimination coefficient of the remote sensing image is obtained according to the difference between the water body attention weight and the non-water body attention weight.
7. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, The specific method for obtaining the weight stability degree of the matching group is as follows: For the first region set of the visible light feature region and the second region set of the near-infrared feature region contained in all matching groups at any two time points, the set further formed by the first region set and the second region set contained in any matching group is taken as a total set of the matching group, the intersection-over-union of the total sets of any two matching groups at different time points is obtained, when the intersection-over-union is maximum, the two matching groups are the same matching group at different time points, and the sequence formed by the comprehensive spatial attention weights corresponding to the same matching group at all continuous time points is taken as an attention weight sequence; The maximum comprehensive spatial attention weight in the attention weight sequence and the mean value of all comprehensive spatial attention weights are obtained, the difference between the comprehensive spatial attention weight of any matching group in the attention weight sequence and the maximum comprehensive spatial attention weight and the mean value is obtained, and the weight stability degree of the matching group is calculated.
8. The wetland water body remote sensing identification system based on large model analysis according to claim 1, characterized in that, The specific method for using the fusion weight to perform image fusion on the feature region in the matching group, taking the fusion result as the input of the neural network, and outputting the wetland water body recognition result is as follows: For any attention weight sequence, the matching group corresponding to the element in the attention weight sequence when the weight stability degree is maximum is taken as a target matching group, the comprehensive spatial attention weight of the target matching group is used as the fusion weight in the alpha fusion algorithm, so that the visible light feature region and the near-infrared feature region in the target matching group are fused to obtain a fusion region; all fusion regions are obtained and taken as the input of the trained convolutional neural network, the convolutional neural network is used for wetland water body recognition, and the recognition result is output by the convolutional neural network.
Citation Information
Patent Citations
Real-time geographic information analysis method and system fusing large model and remote sensing technology
CN119206519A
Coastal wetland damage dynamic identification method and system based on remote sensing technology
CN120259886A