Remote sensing image self-learning segmentation method

By simulating a multi-feature deformation lightweight neural network that simulates non-classical receptive fields and visual attention mechanisms, combined with superpixel division and clustering, the adaptability problem of remote sensing image segmentation technology in complex environments is solved, and high-precision automated segmentation and classification are achieved.

CN120451525APending Publication Date: 2025-08-08NORTHWEST ENGINEERING CORPORATION LIMITED
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510291233.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing high-resolution remote sensing image segmentation technology is difficult to achieve adaptability, self-learning ability and self-checking mechanism in complex and changeable ground environments, especially in scenarios where samples are scarce and interference rich in information, resulting in insufficient accuracy and completeness of segmentation results, which limits its wide application.

Method used

Edge recognition and superpixel division methods that simulate non-classical receptive fields are used, and a multi-feature deformation lightweight neural network based on visual attention mechanism is combined with clustering and a multi-feature deformation lightweight neural network to generate a superpixel area recognition model, perform remote sensing image segmentation, and error repair is performed through the eyeball micro-vibration perception mechanism.

Benefits of technology

It realizes the automatic and intelligent acquisition of high-precision ground scene segmentation results while reducing manual intervention, improves classification accuracy and processing efficiency, and enhances the adaptability and reliability of remote sensing images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451525A_ABST
    Figure CN120451525A_ABST
Patent Text Reader

Abstract

The invention provides a remote sensing image self-learning segmentation method, belongs to the technical field of image processing, and aims to realize automatic, accurate, sufficient and reliable high-resolution remote sensing image segmentation. Comprising the following steps: responding to an input remote sensing image, and performing spectrum correction operation on the remote sensing image to obtain a first color image; carrying out super-pixel division on the first color image by utilizing an edge recognition and super-pixel division method for simulating a non-classical receptive field to obtain super-pixel plaques under different scales; performing clustering operation on the superpixel plaques to obtain a clustering result; training a multi-feature deformation lightweight neural network based on a visual attention mechanism by using a clustering result, and generating a super-pixel region recognition model; and performing remote sensing image segmentation by using the super-pixel region recognition model. According to the method, the self-learning ability of human eye visual perception can be simulated, then image segmentation is carried out through self-adaptive analysis, self-learning identification and self-checking correction, and a segmentation result is rapidly and accurately obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular to a remote sensing image self-learning segmentation method. Background Art

[0002] In recent years, with the continuous improvement of integrated air-space-ground observation systems, high-resolution remote sensing imagery has become a vital component of geographic information science and Earth observation technology, providing indispensable data support and real-world evidence for a wide range of fields, including ecological and environmental monitoring, urban planning, disaster assessment, and agricultural management. Image segmentation technology, a core component of remote sensing image analysis, uses algorithms to automatically identify and distinguish different types of land features in an image, effectively reducing the difficulty of remote sensing image interpretation, optimizing data processing, and improving information extraction efficiency.

[0003] Currently, high-resolution remote sensing image segmentation technology, with sufficient sample support, can accurately segment typical land features, significantly enhancing the application value of remote sensing data. However, faced with complex and changing ground environments, especially in scenarios with scarce samples and abundant interference information, traditional segmentation methods are often limited by the limitations of supervised learning, making it difficult to effectively extract the semantic features of land features, resulting in a significant reduction in the accuracy and completeness of segmentation results.

[0004] Although academia and industry have invested significant resources in exploring numerous improvement strategies, such as deep learning, feature fusion, and multi-scale analysis, to enhance segmentation performance, significant shortcomings remain in adaptability, self-learning capabilities, and self-verification mechanisms. These methods often rely excessively on specific types of training samples and pre-set models, failing to fully account for the inherent heterogeneity and complexity of remote sensing imagery. Consequently, they are unable to achieve comprehensive, accurate, and robust image segmentation in an environment without sample guidance, limiting their practicality and reliability in a wide range of application scenarios. Summary of the Invention

[0005] In order to overcome the problems existing in the related art, the present invention provides a remote sensing image self-learning segmentation method.

[0006] According to a first aspect of an embodiment of the present invention, a remote sensing image self-learning segmentation method is provided, comprising: In response to an input remote sensing image, performing a spectral correction operation on the remote sensing image to obtain a first color image; Performing superpixel division on the first color image using an edge recognition and superpixel division method simulating a non-classical receptive field to obtain superpixel patches at different scales; Performing a clustering operation on the superpixel patches to obtain a clustering result; Using the clustering results to train a multi-feature deformable lightweight neural network based on a visual attention mechanism to generate a superpixel region recognition model; The superpixel region recognition model is used to perform remote sensing image segmentation.

[0007] In some example embodiments of the present invention, based on the aforementioned solution, the step of responding to the input remote sensing image and performing a spectral correction operation on the remote sensing image to obtain the first color image includes: According to the dimension of the remote sensing image, based on the principal component analysis method, the remote sensing image is subjected to dimensionality reduction processing to obtain a three-channel color image; The three-channel color image is transformed using a method specified by the International Illumination Commission to generate a first color image.

[0008] In some example embodiments of the present invention, based on the aforementioned solution, the first color image is subjected to superpixel division using an edge recognition and superpixel division method simulating a non-classical receptive field, and superpixel patches at different scales are obtained, including: Measuring the grayscale gradient of the first color image using a Laplacian operator to obtain a measurement result; The non-classical receptive field of the optic nerve is simulated to identify edge features of different scales, and Laplacian operators of different scales are used to obtain superpixel patches composed of different scales; among them, the independent area wrapped by the edge is defined as a superpixel.

[0009] In some exemplary embodiments of the present invention, based on the aforementioned solution, clustering operation is performed on the superpixel patches to obtain clustering results, including: generating a representation map of independent features based on the superpixel patches at different scales; Analyze the spectral characteristics of all representation images to determine the categories of independent features; Constructing a mapping between remote sensing image data and the categories of the independent features to generate homologous reliable samples; The homologous reliable samples are output as the clustering results.

[0010] In some example embodiments of the present invention, based on the aforementioned solution, generating a representation map of an independent feature based on the superpixel patches at different scales includes: Performing scale normalization on super-pixel patches at different scales to obtain several normalized super-pixel patches; Eroding a number of normalized super-pixel patches to obtain a number of eroded super-pixel patches; Overlaying the eroded superpixel patches to obtain an overlay result; Screening out small-area noise in the superposition result to obtain a screening result; Different superpixel patches in the screening results are numbered based on preset rules to form a representation map of independent features.

[0011] In some exemplary embodiments of the present invention, based on the aforementioned scheme, analyzing the spectral characteristics of all representation images to determine the categories of independent features includes: Credible scale determination step: determining a credible scale of the current representation image based on the scale of each superpixel patch in the current representation image, wherein, when there are superpixel patches of multiple scales in the representation image, the median of the scales is used as the credible scale; Spectral band determination step: determining the spectral band of the current representation image according to the credible scale; Repeating step: repeating the trust scale determination step and the spectral band determination step to obtain spectral bands of all feature maps; Adaptive clustering step: adaptively cluster the spectral bands of all feature maps to obtain clustering results; Category determination step: Based on the clustering results, the category of the independent features is determined.

[0012] In some exemplary embodiments of the present invention, based on the aforementioned solution, the multi-feature deformable lightweight neural network based on the visual attention mechanism includes: A multi-scale differential encoding part, the multi-scale differential encoding part includes an encoding input layer, a first transposed convolution amplification layer to a fourth transposed convolution amplification layer, a first feature mining unit to a fifth feature mining unit, a first self-attention weight assignment unit to a fifth self-attention weight assignment unit, a first average pooling layer, a second average pooling layer and an encoding output layer; wherein, the output of the encoding input layer is simultaneously used as the input of the first transposed convolution amplification layer, the input of the first average pooling layer, the input of the first feature mining unit and the input of the first self-attention weight assignment unit, the output of the first transposed convolution amplification layer is simultaneously used as the input of the second transposed convolution amplification layer, the input of the second feature mining unit and the input of the second self-attention weight assignment unit, the output of the second transposed convolution amplification layer is simultaneously used as the input of the third feature mining unit and the input of the third self-attention weight assignment unit, the output of the first average pooling layer is simultaneously used as the input of the fourth feature mining unit and the input of the fourth self-attention weight assignment unit, the output of the second average pooling layer is simultaneously used as the input of the fifth feature mining unit and the input of the fifth self-attention weight assignment unit, the output of the third feature mining unit and the third self-attention weight assignment unit The output of the redistribution unit performs a first channel multiplication operation, and the results of the first channel multiplication perform a first local maximum downsampling operation and a first pooling reduction operation respectively. The output of the second feature mining unit and the output of the first self-attention weight assignment unit perform a second channel multiplication operation, the result of the second channel multiplication and the result of the first local maximum downsampling perform a first feature subtraction operation, and perform a first feature channel superposition operation with the result of the first pooling reduction, the result of the first feature subtraction performs a second local maximum downsampling operation, and the first feature channel superposition result performs a second pooling reduction operation, the output of the fifth feature mining unit and the output of the fifth self-attention weight assignment unit perform a third channel multiplication operation, the result of the third channel multiplication serves as the input of the third transposed convolutional amplification layer, and performs a first nearest neighbor amplification operation; the output of the fourth feature mining unit and the output of the fourth self-attention weight assignment unit perform a fourth channel multiplication operation, the result of the fourth channel multiplication and the result of the first nearest neighbor amplification perform a second feature subtraction operation, and perform a second feature channel superposition operation with the output of the third transposed convolutional amplification layer, the result of the second feature subtraction performs a second nearest neighbor amplification operation, and the result of the second feature channel superposition serves as the input of the fourth transposed convolutional amplification layer;The output of the first feature mining unit and the output of the first self-attention weight allocation unit are multiplied by a fifth channel. The result of the fifth channel multiplication is subtracted from the result of the second local maximum downsampling by a third feature, and is subtracted from the result of the second nearest neighbor amplification by a fourth feature. The result of the third feature subtraction, the result of the second pooling reduction of the result of the fourth feature subtraction, and the output of the fourth transposed convolution amplification layer are superimposed on the third feature channel. The result of the third feature channel superposition is used as the input of the encoding output layer, and the output of the encoding output layer is the output of the multi-scale differential encoding part. The multi-scale receptive field decoding part includes a decoding input layer, a first 1×1 convolutional layer, a first 3×3 convolutional layer to a sixth 3×3 convolutional layer, a channel superposition layer of the first feature, a first ReLU activation layer, a second ReLU activation layer, and a first Softmax probability layer; wherein the output of the encoding output layer is used as the input of the decoding input layer, and the output of the decoding input layer is also used as the input of the first 1×1 convolutional layer and the first 3×3 convolutional layer. The first 3×3 convolutional layer to the fifth 3×3 convolutional layer are arranged in sequence, and the output of the first 3×3 convolutional layer to the output of the fifth 3×3 convolutional layer and the output of the first 1×1 convolutional layer are superimposed in the channel superposition layer of the first feature, and then a self-attention weight allocation operation is performed, and the operation result is multiplied by the feature channel. The multiplication result is used as the input of the first ReLU activation layer. The first ReLU activation layer, the sixth 3×3 convolutional layer, the second ReLU activation layer, and the first Softmax probability layer are arranged in sequence, and the output of the first Softmax probability layer is the output of the multi-scale receptive field decoding part.

[0013] In some exemplary embodiments of the present invention, based on the aforementioned solution, the first to fifth feature mining units are configured to have the same structure; The feature mining unit includes: the seventh 3×3 convolution layer to the tenth 3×3 convolution layer, the third ReLU activation layer to the sixth ReLU activation layer, the channel superposition layer of the second feature and the channel superposition layer of the third feature; Among them, the input of the seventh 3×3 convolutional layer is used as the input of the feature mining unit, and the seventh 3×3 convolutional layer, the third ReLU activation layer, the eighth 3×3 convolutional layer, the fourth ReLU activation layer and the channel superposition layer of the second feature are arranged in sequence; the ninth 3×3 convolutional layer, the fifth ReLU activation layer, the tenth 3×3 convolutional layer, the sixth ReLU activation layer and the channel superposition layer of the third feature are arranged in sequence; the output of the channel superposition layer of the second feature is also used as the input of the channel superposition layer of the third feature, and the output of the channel superposition layer of the third feature is used as the output of the feature mining unit.

[0014] In some exemplary embodiments of the present invention, based on the aforementioned solution, the first to fifth self-attention weight assignment units are configured to have the same structure; Each attention weight allocation unit includes an eleventh 3×3 convolution layer, a twelfth 3×3 convolution layer, a seventh ReLU activation layer, a second 1×1 convolution layer, an eighth ReLU activation layer and a second Softmax probability layer, which are arranged in sequence.

[0015] In some example embodiments of the present invention, based on the aforementioned solution, after performing remote sensing image segmentation using the superpixel region recognition model, the remote sensing image self-learning segmentation method further includes: Based on the eye tremor perception mechanism, errors in the segmentation results and superpixel patches at different scales are repaired.

[0016] In some example embodiments of the present invention, based on the aforementioned solution, performing error repair on the clustering results and the superpixel patches at different scales based on the eye tremor perception mechanism includes: Map the segmentation results to each superpixel patch; The category of the independent feature that appears most frequently on each superpixel patch is regarded as the category of the main independent feature of the current superpixel patch; The segmentation results are optimized according to the categories of the main independent objects of all superpixel patches.

[0017] According to a second aspect of an embodiment of the present invention, a remote sensing image self-learning segmentation device is provided, the remote sensing image self-learning segmentation device comprising: a spectral correction module, configured to respond to an input remote sensing image and perform a spectral correction operation on the remote sensing image to obtain a first color image; a superpixel division module, configured to perform superpixel division on the first color image by using an edge recognition and superpixel division method simulating a non-classical receptive field to obtain superpixel patches at different scales; A clustering operation module, configured to perform a clustering operation on the superpixel patches to obtain a clustering result; A model building module is used to train a multi-feature deformable lightweight neural network based on a visual attention mechanism using the clustering results to generate a superpixel region recognition model; The image segmentation module is used to perform remote sensing image segmentation using the superpixel region recognition model.

[0018] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising: a processor; and a memory, wherein the memory stores computer-readable instructions, and the computer-readable instructions implement the method in the first aspect when executed by the processor.

[0019] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method in the first aspect when executed by a processor.

[0020] The technical solutions provided by the embodiments of the present invention may have the following beneficial effects: The present invention can avoid complicated processes such as sample data collection and sample reliability discussion, simulate the self-learning ability of human visual perception, and automatically and intelligently obtain high-precision ground scene segmentation results, thereby effectively improving classification accuracy and processing efficiency while reducing human intervention.

[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0023] Figure 1 A schematic diagram showing a system architecture of an exemplary application environment in which a remote sensing image self-learning segmentation method and apparatus according to an embodiment of the present invention can be applied; Figure 2 Schematically illustrating a flow chart of a remote sensing image self-learning segmentation method according to some embodiments of the present invention; Figure 3 Schematically illustrates the overall structure of a multi-feature deformation lightweight neural network based on visual attention according to some embodiments of the present invention; Figure 4 Schematically shows Figure 3 Schematic diagram of the structure of the multi-scale differential coding part; Figure 5 Schematically shows Figure 3 Schematic diagram of the structure of the multiple receptive field decoding part; Figure 6 Schematically shows Figure 4 Schematic diagram of the structure of the feature mining unit; Figure 7 Schematically shows Figure 4 Schematic diagram of the structure of the self-attention weight allocation unit; Figure 8 Schematically illustrates a schematic diagram of a remote sensing image self-learning segmentation device according to some embodiments of the present invention; Figure 9 Schematically illustrates a structural diagram of a computer system of an electronic device according to some embodiments of the present invention; Figure 10A schematic diagram of a computer-readable storage medium according to some embodiments of the present invention is schematically shown. DETAILED DESCRIPTION

[0024] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, like numbers in different figures represent like or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present invention, as detailed in the appended claims.

[0025] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a," "the," and "the" used in this invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.

[0026] It should be understood that although the terms "first," "second," "third," etc. may be used in the present invention to describe various information, such information should not be limited to these terms. These terms are merely used to distinguish information of the same type from one another. For example, first information may also be referred to as second information, and similarly, second information may also be referred to as first information, without departing from the scope of the present invention. Depending on the context, the term "if" as used herein may be interpreted as "when," "when," or "in response to determining."

[0027] Figure 1 A schematic diagram shows the system architecture of an exemplary application environment in which a remote sensing image self-learning segmentation method and apparatus according to an embodiment of the present invention can be applied.

[0028] like Figure 1 As shown, the system architecture 100 may include one or more terminal devices such as a desktop computer 101, a portable computer 102, a smart phone 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the terminal device and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc. The terminal device may be any electronic device with a data processing function, which has a display screen for displaying target structure feature points or three-dimensional spatial information of each independent pipeline to the user, including but not limited to the above-mentioned desktop computers, portable computers, smart phones, etc. It should be understood that Figure 1The number of terminal devices, networks, and servers in the embodiment is merely illustrative. Any number of terminal devices, networks, and servers may be provided as needed. For example, server 105 may be a sub-server cluster consisting of multiple sub-servers.

[0029] The remote sensing image self-learning segmentation method provided in the embodiments of the present invention can generally be executed by a terminal device, and accordingly, the remote sensing image self-learning segmentation device is generally provided in the terminal device. However, those skilled in the art will readily appreciate that the remote sensing image self-learning segmentation method provided in the embodiments of the present invention can also be executed by the server 105, and accordingly, the remote sensing image self-learning segmentation device can also be provided in the server 105, and this exemplary embodiment does not specifically limit this.

[0030] In addition, it should be understood that the remote sensing image self-learning segmentation method of the embodiment of the present invention can be configured as a software module. In some implementation scenarios, the remote sensing image self-learning segmentation scheme of the present invention can be deployed separately to achieve self-learning segmentation of various remote sensing images. In other implementation scenarios, the remote sensing image self-learning segmentation scheme of the present invention can be deployed in other software as a functional module of the software, such as deployed in remote sensing image self-learning segmentation analysis software. The present invention does not place any particular restrictions on the application of the remote sensing image self-learning segmentation method.

[0031] Next, the embodiments of the present invention are described in detail.

[0032] like Figure 2 As shown, Figure 2 1 is a flowchart of a remote sensing image self-learning segmentation method according to an exemplary embodiment of the present invention, comprising the following steps: S210: Responding to an input remote sensing image, performing a spectral correction operation on the remote sensing image to obtain a first color image; S220: performing superpixel division on the first color image using an edge recognition and superpixel division method simulating a non-classical receptive field to obtain superpixel patches at different scales; S230: performing a clustering operation on the superpixel patches to obtain a clustering result; S240: Using the clustering results to train a multi-feature deformable lightweight neural network based on a visual attention mechanism to generate a superpixel region recognition model; S250: Perform remote sensing image segmentation using the superpixel region recognition model.

[0033] In an embodiment of the present invention, the spectral correction operation can eliminate or reduce errors caused by factors such as atmosphere and lighting during the data acquisition process of the remote sensing sensor, ensuring that the spectral information of the remote sensing image is more real and reliable, so that the obtained first color image has more accurate color information and higher visual quality, providing a good data basis for subsequent processing; the edge recognition method that simulates the non-classical receptive field can more effectively identify the edges and details in the image. Combined with superpixel division, a set of pixels with similar features, namely superpixel patches, can be generated. This method can reduce the amount of data for image processing while maintaining the structural information of the image. Superpixel patches of different scales help to adapt to the features of land objects of different sizes; the clustering operation can group superpixel patches with similar features into a group, thereby classifying the land objects in the remote sensing image according to feature similarity. The generated clustering results provide meaningful training labels for subsequent neural network training, which helps to improve the classification accuracy of the model; the neural network based on the visual attention mechanism can simulate the human visual system, give priority to important areas in the image, and improve processing efficiency and accuracy. The design of a multi-feature deformable lightweight neural network can extract richer feature information while maintaining the network's lightweight nature, facilitating rapid training and deployment. The resulting superpixel region recognition model can more accurately identify and classify superpixel regions in remote sensing imagery. Using the trained superpixel region recognition model for remote sensing image segmentation enables precise identification and differentiation of different ground features in remote sensing imagery, providing a foundation for further applications of remote sensing imagery (such as resource surveys, environmental monitoring, and urban planning). In summary, this invention avoids the complex processes of sample data collection and sample reliability research, simulating the self-learning capabilities of human visual perception to automatically and intelligently obtain high-precision ground scene segmentation results.

[0034] In S210 , in response to an input remote sensing image, a spectral correction operation is performed on the remote sensing image to obtain a first color image.

[0035] Here, once the remote sensing image is input, the process begins. Spectral correction is performed to correct spectral distortion in remote sensing images, making the image data more closely resemble the actual reflectance or emission characteristics of ground objects. The present invention does not specifically limit the method for implementing spectral manipulation; for example, in some embodiments, radiometric correction, atmospheric correction, sensor correction, and other methods may be employed.

[0036] However, in order to ensure color consistency and visual quality of the color image while improving the authenticity and reliability of the data, in some exemplary embodiments of the present invention, the step of responding to the input remote sensing image and performing a spectral correction operation on the remote sensing image to obtain the first color image includes: According to the dimension of the remote sensing image, based on the principal component analysis method, the remote sensing image is subjected to dimensionality reduction processing to obtain a three-channel color image; The three-channel color image is transformed using a method specified by the International Illumination Commission to generate a first color image.

[0037] Principal component analysis (PCA) converts multi-band remote sensing images into three-channel color maps, reducing the amount of data to be processed and improving computational efficiency. Furthermore, PCA ensures that the converted three-channel color maps contain most of the information in the original image, facilitating subsequent analysis and processing. Furthermore, by highlighting the principal components, PCA enhances key features in the image and improves the accuracy of object recognition. Therefore, PCA can extract the principal components from remote sensing images, reducing data dimensionality while retaining the majority of the information.

[0038] Using the method specified by the International Commission on Illumination (CIE) for color transformation can, on the one hand, standardize the three-channel color images so that the processed images have consistent color performance on different display devices; on the other hand, it can optimize the visual effects of the three-channel color images, making the features of the ground objects clearer and easier for the human eye to recognize and analyze; in addition, adhering to international standards can make color images more suitable for a wide range of remote sensing applications, such as urban planning, environmental monitoring, resource management, etc.

[0039] In S220, the first color image is divided into superpixels using an edge recognition and superpixel division method that simulates a non-classical receptive field, and superpixel patches at different scales are obtained, including: Measuring the grayscale gradient of the first color image using a Laplacian operator to obtain a measurement result; The non-classical receptive field of the optic nerve is simulated to identify edge features of different scales, and Laplacian operators of different scales are used to obtain superpixel patches composed of different scales; among them, the independent area wrapped by the edge is defined as a superpixel.

[0040] The Laplacian operator is a second-order derivative operator used to detect grayscale changes in an image, particularly edges and corners. By measuring grayscale gradients, edge features can be effectively identified. The resulting measurement reflects the contrast and texture information of different image regions, helping to distinguish different ground features.

[0041] Methods that simulate the non-classical receptive field of the optic nerve can better mimic the human visual system's ability to detect edges, thereby identifying edge features in images at multiple scales. By using Laplacian operators at different scales, edges can be identified at different scales, helping to capture a variety of ground features, from fine to coarse. Based on multi-scale edge features, images can be segmented into superpixel patches with similar characteristics. These patches have a hierarchical scale and can better adapt to ground features of different scales.

[0042] Defining the edge-wrapped independent regions as superpixels can ensure the consistency within each superpixel, thereby ensuring that each superpixel has similar spectral and texture characteristics while being clearly distinguishable from adjacent regions.

[0043] In S230 , a clustering operation is performed on the superpixel patches to obtain a clustering result.

[0044] The present invention does not specifically limit the method for implementing superpixel patch clustering. For example, in some embodiments, a number of clustering algorithms may be employed, such as the K-means clustering algorithm, the hierarchical clustering algorithm, or the fuzzy C-means clustering algorithm. The selection and parameter settings of the clustering algorithm can be adjusted based on the characteristics of the remote sensing image, the desired clustering accuracy, and the intended application. For example, if the features in the remote sensing image are clearly defined and their number is known, K-means may be a suitable choice; if the features are complexly distributed and have diverse shapes, spectral clustering may be more appropriate.

[0045] Furthermore, the present invention allows these clustering methods to be combined with prior knowledge or auxiliary data (such as geographic information system data) to further improve the accuracy and practicality of clustering. By not limiting the specific clustering method, the present invention provides a wide range of possibilities for superpixel patch clustering, enhancing the versatility and adaptability of the technology.

[0046] Taking into account the accuracy and reliability of object recognition and the purpose of simplifying the data processing process, in the embodiment provided by the present invention, the superpixel patches are clustered to obtain the clustering results including: generating a representation map of independent features based on the superpixel patches at different scales; Analyze the spectral characteristics of all representation images to determine the categories of independent features; Constructing a mapping between remote sensing image data and the categories of the independent features to generate homologous reliable samples; The homologous reliable samples are output as the clustering results.

[0047] Here, the representation map highlights the characteristics of each individual feature, such as shape, size, and texture. Superpixel patches of varying scales can better adapt to features of varying sizes, ensuring the accuracy and comprehensiveness of the representation map. Spectral characteristic analysis can determine important information for feature identification, helping to improve classification accuracy. The analysis process can also reveal differences in the spectral characteristics of features, providing a basis for determining feature categories. By establishing a mapping relationship between remote sensing image data and individual feature categories, the consistency of homologous reliable samples with the original remote sensing image data can be ensured, improving sample representativeness.

[0048] The present invention does not limit the specific method of generating a representation map of an independent feature based on superpixel patches at different scales. For example, in some embodiments, a series of operations such as multi-scale segmentation, feature extraction, and feature fusion can be used to achieve the generation of a representation map of an independent feature.

[0049] In order to improve the accuracy and clarity requirements of the representation map of the independent feature, in some exemplary embodiments of the present invention, generating the representation map of the independent feature based on the superpixel patches at different scales includes: Performing scale normalization on super-pixel patches at different scales to obtain several normalized super-pixel patches; Eroding a number of normalized super-pixel patches to obtain a number of eroded super-pixel patches; Overlaying the eroded superpixel patches to obtain an overlay result; Screening out small-area noise in the superposition result to obtain a screening result; Different superpixel patches in the screening results are numbered based on preset rules to form a representation map of independent features.

[0050] Normalized superpixel patches have a uniform scale, facilitating unified processing and analysis. They also facilitate comparison and identification of features at different scales, improving the accuracy of feature extraction. Therefore, scale normalization ensures consistency between superpixel patches of different scales in subsequent processing, eliminating the impact of scale differences on analysis results.

[0051] The erosion operation helps remove unnecessary small-area noise (such as small holes and fine noise), making superpixel patches purer. It also smooths feature edges, making superpixel edges clearer and more regular, facilitating subsequent identification and classification. It's important to emphasize that the erosion radius is proportional to the scale, ensuring that small-scale edges are preserved when overlapping features are present.

[0052] The overlay results can show the spatial relationship between different superpixel patches, which helps to understand the spatial structure of the ground objects. Moreover, through overlay, features at different scales can be integrated, enhancing the overall representation of the ground objects and helping to reveal the spatial distribution and relationship of ground objects at different scales.

[0053] Screening out small-area noise can further purify the overlay results, ensuring that the information in the representation map mainly comes from actual ground objects rather than noise. On the one hand, it can effectively reduce the possibility of misclassification, and on the other hand, it can more focus on analyzing important ground object features.

[0054] The numbering operation assigns a unique identifier to each superpixel patch, which helps to track and identify specific features in subsequent analysis. Here, the present invention numbers different superpixel patches based on 4 together to form a representation map of independent features.

[0055] In summary, the method provided by the present invention for generating representation maps of independent land objects based on superpixel patches at different scales can effectively generate representation maps of independent land objects, improve the extraction accuracy of land object features, optimize the spatial relationship expression of land objects, enhance the purity of the results, and provide strong support for the in-depth analysis and application of remote sensing images.

[0056] In addition, the present invention also does not specifically limit the specific analysis content of analyzing the spectral characteristics of all representation images to determine the category of independent features. For example, in some embodiments, spectral index calculation, spectral angle matching, spectral matched filter, spectral cluster analysis, etc. can be used for determination.

[0057] In some exemplary embodiments of the present invention, analyzing the spectral characteristics of all representation images to determine the categories of independent features includes: Credible scale determination step: determining a credible scale of the current representation image based on the scale of each superpixel patch in the current representation image, wherein, when there are superpixel patches of multiple scales in the representation image, the median of the scales is used as the credible scale; Spectral band determination step: determining the spectral band of the current representation map based on the credible scale (i.e., based on the credible scale of each independent feature, finding the spectral distribution of the superpixel patch area from the image of the corresponding scale. Since the spectral fluctuation in this area is small, the average value can be taken as the spectral band of the current representation map); Repeating step: repeating the trust scale determination step and the spectral band determination step to obtain spectral bands of all feature maps; Adaptive clustering step: adaptively cluster the spectral bands of all feature maps to obtain clustering results; Category determination step: Based on the clustering results, the category of the independent features is determined.

[0058] By determining the credible scale of the representation map, we can identify the features of the objects represented by superpixel patches at specific scales. Identifying the scale of each superpixel patch helps us understand the behavior of objects at different scales. When superpixel patches exist at multiple scales, selecting the median scale as the credible scale can balance the influence of different scales and improve the robustness of classification.

[0059] Selecting spectral bands corresponding to the credible scale helps to extract spectral features related to the scale of the object; by matching bands, the spectral features of the object can be optimized to make it more suitable for classification purposes; therefore, determining the spectral bands according to the credible scale can ensure that the analyzed spectral information matches the actual scale of the object, thereby improving the accuracy of spectral analysis.

[0060] Repeating the trust scale determination step and the spectral band determination step ensures that the spectral bands of all feature maps are accurately identified, which is crucial to maintaining the consistency and accuracy of the entire analysis process.

[0061] Here, the present invention uses the Michelson contrast threshold of 15.27% provided by the Rayleigh criterion to adaptively cluster the spectral bands of all independent objects. Therefore, it can simulate the light and dark cognition ability of visual perception based on the visual Rayleigh brightness criterion, analyze the light intensity and spectrum of the scene and image, and realize accurate identification of independent object categories in remote sensing images.

[0062] In S240, the clustering result is used to train a multi-feature deformable lightweight neural network based on a visual attention mechanism to generate a superpixel region recognition model.

[0063] Structural reference of multi-feature deformable lightweight neural network based on visual attention mechanism Figure 3 shown.

[0064] The hierarchical cognitive nature of vision is that the cerebral visual cortex is divided into multiple regions according to their complexity. These regions do not process all visual electrical signals at once, but instead decompose multiple levels of individual elements from complex scenes in a sequence from simple to complex, thereby efficiently and multi-layeredly interpreting visual signals and understanding visual scenes. The visual attention mechanism mainly refers to the phenomenon that when the human eye observes a scene, its visual attention focuses on the target of interest or important area in the scene to form a fixation. Based on the visual attention mechanism, the present invention designs the encoding part of a multi-feature deformable lightweight neural network.

[0065] Multi-scale differential coding part reference Figure 4As shown, the multi-scale differential coding part includes an encoding input layer, a first transposed convolution amplification layer to a fourth transposed convolution amplification layer, a first feature mining unit to a fifth feature mining unit, a first self-attention weight assignment unit to a fifth self-attention weight assignment unit, a first average pooling layer, a second average pooling layer and an encoding output layer; wherein, the output of the encoding input layer is simultaneously used as the input of the first transposed convolution amplification layer, the input of the first average pooling layer, the input of the first feature mining unit and the input of the first self-attention weight assignment unit, the output of the first transposed convolution amplification layer is simultaneously used as the input of the second transposed convolution amplification layer, the input of the second feature mining unit and the input of the second self-attention weight assignment unit, the output of the second transposed convolution amplification layer is simultaneously used as the input of the third feature mining unit and the input of the third self-attention weight assignment unit, the output of the first average pooling layer is simultaneously used as the input of the fourth feature mining unit and the input of the fourth self-attention weight assignment unit, the output of the second average pooling layer is simultaneously used as the input of the fifth feature mining unit and the input of the fifth self-attention weight assignment unit, the output of the third feature mining unit and the output of the third self-attention weight assignment unit The output of the element is subjected to a first channel multiplication operation, and the results of the first channel multiplication are subjected to a first local maximum downsampling operation and a first pooling reduction operation respectively. The output of the second feature mining unit and the output of the first self-attention weight assignment unit are subjected to a second channel multiplication operation. The result of the second channel multiplication and the result of the first local maximum downsampling are subjected to a first feature subtraction operation, and a first feature channel superposition operation is performed on the result of the first pooling reduction. The result of the first feature subtraction is subjected to a second local maximum downsampling operation, and the first feature channel superposition result is subjected to a second pooling reduction operation. The output of the fifth feature mining unit and the output of the fifth self-attention weight assignment unit are subjected to a third channel multiplication operation. The result of the third channel multiplication is used as the input of the third transposed convolutional amplification layer, and a first nearest neighbor amplification operation is performed; the output of the fourth feature mining unit and the output of the fourth self-attention weight assignment unit are subjected to a fourth channel multiplication operation. The result of the fourth channel multiplication and the result of the first nearest neighbor amplification are subjected to a second feature subtraction operation, and a second feature channel superposition operation is performed on the output of the third transposed convolutional amplification layer. The result of the second feature subtraction is subjected to a second nearest neighbor amplification operation, and the result of the second feature channel superposition is used as the input of the fourth transposed convolutional amplification layer;The output of the first feature mining unit and the output of the first self-attention weight assignment unit are multiplied by a fifth channel. The result of the fifth channel multiplication is subtracted from the result of the second local maximum downsampling by a third feature, and then subtracted from the result of the second nearest neighbor amplification by a fourth feature. The result of the third feature subtraction, the result of the fourth feature subtraction, the result of the second pooling reduction, and the output of the fourth transposed convolution amplification layer are superimposed on the third feature channel. The result of the third feature channel superposition serves as the input of the encoding output layer, and the output of the encoding output layer serves as the output of the multi-scale differential encoding part.

[0066] The multi-scale differential coding part provided by the present invention uses average pooling to expand the scale to interpret low-resolution features, combined with transposed convolution to reduce the scale to analyze super-resolution features, and cooperates with the original features to perform multi-scale feature mining. Taking the original features as the baseline, it uses nearest neighbor sampling to amplify low-resolution features, and uses local maximum downsampling to reduce super-resolution features to construct a multi-scale feature differential structure. Finally, the super-resolution features are reduced by pooling, and the low-resolution features are amplified by transposed convolution. Through residual connections, various feature information is aggregated to establish a multi-scale differential coding structure and output the mined hidden features.

[0067] Multi-scale receptive field decoding part reference Figure 5 As shown, the multi-scale receptive field decoding part includes a decoding input layer, a first 1×1 convolutional layer, a first 3×3 convolutional layer to a sixth 3×3 convolutional layer, a channel superposition layer of the first feature, a first ReLU activation layer, a second ReLU activation layer and a first Softmax probability layer; wherein, the output of the encoding output layer is used as the input of the decoding input layer, and the output of the decoding input layer is also used as the input of the first 1×1 convolutional layer and the first 3×3 convolutional layer. The first 3×3 convolutional layer to the fifth 3×3 convolutional layer are arranged in sequence, and the output of the first 3×3 convolutional layer to the output of the fifth 3×3 convolutional layer and the output of the first 1×1 convolutional layer are superimposed in the channel superposition layer of the first feature, and then a self-attention weight allocation operation is performed, and the operation result is multiplied by the feature channel, and the multiplication result is used as the input of the first ReLU activation layer. The first ReLU activation layer, the sixth 3×3 convolutional layer, the second ReLU activation layer and the first Softmax probability layer are arranged in sequence, and the output of the first Softmax probability layer is the output of the multi-scale receptive field decoding part.

[0068] Specifically, it combines 1×1 convolutions to construct multi-convolutional templates, residual connections, and channel stacking, along with a self-attention weight distribution structure. This is combined with a Softmax probability function segmenter, using cross-entropy as the loss function for decoding and segmentation. This multi-receptive field decoding structure maps the association between features and segmentation, forming a complete chain of backward and forward propagation, and constructing a top-down observation neural network to achieve detailed observation of surface scenes.

[0069] On the basis of the above solution, the first to fifth feature mining units are constructed to have the same structure.

[0070] refer to Figure 6 As shown, the feature mining unit includes: the seventh 3×3 convolution layer to the tenth 3×3 convolution layer, the third ReLU activation layer to the sixth ReLU activation layer, the channel superposition layer of the second feature and the channel superposition layer of the third feature; Among them, the input of the seventh 3×3 convolutional layer is used as the input of the feature mining unit, and the seventh 3×3 convolutional layer, the third ReLU activation layer, the eighth 3×3 convolutional layer, the fourth ReLU activation layer and the channel superposition layer of the second feature are arranged in sequence; the ninth 3×3 convolutional layer, the fifth ReLU activation layer, the tenth 3×3 convolutional layer, the sixth ReLU activation layer and the channel superposition layer of the third feature are arranged in sequence; the output of the channel superposition layer of the second feature is also used as the input of the channel superposition layer of the third feature, and the output of the channel superposition layer of the third feature is used as the output of the feature mining unit.

[0071] On the basis of the above scheme, the first to fifth self-attention weight allocation units are constructed to have the same structure; refer to Figure 7 As shown, each attention weight allocation unit includes an eleventh 3×3 convolution layer, a twelfth 3×3 convolution layer, a seventh ReLU activation layer, a second 1×1 convolution layer, an eighth ReLU activation layer and a second Softmax probability layer, which are arranged in sequence.

[0072] In some exemplary embodiments of the present invention, referring to the perception mechanism of eye tremor, which recognizes context from the bottom up and analyzes details from the top down, after using the superpixel region recognition model to perform remote sensing image segmentation, the remote sensing image self-learning segmentation method further includes: Based on the eye tremor perception mechanism, errors in the segmentation results and superpixel patches at different scales are repaired.

[0073] On this basis, the steps specifically include: Map the clustering results to each superpixel patch; The category of the independent feature that appears most frequently on each superpixel patch is regarded as the category of the main independent feature of the current superpixel patch; The segmentation results are optimized according to the categories of the main independent objects of all superpixel patches.

[0074] The mechanism of eye tremor perception is that during microtremors, peripheral rods form a blurred vision, providing a fuzzy initial perception of the scene. Cones in the fovea, building on this initial perception of blurred vision, then conduct a detailed observation of details such as edges, textures, and colors, thereby forming a clear field of vision and completing scene perception. Based on a bottom-up global parsing process to create superpixels, each superpixel is considered a unit containing only a single type of object.

[0075] The object classification results from the top-down neural network are mapped to each superpixel. The object classification with the highest number of objects in the superpixel is then used as the segmentation result for that superpixel, obtaining the segmentation results for all superpixels. Through a top-down, detailed neural network observation, combined with the reference information provided by the superpixel segmentation, errors are removed and edges are repaired.

[0076] According to a second aspect of an embodiment of the present invention, a remote sensing image self-learning segmentation device 800 is provided, referring to Figure 8 As shown, the remote sensing image self-learning segmentation device 800 includes: The spectral correction module 810 is configured to respond to an input remote sensing image and perform a spectral correction operation on the remote sensing image to obtain a first color image; A superpixel division module 820 is configured to perform superpixel division on the first color image by using an edge recognition and superpixel division method simulating a non-classical receptive field to obtain superpixel patches at different scales; A clustering operation module 830 is used to perform a clustering operation on the superpixel patches to obtain a clustering result; A model building module 840 is configured to use the clustering results to train a multi-feature deformable lightweight neural network based on a visual attention mechanism to generate a superpixel region recognition model; The image segmentation module 850 is used to perform remote sensing image segmentation using the superpixel region recognition model.

[0077] It should be noted that although several modules of the remote sensing image self-learning segmentation device 700 are mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules or submodules described above can be embodied in a single module or unit. Conversely, the features and functions of a single module described above can be further divided and embodied by multiple modules or submodules.

[0078] In addition, in an exemplary embodiment of the present invention, an electronic device capable of implementing the above-mentioned remote sensing image self-learning segmentation method is also provided.

[0079] Those skilled in the art will appreciate that various aspects of the present invention may be implemented as systems, methods, or program products. Therefore, various aspects of the present invention may be implemented as a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0080] Refer to the following Figure 9 An electronic device 900 according to such an embodiment of the present invention will be described. Figure 9 The electronic device 900 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0081] like Figure 9 As shown, electronic device 900 is implemented as a general-purpose computing device. Components of electronic device 900 may include, but are not limited to, the aforementioned at least one processing unit 910, the aforementioned at least one storage unit 920, a bus 930 connecting various system components (including storage unit 920 and processing unit 910), and a display unit 940.

[0082] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 performs the steps of various exemplary embodiments of the present invention described in the above “Exemplary Method” section. For example, the processing unit 910 can perform the following steps: Figure 2 As shown in S210, in response to the input remote sensing image, a spectral correction operation is performed on the remote sensing image to obtain a first color image; S220, superpixel division is performed on the first color image using an edge recognition and superpixel division method that simulates a non-classical receptive field to obtain superpixel patches at different scales; S230, a clustering operation is performed on the superpixel patches to obtain clustering results; S240, a multi-feature deformable lightweight neural network based on a visual attention mechanism is trained using the clustering results to generate a superpixel region recognition model; S250, remote sensing image segmentation is performed using the superpixel region recognition model.

[0083] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 921 and / or a cache memory unit 922 , and may further include a read-only memory unit (ROM) 923 .

[0084] The storage unit 920 may also include a program / utility 924 having a set (at least one) of program modules 925, such program modules 925 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0085] Bus 930 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0086] The electronic device 900 can also communicate with one or more external devices 970 (e.g., a keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 900, and / or any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., a router, modem, etc.). This communication can occur via an input / output (I / O) interface 950. Furthermore, the electronic device 900 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 900, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0087] From the above description of the embodiments, those skilled in the art will readily appreciate that the exemplary embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored on a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes instructions for causing a computing device (such as a personal computer, server, terminal device, or network device) to execute the methods according to the embodiments of the present invention.

[0088] In an exemplary embodiment of the present invention, a computer-readable storage medium is further provided, on which is stored a program product capable of implementing the above-described method of the present invention. In some possible embodiments, various aspects of the present invention may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is used to cause the terminal device to perform the steps according to the various exemplary embodiments of the present invention described in the "Exemplary Method" section above.

[0089] refer to Figure 10 FIG. 1 illustrates a program product 1000 for implementing the aforementioned underground pipeline detection method according to an embodiment of the present invention. The program product 1000 may be implemented in a portable compact disc read-only memory (CD-ROM) and include program code, and may be executed on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In the present invention, a readable storage medium may be any tangible medium containing or storing a program, which may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0090] The program product may utilize any combination of one or more readable storage media. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0091] Program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0092] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0093] From the above description of the embodiments, it will be readily apparent to those skilled in the art that the exemplary embodiments described herein can be implemented via software or via a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present invention can be embodied in the form of a software product, which can be stored on a non-volatile storage medium (such as a CD-ROM, USB flash drive, or mobile hard drive) or on a network and includes instructions for causing a computing device (such as a personal computer, server, touchscreen terminal, or network device) to execute the methods according to the embodiments of the present invention.

[0094] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the present invention that follow from the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0095] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A remote sensing image self-learning segmentation method, characterized in that: include: In response to an input remote sensing image, performing a spectral correction operation on the remote sensing image to obtain a first color image; Performing superpixel division on the first color image using an edge recognition and superpixel division method simulating a non-classical receptive field to obtain superpixel patches at different scales; Performing a clustering operation on the superpixel patches to obtain a clustering result; Using the clustering results to train a multi-feature deformable lightweight neural network based on a visual attention mechanism to generate a superpixel region recognition model; The superpixel region recognition model is used to perform remote sensing image segmentation.

2. The remote sensing image self-learning segmentation method according to claim 1, characterized in that: The step of responding to the input remote sensing image and performing a spectral correction operation on the remote sensing image to obtain a first color image includes: According to the dimension of the remote sensing image, based on the principal component analysis method, the remote sensing image is subjected to dimensionality reduction processing to obtain a three-channel color image; Transforming the three-channel color image using a method specified by the International Illumination Commission to generate a first color image; The first color image is divided into superpixels using an edge recognition and superpixel division method simulating a non-classical receptive field, and superpixel patches at different scales are obtained, including: Measuring the grayscale gradient of the first color image using a Laplacian operator to obtain a measurement result; The non-classical receptive field of the optic nerve is simulated to identify edge features of different scales, and Laplacian operators of different scales are used to obtain superpixel patches composed of different scales; among them, the independent area wrapped by the edge is defined as a superpixel.

3. The remote sensing image self-learning segmentation method according to claim 1, characterized in that: The superpixel patches are clustered to obtain clustering results including: generating a representation map of independent features based on the superpixel patches at different scales; Analyze the spectral characteristics of all representation images to determine the categories of independent features; Constructing a mapping between remote sensing image data and the categories of the independent features to generate homologous reliable samples; The homologous reliable samples are output as the clustering results.

4. The remote sensing image self-learning segmentation method according to claim 3, characterized in that: Generating a representation map of an independent feature based on the superpixel patches at different scales includes: Performing scale normalization on super-pixel patches at different scales to obtain several normalized super-pixel patches; Eroding a number of normalized super-pixel patches to obtain a number of eroded super-pixel patches; Overlaying the eroded superpixel patches to obtain an overlay result; Screening out small-area noise in the superposition result to obtain a screening result; Different superpixel patches in the screening results are numbered based on preset rules to form a representation map of independent features.

5. The remote sensing image self-learning segmentation method according to claim 3, characterized in that: The spectral characteristics of all representation images are analyzed to determine the categories of independent objects including: Credible scale determination step: determining a credible scale of the current representation image based on the scale of each superpixel patch in the current representation image, wherein, when there are superpixel patches of multiple scales in the representation image, the median of the scales is used as the credible scale; Spectral band determination step: determining the spectral band of the current representation image according to the credible scale; Repeating step: repeating the trust scale determination step and the spectral band determination step to obtain spectral bands of all feature maps; Adaptive clustering step: adaptively cluster the spectral bands of all feature maps to obtain clustering results; Category determination step: Based on the clustering results, the category of the independent features is determined.

6. The remote sensing image self-learning segmentation method according to any one of claims 1 to 5, characterized in that: The multi-feature deformable lightweight neural network based on the visual attention mechanism includes: A multi-scale differential encoding part, the multi-scale differential encoding part includes an encoding input layer, a first transposed convolution amplification layer to a fourth transposed convolution amplification layer, a first feature mining unit to a fifth feature mining unit, a first self-attention weight assignment unit to a fifth self-attention weight assignment unit, a first average pooling layer, a second average pooling layer and an encoding output layer; wherein, the output of the encoding input layer is simultaneously used as the input of the first transposed convolution amplification layer, the input of the first average pooling layer, the input of the first feature mining unit and the input of the first self-attention weight assignment unit, the output of the first transposed convolution amplification layer is simultaneously used as the input of the second transposed convolution amplification layer, the input of the second feature mining unit and the input of the second self-attention weight assignment unit, the output of the second transposed convolution amplification layer is simultaneously used as the input of the third feature mining unit and the input of the third self-attention weight assignment unit, the output of the first average pooling layer is simultaneously used as the input of the fourth feature mining unit and the input of the fourth self-attention weight assignment unit, the output of the second average pooling layer is simultaneously used as the input of the fifth feature mining unit and the input of the fifth self-attention weight assignment unit, the output of the third feature mining unit and the third self-attention weight assignment unit The output of the redistribution unit performs a first channel multiplication operation, and the results of the first channel multiplication perform a first local maximum downsampling operation and a first pooling reduction operation respectively. The output of the second feature mining unit and the output of the first self-attention weight assignment unit perform a second channel multiplication operation, the result of the second channel multiplication and the result of the first local maximum downsampling perform a first feature subtraction operation, and perform a first feature channel superposition operation with the result of the first pooling reduction, the result of the first feature subtraction performs a second local maximum downsampling operation, and the first feature channel superposition result performs a second pooling reduction operation, the output of the fifth feature mining unit and the output of the fifth self-attention weight assignment unit perform a third channel multiplication operation, the result of the third channel multiplication serves as the input of the third transposed convolutional amplification layer, and performs a first nearest neighbor amplification operation; the output of the fourth feature mining unit and the output of the fourth self-attention weight assignment unit perform a fourth channel multiplication operation, the result of the fourth channel multiplication and the result of the first nearest neighbor amplification perform a second feature subtraction operation, and perform a second feature channel superposition operation with the output of the third transposed convolutional amplification layer, the result of the second feature subtraction performs a second nearest neighbor amplification operation, and the result of the second feature channel superposition serves as the input of the fourth transposed convolutional amplification layer;The output of the first feature mining unit and the output of the first self-attention weight allocation unit are multiplied by a fifth channel. The result of the fifth channel multiplication is subtracted from the result of the second local maximum downsampling by a third feature, and is subtracted from the result of the second nearest neighbor amplification by a fourth feature. The result of the third feature subtraction, the result of the second pooling reduction of the result of the fourth feature subtraction, and the output of the fourth transposed convolution amplification layer are superimposed on the third feature channel. The result of the third feature channel superposition is used as the input of the encoding output layer, and the output of the encoding output layer is the output of the multi-scale differential encoding part. The multi-scale receptive field decoding part includes a decoding input layer, a first 1×1 convolutional layer, a first 3×3 convolutional layer to a sixth 3×3 convolutional layer, a channel superposition layer of the first feature, a first ReLU activation layer, a second ReLU activation layer, and a first Softmax probability layer; wherein the output of the encoding output layer is used as the input of the decoding input layer, and the output of the decoding input layer is also used as the input of the first 1×1 convolutional layer and the first 3×3 convolutional layer. The first 3×3 convolutional layer to the fifth 3×3 convolutional layer are arranged in sequence, and the output of the first 3×3 convolutional layer to the output of the fifth 3×3 convolutional layer and the output of the first 1×1 convolutional layer are superimposed in the channel superposition layer of the first feature, and then a self-attention weight allocation operation is performed, and the operation result is multiplied by the feature channel. The multiplication result is used as the input of the first ReLU activation layer. The first ReLU activation layer, the sixth 3×3 convolutional layer, the second ReLU activation layer, and the first Softmax probability layer are arranged in sequence, and the output of the first Softmax probability layer is the output of the multi-scale receptive field decoding part.

7. The remote sensing image self-learning segmentation method according to claim 6, characterized in that: The first to fifth feature mining units are constructed to have the same structure; The feature mining unit includes: the seventh 3×3 convolution layer to the tenth 3×3 convolution layer, the third ReLU activation layer to the sixth ReLU activation layer, the channel superposition layer of the second feature and the channel superposition layer of the third feature; Among them, the input of the seventh 3×3 convolutional layer is used as the input of the feature mining unit, and the seventh 3×3 convolutional layer, the third ReLU activation layer, the eighth 3×3 convolutional layer, the fourth ReLU activation layer and the channel superposition layer of the second feature are arranged in sequence; the ninth 3×3 convolutional layer, the fifth ReLU activation layer, the tenth 3×3 convolutional layer, the sixth ReLU activation layer and the channel superposition layer of the third feature are arranged in sequence; the output of the channel superposition layer of the second feature is also used as the input of the channel superposition layer of the third feature, and the output of the channel superposition layer of the third feature is used as the output of the feature mining unit.

8. The remote sensing image self-learning segmentation method according to claim 6, characterized in that: The first to fifth self-attention weight allocation units are constructed to have the same structure; Each attention weight allocation unit includes an eleventh 3×3 convolution layer, a twelfth 3×3 convolution layer, a seventh ReLU activation layer, a second 1×1 convolution layer, an eighth ReLU activation layer and a second Softmax probability layer, which are arranged in sequence.

9. The remote sensing image self-learning segmentation method according to claim 1, characterized in that: After performing remote sensing image segmentation using the superpixel region recognition model, the remote sensing image self-learning segmentation method further includes: Based on the eye tremor perception mechanism, errors in the segmentation results and superpixel patches at different scales are repaired.

10. The remote sensing image self-learning segmentation method according to claim 9, characterized in that: The error repair of the segmentation results and superpixel patches at different scales based on the eye tremor perception mechanism includes: Map the clustering results to each superpixel patch; The category of the independent feature that appears most frequently on each superpixel patch is regarded as the category of the main independent feature of the current superpixel patch; The segmentation results are optimized according to the categories of the main independent objects of all superpixel patches.

Citation Information

Cited By

  • Multi-source data collaborative snake-green hybrid rock air-space-ground integrated interpretation method and system

    CN121789077A