Remote sensing image ground object migration identification method, system, device and storage medium
The remote sensing image feature migration identification method based on manifold topological induction and vector field theory physical constraints solves the problems of data distribution offset and topological representation difficulties in remote sensing image identification, realizes accurate identification and high-fidelity reconstruction of feature skeletons, and improves the accuracy and efficiency of remote sensing image identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WENZHOU ELECTRIC POWER BUREAU
- Filing Date
- 2026-04-21
- Publication Date
- 2026-05-29
Smart Images

Figure CN122116185A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing image recognition technology, and in particular to a method, system, device, and storage medium for recognizing the migration of ground features in remote sensing images. Background Technology
[0002] With the rapid development of remote sensing Earth observation technology, massive amounts of high-resolution remote sensing imagery have become a fundamental data source for Earth system science and smart city construction. However, in practical applications of cross-domain ground feature identification, existing deep learning paradigms are still constrained by several fundamental theoretical bottlenecks.
[0003] On the one hand, manifold shifts in data distribution severely hinder the generalization ability of models. High-performance models typically assume that the source and target domain data follow the same statistical distribution. However, remote sensing images are affected by illumination, season, and sensor imaging mechanisms, resulting in significant nonlinear manifold distortions in the feature space. Traditional convolutional neural networks mostly fit statistical textures in Euclidean space, lacking the ability to generalize the physical essence of images. When texture noise not seen in the source domain appears in the target domain, the model is prone to misclassifying it as ground feature, leading to a severe drop in cross-domain recognition performance. On the other hand, the physical field representation difficulties of geometric topology limit the extraction accuracy of subtle ground features. In complex surface scenes, linear features such as roads and rivers physically represent continuous flowing helical fields, while background noise often represents discrete divergent active fields. Existing feature extraction operators are mostly mathematically represented as isotropic filters, lacking the ability to perceive the curl and divergence of vector fields. This prevents models from effectively distinguishing between linear skeletons and speckle noise at the physical topology level, resulting in road breaks or false alarms in the background.
[0004] Furthermore, the inverse mapping from low-frequency semantics to high-frequency details presents an inherent well-posed challenge. While commonly used large-scale visual models have established powerful general semantic representations, their output features are essentially highly compressed low-frequency potential fields with extremely low resolution. Existing decoders often employ simple bilinear interpolation or deconvolution for upsampling, resulting in often blurred boundaries of reconstructed ground features, which fails to meet the requirements of high-precision mapping. Summary of the Invention
[0005] To address the aforementioned technical problems, this invention provides a method, system, device, and storage medium for remote sensing image feature migration identification. Through manifold topological induction, vector field theory physical constraints, and dynamic evolution mechanisms, remote sensing feature migration identification can achieve the technical effect of physical-level decoupling and high-fidelity reconstruction of the feature skeleton under zero-sample conditions.
[0006] In a first aspect, the present invention provides a method for identifying the migration of ground features in remote sensing images, the method comprising: Acquire remote sensing images of the area to be measured, and divide the remote sensing images into several local images using an overlapping sliding window mechanism; A feature extraction model is used to extract multi-level features from each local image, and noise filtering and dimension reshaping are performed on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The deep features of the two-dimensional feature pyramid are used as a low-frequency potential energy field, and spiral vector field induction and potential field decoupling enhancement are performed on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The high-frequency singularity field and the low-frequency potential energy field are coupled and hyperspectral manifold embedded mapping is performed to obtain a high-dimensional hyperspectral manifold. The high-dimensional hyperspectral manifold is then subjected to singularity-driven anisotropic gradient flow evolution and inverse mapping decoding to obtain the decoding features of each local image. Linear classification is performed on the decoded features of each local image, and the predicted tensor probabilities obtained from the classification are accumulated and normalized to obtain the ground feature classification and recognition results.
[0007] Furthermore, the step of using a feature extraction model to perform multi-level feature extraction on each local image, and then performing noise filtering and dimension reshaping on the extracted multi-level features to obtain a two-dimensional feature pyramid includes: The standardized local image is input into a feature extraction model built on the DINOv3 network to perform multi-level feature extraction, resulting in multi-level features, which include one deep feature and several shallow features. A semantic standing wave resonance mechanism is adopted to calculate the resonance coefficient between each shallow feature and the deep feature, and to perform semantic modulation on the shallow feature based on the resonance coefficient to obtain the semantic resonance feature corresponding to each shallow feature. A layer-by-layer differential entropy screening mechanism is used to filter noise from each semantic resonance feature, resulting in multi-level purified features. For each layer of purification features in the multi-level purification features, target domain structure manifold embedding and physical dimension inverse mapping are performed to obtain an initial two-dimensional feature map; The initial two-dimensional feature map is semantically repaired based on the anisotropic topological healing mechanism to obtain a two-dimensional feature map. The two-dimensional feature map is then dimensionally projected to obtain the solidified features corresponding to each layer of purified features. The solidified features of each layer are then assembled to obtain a two-dimensional feature pyramid.
[0008] Furthermore, the step of inducing a spiral vector field and decoupling the potential field to enhance each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field includes: Perform a normalized field mapping on each shallow feature of the two-dimensional feature pyramid to generate a vector potential; The vector potential is convolved with a discrete curl operator having an antisymmetric cross topology to generate a vortex vector field, and the vortex vector field is dynamically shaped with an anisotropic induced tensor to generate a spiral characteristic field. The divergence source density of the spiral characteristic field is calculated using a discrete divergence operator, and a divergence potential energy surface is constructed based on the divergence source density. Based on the divergence potential energy surface, a purity-gated mask is calculated, and noise is stripped from the spiral feature field using the purity-gated mask to obtain a purified feature field. The purified feature field is reconstructed using the Poisson equation to obtain the purified feature vector field, and then the purified feature vector field is scalar transformed to obtain the scalar response features. The scalar response features of each shallow feature are spliced and recombined to obtain a high-frequency singularity field.
[0009] Furthermore, the step of coupling the high-frequency singularity field and the low-frequency potential energy field and embedding them into a hyperspectral manifold to obtain a high-dimensional hyperspectral manifold includes: According to the high-frequency singularity field, the low-frequency potential energy field is aligned with the spatial coordinate system to obtain the initial potential energy field, and the high-frequency singularity field and the initial potential energy field are coupled to obtain the coupled field tensor. Hyperspectral manifold dilation is performed on the coupled field tensor to obtain an initial high-dimensional hyperspectral manifold; The initial high-dimensional hyperspectral manifold is subjected to manifold statistical normalization and Gaussian error linear mapping to obtain the high-dimensional hyperspectral manifold.
[0010] Furthermore, the step of performing singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image includes: The manifold spatial gradient field is calculated for the high-dimensional hyperspectral manifold to obtain the manifold gradient vector field; Based on the high-frequency singularity field, a singularity gate is generated, and based on the singularity gate, anisotropic impact evolution is performed on the manifold gradient vector field to obtain evolution characteristics. The evolutionary features are orthogonally projected and collapsed to obtain low-dimensional features, and the low-dimensional features are added to the initial potential field to obtain the decoded features.
[0011] Further, the step of performing anisotropic impact evolution on the manifold gradient vector field based on the singularity gating to obtain evolution characteristics includes: The singularity gate is nonlinearly combined with the manifold gradient vector field to obtain the evolution increment; The evolutionary features are obtained by adding the high-dimensional hyperspectral manifold to the evolutionary increment.
[0012] Furthermore, the steps of performing linear classification on the decoded features of each local image, and accumulating and normalizing the predicted tensor probabilities obtained from the classification to obtain the land cover classification and recognition results include: Linear classification is performed on the decoded features of each local image to obtain the predicted probability tensor of each local image; According to the coordinate region of the remote sensing image, each predicted probability tensor is added together and counted to obtain the cumulative probability tensor and the cumulative count tensor. Based on the accumulated probability tensor and the accumulated count tensor, a pixel-by-pixel probability averaging calculation is performed, and a full-map pixel-level land feature classification mask is generated according to the maximum probability principle. Based on the pixel-level land feature classification mask of the entire image, the land feature classification and recognition results are obtained.
[0013] In a second aspect, the present invention provides a remote sensing image feature migration identification system, the system comprising: The image segmentation module is used to acquire remote sensing images of the area to be tested and to divide the remote sensing images into several local images using an overlapping sliding window mechanism. The feature extraction module is used to perform multi-level feature extraction on each local image using a feature extraction model, and to perform noise filtering and dimension reshaping on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The feature decoupling module is used to take the deep features of the two-dimensional feature pyramid as a low-frequency potential energy field, and to perform spiral vector field induction and potential field decoupling enhancement on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The cascaded reconstruction module is used to couple and embed the high-frequency singularity field and the low-frequency potential energy field into a hyperspectral manifold to obtain a high-dimensional hyperspectral manifold, and to perform singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image. The sliding window inference module is used to perform linear classification on the decoded features of each local image, and to accumulate and normalize the predicted tensor probabilities obtained from the classification to obtain the ground feature classification and recognition results.
[0014] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0015] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0016] This invention provides a method, system, device, and storage medium for remote sensing image feature migration identification. Through a sequence spatial inverse reshaping mechanism based on target domain manifold induction, this invention provides a semantically dense and topologically complete manifold basis for subsequent processing. Through a decoupling mechanism between divergence- and curl-induced spiral vector field convolution and Helmholtz potential field, it achieves zero-sample, extremely pure extraction in unlabeled target domains. Through a manifold inverse mapping decoding architecture based on singularity potential coupling, it achieves accurate detail reconstruction. This invention, through manifold topological induction, vector field theory physical constraints, and dynamic evolution mechanisms, enables physical-level decoupling and high-fidelity reconstruction of the feature skeleton under zero-sample conditions, thereby achieving accurate and efficient remote sensing feature migration identification. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating the remote sensing image feature migration identification method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the antisymmetric cross-vortex kernel structure of the discrete curl convolution operator in this embodiment of the invention; Figure 3 This is a schematic diagram of the structure of the remote sensing image ground feature migration recognition system in an embodiment of the present invention; Figure 4 This is an internal structural diagram of the computer device in an embodiment of the present invention.
[0018] Figure label: 10. Image segmentation module; 20. Feature extraction module; 30. Feature decoupling module; 40. Cascaded reconstruction module; 50. Sliding window inference module. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Please see Figure 1 The first embodiment of the present invention proposes a method for identifying the migration of ground features in remote sensing images, comprising steps S10 to S50: Step S10: Acquire remote sensing images of the area to be tested, and divide the remote sensing images into several local images using an overlapping sliding window mechanism; Step S20: A feature extraction model is used to extract multi-level features from each local image, and noise filtering and dimension reshaping are performed on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. Step S30: The deep features of the two-dimensional feature pyramid are used as a low-frequency potential energy field, and spiral vector field induction and potential field decoupling enhancement are performed on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. Step S40: Couple the high-frequency singularity field and the low-frequency potential energy field and perform hyperspectral manifold embedding mapping to obtain a high-dimensional hyperspectral manifold. Then, perform singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image. In step S50, the decoded features of each local image are linearly classified, and the predicted tensor probabilities obtained from the classification are accumulated and normalized to obtain the ground feature classification and recognition results.
[0021] In this embodiment, remote sensing images of the area to be tested are first acquired. Then, addressing the issue that the large size of the remote sensing images prevents them from being input into the network all at once, an overlapping sliding window strategy is employed for image segmentation. Specifically, the size and step size of the sliding window are set. Preferably, the step size is set to half the size length, i.e., a 50% overlap rate. A set of coordinates covering the top-left corner of the entire image is generated using grid coordinate calculations. For areas with edges less than one step size, coordinates are forcibly backtracked to ensure the sliding window does not exceed its boundaries. For each sliding window position in the top-left corner coordinate set, a local image is cropped, thus dividing the remote sensing image into multiple local images. Probabilistic prediction is then performed on each local image. The prediction steps for the local images are described in detail below.
[0022] In this embodiment, the local image is first standardized to eliminate the dimensional differences caused by illumination intensity, mapping it to standard dimensionless physical quantities. Then, a large visual model is used to extract deep semantics, and a dynamic standing wave resonance mechanism is used to filter out high-frequency background noise irrelevant to the task at the one-dimensional sequence level. The one-dimensional sequence after noise filtering is then reshaped to remap it into topologically continuous two-dimensional image features. The specific steps include: The standardized local image is input into a feature extraction model built on the DINOv3 network to perform multi-level feature extraction, resulting in multi-level features, which include one deep feature and several shallow features. A semantic standing wave resonance mechanism is adopted to calculate the resonance coefficient between each shallow feature and the deep feature, and to perform semantic modulation on the shallow feature based on the resonance coefficient to obtain the semantic resonance feature corresponding to each shallow feature. A layer-by-layer differential entropy screening mechanism is used to filter noise from each semantic resonance feature, resulting in multi-level purified features. For each layer of purification features in the multi-level purification features, target domain structure manifold embedding and physical dimension inverse mapping are performed to obtain an initial two-dimensional feature map; The initial two-dimensional feature map is semantically repaired based on the anisotropic topological healing mechanism to obtain a two-dimensional feature map. The two-dimensional feature map is then dimensionally projected to obtain the solidified features corresponding to each layer of purified features. The solidified features of each layer are then assembled to obtain a two-dimensional feature pyramid.
[0023] In this embodiment, a feature extraction model is built using the DINOv3 backbone network. A standardized local image is input into the feature extraction model, and the input image is discretized into an unordered sequence containing N tokens through the DINOv3 patch embedding layer. At this time, the spatial neighborhood relationship is broken, and then multi-level feature extraction is performed. Taking a 12-layer network architecture (12 Transformer layers) as an example, the intermediate hidden states of the 3rd, 6th, 9th, and 12th layers of the Transformer Block are specifically extracted, which correspond to low, medium, medium-high, and high-level features, respectively. Of course, if the model has 24 layers, multi-level features can be extracted from the 6th, 12th, 18th, and 24th layers of the Transformer. For the sake of consistent description, this embodiment refers to high-level features as deep features and the remaining layer features as shallow features.
[0024] Taking a 12-layer structure as an example, the deep features of the 12th layer are... Defined as a strong semantic guidance signal, it incorporates shallow features from layers 3, 6, and 9. Defined as a rich-textured basic signal, where, .
[0025] The conventional approach to generating two-dimensional image features from the output of the DINOv3 model is to use a fixed-level extraction and simple dimensional reshaping paradigm. However, the original output of the DINOv3 model is limited to containing only high-level abstract semantics and losing spatial topology. To overcome this problem and considering the large scale span of remote sensing images, the abundance of background redundancy, and the unique geometric patterns of spatial distribution under a top-down view, this embodiment abandons the traditional fixed-level extraction and simple dimensional reshaping paradigm and designs two cascaded mechanisms: "dynamic pyramid construction based on semantic standing wave resonance" and "sequence-spatial inverse reshaping based on target domain manifold induction." This mechanism aims to establish a physical mapping bridge from the general semantic space to the geometric space of remote sensing objects, which will be explained below: S21. Pyramid Construction Based on Semantic Standing Wave Resonance and Interlayer Entropy Screening Traditional methods typically employ static indexing to extract features. This "hard truncation" approach severs the inherent semantic flow between Transformer layers, resulting in shallow features containing a large amount of high-frequency noise irrelevant to the task (such as cloud and wave textures), while deep features lose crucial edge details. In the Transformer architecture, deep features contain the strongest category semantic representation (such as "identified as a bridge") but lack spatial localization capabilities; while shallow features contain rich texture details but have a low signal-to-noise ratio. Therefore, this embodiment constructs an inter-layer filtering mechanism based on inverse semantic standing waves, utilizing deep features as a "semantic standing wave source" to inversely stimulate and filter shallow features.
[0026] Specifically, to suppress shallow features Irrelevant backgrounds (such as clouds and wave textures) are utilized by deep features. Using the reference signal, a reverse resonant wave is sent to the shallow features. Define the first... The semantic resonance coefficient matrix R of the layer i as follows: in, It is the Sigmoid activation function. The StopGradient operation represents the action that blocks gradient backpropagation. This is the scaling factor.
[0027] The physical meaning of the semantic resonance coefficient matrix lies in its quantification of the resonance strength between each token in the shallow layer and the target semantics in the deep layer. If a texture patch in the shallow layer is identified as belonging to the target feature in the deep layer, the resonance coefficient approaches 1; if it is an irrelevant background, it approaches 0. The resonance coefficient is used to semantically modulate the shallow features, generating semantic resonance features. Its expression is: In the formula, This represents the semantic resonance feature of the i-th layer. This represents the shallow features of the i-th layer. It is the product of Hadama.
[0028] Through the above formula, the model introduces top-level semantic design at the source of feature extraction, forcing shallow features to retain only the texture details that resonate with the final recognition task.
[0029] Then, high-energy region extraction is performed using inter-layer differential entropy. To address the computational redundancy caused by large areas of low-information regions (such as continuous water bodies or bare land) in remote sensing imagery, this embodiment introduces information entropy filtering. Unlike retaining all tokens, this embodiment calculates the local information entropy density of features, retaining only high-energy regions. Specifically, the local information entropy of the p-th token in the i-th feature layer is calculated. Then, a dynamic entropy mask is constructed based on the local information entropy. : In the formula, For indicator functions, and Let be the mean and standard deviation of the local information entropy of the i-th feature layer, respectively. Let H be the sensitivity coefficient. For simplicity, in the above formula, H is used as... i The local information entropy of the i-th feature layer is uniformly represented.
[0030] Finally, the calculated dynamic entropy mask is applied to the semantic resonance features generated in the above steps. This step aims to filter out background noise that, while resonant, has extremely low information content, ensuring the purity of the features. The final feature pyramid hierarchy is calculated as follows: In the formula, This represents the purified feature of the i-th feature layer. Represents the tensor product. .
[0031] For deep features Then it is directly defined as a purified sequence. Thus, multi-level one-dimensional purification characteristics are obtained. .
[0032] Through the interlayer filtering mechanism based on inverse semantic standing waves in this embodiment, the energy level filtering of features is realized, low-frequency redundant regions are automatically eliminated, and a highly efficient feature pyramid with semantic focus and dense information is output.
[0033] S22. Sequence-spatial inverse reshaping based on target domain manifold induction To address the issue that directly reshaping the one-dimensional sequence output by DINOv3 leads to spatial neighborhood fragmentation and fails to adapt to the spatial distribution characteristics of remote sensing from a top-down perspective, this embodiment defines reshaping as a spatial collapse and repositioning process of a high-dimensional semantic manifold under the prior guidance of the target domain structure. Specifically, it includes three cascaded sub-steps, targeting sets... Each purified sequence in the process (collectively referred to here as the input sequence) ), and execute the following cascading operations in sequence: First, a target domain structure manifold embedding is constructed. Since the feature sequences output by existing large-scale visual models lack perception of specific spatial structures in remote sensing images (such as the continuity of road grids and the regular arrangement of buildings), this embodiment introduces a learnable global parameter tensor, defined as the target domain manifold embedding parameter. This is used to implicitly encode the absolute coordinates and relative distance priors of ground features. The spatial priors are injected into the input sequence using the following manifold entanglement formula: In the formula, This represents the inductive strength factor, used to adjust the degree to which prior knowledge interferes with the original features. Represents a spatial potential energy sequence. Indicates normalization, Represents the hyperbolic tangent function. This represents the input sequence.
[0034] By embedding the target domain structure manifold, each abstract semantic token is endowed with the spatial position potential unique to the target domain, which can prevent misalignment during reshaping.
[0035] Secondly, based on the minimum action manifold expansion, and guided by manifold embedding, the inverse mapping of the physical dimension is performed. This dimensional reshaping process is the repositioning of semantic features in two-dimensional Euclidean space. The rearrange operation is then used to transform the sequence with spatial potential... Forced mapping back to two-dimensional Euclidean space to construct an initial two-dimensional feature map .
[0036] Because of the injection of target domain manifold embedding parameters, the initial 2D feature map has lower spatial entropy and stronger semantic continuity between adjacent patches. Despite manifold unrolling, semantic gaps may still exist at the physical boundaries of the patches due to Transformer discretization. To repair these gaps without obscuring the edges of features, this embodiment utilizes grouped convolution to simulate anisotropic diffusion processes in thermodynamics.
[0037] The channels of the initial two-dimensional feature map are divided into several groups, and each group is applied independently. Convolutional kernels perform small spatial neighborhood interactions to compute healing residuals. The healing residual is then added to the initial two-dimensional feature map to generate the healed features, thus obtaining the two-dimensional feature map. : In the formula, GeLU represents the Gaussian error linear unit, and BN represents normalization.
[0038] Grouped convolution ensures that features with different attributes (such as water ripples and road lines) are spatially smoothed independently within their respective semantic channels without interfering with each other.
[0039] Finally, dimensional projection and spatial solidification are performed, through... Convolution projects the repaired 2D feature map onto the channel dimension required by the downstream vector field convolution module, thus solidifying the features of the current layer and obtaining the solidified features. : right The above operations are performed on all four layers to obtain the solidified features of each layer, which are represented as follows: and These solidified features are then combined to ultimately obtain a two-dimensional feature pyramid with a complete spatial structure. It should be noted that in this embodiment and subsequent embodiments, layers 3, 6, 9, and 12 of the Transformer Block are used as examples for illustration.
[0040] The above embodiments can transform the semantic sequence of a large visual model that lacks spatial structure into a high-quality two-dimensional feature map with the topological structure of remote sensing objects and the distribution characteristics of the target domain, laying a solid foundation for subsequent asymmetric convolution enhancement.
[0041] This embodiment overcomes the topological bottleneck of one-dimensional sequence features in Visual Large Models (ViT) being difficult to adapt to dense prediction tasks by employing a sequence spatial inverse reshaping mechanism based on target domain manifold induction. Addressing the issues of the abstract semantic sequences output by DINOv3 lacking spatial structure and the tendency for direct reshaping to lead to neighborhood fragmentation, a dynamic pyramid is established through semantic standing wave resonance. Furthermore, the disordered token sequence is mapped to a two-dimensional multi-scale feature pyramid with topological continuity using a priori target domain structure. This design injects the spatial inductive bias of the target domain at the source of feature input, effectively solving the problem of excessively high misclassification rates for general large model features in remote sensing top-down perspective tasks, and providing a semantically dense and topologically complete manifold basis for subsequent processing.
[0042] To address the inherent limitation of existing convolution operations in distinguishing linear features (roads / rivers) from speckle noise (background / artifacts) at the physical topology level, this embodiment introduces the Helmholtz decomposition theorem from vector field theory to construct a divergence-free curl-induced spiral vector field convolution mechanism. This mechanism utilizes the conservation properties of the spiral field to anchor the feature skeleton and mathematically removes noise through a divergence-zeroing projection mechanism. Specific steps include: Perform a normalized field mapping on each shallow feature of the two-dimensional feature pyramid to generate a vector potential; The vector potential is convolved with a discrete curl operator having an antisymmetric cross topology to generate a vortex vector field, and the vortex vector field is dynamically shaped with an anisotropic induced tensor to generate a spiral characteristic field. The divergence source density of the spiral characteristic field is calculated using a discrete divergence operator, and a divergence potential energy surface is constructed based on the divergence source density. Based on the divergence potential energy surface, a purity-gated mask is calculated, and noise is stripped from the spiral feature field using the purity-gated mask to obtain a purified feature field. The purified feature field is reconstructed using the Poisson equation to obtain the purified feature vector field, and then the purified feature vector field is scalar transformed to obtain the scalar response features. The scalar response features of each shallow feature are spliced and recombined to obtain a high-frequency singularity field.
[0043] The spiral vector field convolution mechanism based on divergence-free curl induced by this embodiment can be divided into two steps: S31. Construct a divergence-free curl-induced convolutional layer The purpose of constructing a divergence-free curl-induced convolutional layer is to extract the framework of ground features with fluid continuity from the perspective of physical fields. Unlike traditional convolutional kernels that fit local geometry, this embodiment constructs a hydrodynamic operator to calculate the curl of the feature field, thereby responding only to structures with vortex characteristics (i.e., continuous extension) and being naturally immune to discrete noise.
[0044] First, extract the first three shallow layers of features from the two-dimensional feature pyramid, namely... And perform normalized field mapping. Specifically, for any channel feature map in these three layers of features, it is uniformly defined as the input feature variable. h And its physical meaning is regarded as a scalar stream function. To perform calculations in three-dimensional vector space, gauge field theory is used to... h (Right now Elevate it to a vector potential perpendicular to the image plane: In the formula, A(*) represents the vector potential. Let x and y represent the scalar stream function, and let x and y represent two orthogonal coordinate axes in a spatial rectangular coordinate system.
[0045] Then, a discrete curl convolution operator is constructed. To simulate circulation calculations in physics, the weight matrix of this operator is rigidly constrained to an antisymmetric structure; that is, the operator adopts an antisymmetric cross-vortex kernel structure, such as... Figure 2 As shown, the horizontal arrows represent the horizontal curl kernel (i.e., horizontal curl classification) of the divergence curl convolution operator. The vertical arrow indicates the vertical curl nucleus (i.e., the vertical curl component). The specific expression is as follows: Using the above operators on vector potential Perform a convolution operation to generate the initial vortex vector field: In the formula, This represents the vortex vector field, which describes the intensity of the fluid's rotation about the vertical axis. The curl field only produces a non-zero response when a continuously extending line skeleton exists, while the theoretical value of the curl of isotropic speckle noise is zero. h Indicates the input feature variables. Represents the scalar stream function. This represents the discrete divergence operator.
[0046] Then, anisotropic tensor induction of the spiral field is performed. To enable the curl field to adapt to roads of different directions and widths in remote sensing imagery, this embodiment introduces a learnable anisotropic induction tensor. The vortex field is dynamically shaped through tensor product operations to generate a spiral characteristic field. In the formula, Let J denote the spiral characteristic field, J denote the anisotropic induced tensor, and b denote the bias term.
[0047] To further enhance the continuity of the flow, a streamline smoothing term is introduced, which is mathematically expressed as a second-order derivative constraint along the streamline direction: In the formula, For the Laplace operator along the local streamline direction, These are the operator coefficients.
[0048] The spiral characteristic field after streamline smoothing It is a localized divergence-free field, which means that the characteristic energy is strictly transported along the ground skeleton (flow tube) and will not leak into the background.
[0049] S32. Divergence-zero decoupling based on Helmholtz potential theory To fundamentally eliminate speckle noise specific to the source domain (such as building shadows and cloud cover), this embodiment transforms the denoising process into an orthogonal decoupling problem of the physical potential field. First, based on the Helmholtz orthogonal decomposition assumption, a smooth vector field over any bounded domain can be uniquely decomposed into the sum of divergence-free and irrotational components. In this embodiment, the feature field is modeled as the sum of divergence-free components (corresponding to continuous ground features) and irrotational components (corresponding to discrete noise): In the formula, F Represents the characteristic field, Indicates no discrete components. Represents the irrotational component. Represents the discrete divergence operator. Represents vector potential, It represents scalar potential.
[0050] In the above modeling formula, For continuous land features such as roads / rivers, the flow lines are continuous and unbroken. Corresponding to discrete noise / clutter, it manifests as the source and sink of gradients.
[0051] Secondly, based on the above model, the divergence potential energy level is extracted. To locate noise, this embodiment utilizes the discrete divergence operator. Calculate the divergence source density by applying the spiral feature field to the input. : In the formula, This represents the partial derivative of the spiral characteristic field with respect to the coordinate x. This represents the partial derivative of the spiral characteristic field with respect to the coordinate y.
[0052] Based on this, a divergence potential surface is constructed to quantify the probability of each pixel being a noise source: In the formula, This represents the noise source probability of a pixel at coordinate point (x, y), also known as the divergence potential surface. Indicates the preset coefficient. This represents the square of the norm.
[0053] The physical significance of the divergence potential surface is that a region with a high divergence potential surface means that the feature field vector has "broken" or "piled up" (i.e., the divergence is non-zero), which is a typical physical characteristic of speckle noise, dead ends, or unstructured textures.
[0054] Secondly, divergence nullification projection and manifold reconstruction are performed. A nonlinear nullification operator is constructed based on the divergence potential surface to forcibly project the feature field back to the divergence-free space. Specifically, the purity-gated mask is first calculated. Introducing coefficients Controlling the sensitivity of inhibition: Then perform vector-level projection to strip away the active components: In the formula, This represents the purified characteristic field.
[0055] Furthermore, to repair any minor topological damage that may be caused by projection, the smoothing property of the Poisson equation is used for manifold reconstruction, resulting in the final enhanced feature vector: In the formula, Represents the purified eigenvector field. Indicates the preset coefficient. Represents the inverse Laplace operator. This represents the discrete divergence operator.
[0056] This embodiment uses the Poisson equation to reconstruct the manifold, which mathematically ensures that the divergence of the output purified eigenvector field strictly approaches zero.
[0057] Then, the L2 norm of the purified feature vector field is calculated, transforming it into a scalar response feature usable for downstream computation. In the formula, Indicates scalar response characteristics, This represents the L2 norm.
[0058] Through the above steps, any background texture that appears as a "source" or "sink" (i.e., high dispersion) is mathematically erased, leaving only the topological skeleton that appears as a "flow" (i.e., zero dispersion).
[0059] Finally, cross-module physical field mapping and parameter handover are performed. The scalar response features obtained from the processing of the first three shallow layers are spliced and recombined to form a pure high-resolution enhanced feature. Here, the recombined enhanced feature is defined as a high-frequency singularity field, which is physically represented as a collection of topological skeleton and edge texture. At the same time, the deep features that did not participate in the above processing are defined as a low-frequency potential energy field, thus realizing physical-level decoupling.
[0060] This embodiment resolves the inherent contradiction between feature extraction and background noise removal from a physical field theory perspective by using a decoupling mechanism based on divergence-free curl-induced spiral vector field convolution and Helmholtz potential field. For feature extraction, an antisymmetric cross-vortex kernel convolution is designed, using curl rather than shape to induce the feature skeleton, ensuring a response only to spiral fields with continuous flow. For denoising, the Helmholtz decomposition theorem is used to perform divergence nullification projection by calculating the divergence potential energy of the feature field, mathematically forcibly stripping away speckle noise that appears as an active field. This physical-level decoupling gives the model natural immunity to source domain texture noise, achieving zero-sample, extremely pure extraction in the unlabeled target domain.
[0061] In the above embodiments, although the deep features provided by the DINOv3 backbone network are semantically rich, they physically manifest as low-frequency potential energy fields. Directly using traditional bilinear interpolation or deconvolution for upsampling leads to energy dissipation, resulting in blurred edges. While the enhanced features obtained through divergence-zero decoupling have weaker semantics, they accurately locate high-frequency topological singularities (i.e., ground feature edges and texture abrupt change points). Therefore, this embodiment provides a manifold inverse mapping mechanism based on singularity-potential energy coupling. This mechanism models the decoding and reconstruction process as an anisotropic diffusion process of the potential energy field guided by singularities. It aims to unfold highly compressed semantic information in a high-dimensional hyperspectral manifold space and forcibly anchor ground feature boundaries using singularity features, thereby achieving high-fidelity resolution restoration. This mechanism can be divided into three sub-steps, each of which is described in detail below.
[0062] S41. Hyperspectral manifold expansion and potential field initialization The purpose of this step is to construct a high-dimensional hyperspectral problem-solving manifold. Traditional decoders typically perform feature recovery in low-dimensional channels, leading to severe feature entanglement. This embodiment, inspired by the Whitney embedding theorem, maps low-dimensional semantic features to a high-dimensional Euclidean space, enabling the feature manifold to achieve topological decoupling in a higher dimension. Specific steps include: According to the high-frequency singularity field, the low-frequency potential energy field is aligned with the spatial coordinate system to obtain the initial potential energy field, and the high-frequency singularity field and the initial potential energy field are coupled to obtain the coupled field tensor. Hyperspectral manifold dilation is performed on the coupled field tensor to obtain an initial high-dimensional hyperspectral manifold; The initial high-dimensional hyperspectral manifold is subjected to manifold statistical normalization and Gaussian error linear mapping to obtain the high-dimensional hyperspectral manifold.
[0063] In this embodiment, firstly, the low-frequency potential field representing the overall semantic category of ground features is... High-frequency singularity fields characterizing the physical boundaries and continuous skeleton of ground features Alignment and coupling are performed. Specifically, to inject potential energy into the singularity space, bilinear interpolation is used to map the low-resolution potential energy field back to the original resolution coordinate system with the same singularity field, thus obtaining the upsampled initial potential energy field. .
[0064] Secondly, the initial potential field and the high-frequency singularity field are concatted along the channel dimension to obtain the coupled field tensor. : The spliced coupled field tensor contains both macroscopic semantic potential energy and microscopic structural singularities, providing complete initial boundary conditions for subsequent manifold evolution.
[0065] Then, hyperspectral manifold embedding mapping is performed. To provide sufficient degrees of freedom for feature evolution, this embodiment performs manifold dilation on the coupled field tensor and defines the manifold embedding matrix. And introduce manifold bias terms ,use Convolution is By projecting onto a higher-dimensional tangent space, an initial higher-dimensional hyperspectral manifold is obtained. : Finally, manifold curvature correction and activation are performed, sequentially applying batch normalization and GELU activation. Specifically, manifold statistical normalization is used to readjust the feature distribution to place it within the stable region of the manifold, and Gaussian error linear mapping is applied to the normalized manifold. The GELU function is then used to introduce smooth nonlinearity, making the manifold surface smoother. The high-dimensional hyperspectral manifold obtained through these steps is then described. They are no longer simple image pixels, but data points in a high-dimensional flow form, possessing extremely high decoupling potential.
[0066] Then, based on the singularity potential coupling of the above steps, manifold inverse mapping decoding is performed, and its second and third sub-steps specifically include: The manifold spatial gradient field is calculated for the high-dimensional hyperspectral manifold to obtain the manifold gradient vector field; Based on the high-frequency singularity field, a singularity gate is generated, and based on the singularity gate, anisotropic impact evolution is performed on the manifold gradient vector field to obtain evolution characteristics. The evolutionary features are orthogonally projected and collapsed to obtain low-dimensional features, and the low-dimensional features are added to the initial potential field to obtain the decoded features.
[0067] The last two sub-steps included in this embodiment will be described in detail below: S42, Singularity-Driven Anisotropic Gradient Flow Evolution Traditional deep convolution typically behaves as an isotropic low-pass filter, which, while eliminating noise, inevitably smooths out the edges of ground features (i.e., the thermal diffusion effect). To overcome this thermodynamic entropy increase process, this step models feature reconstruction as a controlled shock wave evolution process under singularity control.
[0068] First, the gradient field of the manifold tangent space is calculated. Specifically, on the high-dimensional hyperspectral manifold mentioned in the previous steps, the local rate of change at each feature point is calculated. To capture the fine geometric structure, this embodiment defines a manifold gradient operator, which consists of a set of orthogonal Sobel differential kernels. and The structure utilizes the manifold gradient operator to calculate along the curves. shaft and The partial derivatives of the axes form the gradient vector field of the manifold. : The gradient vector field of this manifold indicates the steepest descent direction of the characteristic manifold surface, and its magnitude represents the sharpness of the local feature edge.
[0069] Secondly, a singularity modulation tensor is constructed. Specifically, it employs... Convolution performs channel adaptation on the high-frequency singularity field, and the Sigmoid activation function is used to normalize the convolutional high-frequency singularity field into a probability distribution, thus constructing a singularity gate. The physical significance of this singularity gating lies in the fact that at the singularity (edge region)... This allows for dramatic evolution of gradient flow; in flat areas (within the region). To maintain feature smoothness and prevent noise amplification.
[0070] Finally, anisotropic shock evolution is performed, with the following specific steps: The singularity gate is nonlinearly combined with the manifold gradient vector field to obtain the evolution increment; The evolutionary features are obtained by adding the high-dimensional hyperspectral manifold to the evolutionary increment.
[0071] In this embodiment, the evolution increment is defined. Approximating it as a nonlinear combination of gradient fields under singularity-gated control, the sharpening residual is directly learned: In the formula, This is a convolutional layer that learns how to combine gradient components to form edge enhancements. A learnable evolutionary step size factor. It is the product of Hadama.
[0072] Then, the evolutionary increment is added to the high-dimensional hyperspectral manifold to obtain the evolutionary features. : This step is mathematically equivalent to performing a controlled "reverse heat conduction process," limited only to... At the specified singularity location, the feature value difference is forcibly amplified at the edge of the ground object (creating a shock wave), thereby eliminating the blur caused by upsampling. In non-edge areas, the evolution term is zero, avoiding the problem of excessive noise enhancement common in traditional sharpening algorithms.
[0073] S43, Potential subspace collapse and residual injection This step aims to safely map the high-dimensional manifold features of anisotropic evolution output from step S32 back to the low-dimensional target semantic space while maintaining low-frequency semantic conservation. To avoid information loss (especially small edges) during the dimensionality reduction process, this embodiment introduces an orthogonal subspace projection and potential energy conservation residual mechanism.
[0074] First, orthogonal projection collapse of the hyperspectral manifold is performed. The evolutionary features output from the above steps are in a high-dimensional hyperspectral space, containing rich edge details. To transform them back into the compact semantic representation required by downstream tasks, manifold collapse needs to be performed. This embodiment utilizes... Convolution simulates orthogonal projection, projecting a high-dimensional manifold onto the desired low-dimensional subspace of the target, thus obtaining low-dimensional features. : This process is equivalent to imprinting the most critical high-resolution topological structure from a complex high-dimensional geometry.
[0075] Secondly, a potential energy conservation residual injection is performed. In deep neural networks, simple nonlinear transformations can easily lead to gradient vanishing, meaning that deep semantic information gradually decays during backpropagation. To prevent the loss of deep semantic information during mapping, this embodiment introduces a potential energy conservation constraint. The initial potential energy field generated in step S31 is injected into the low-dimensional features to construct the final decoded features with high resolution. In the formula, Indicates decoding features, This represents the channel-aligned linear transformation operator. This represents the initial potential energy field.
[0076] In the actual implementation, if the input and output dimensions are the same, addition is performed directly; if they are not the same (usually occurring in the channel adjustment layer), a linear transformation is used. Alignment dimension.
[0077] The physical essence of the manifold inverse mapping mechanism based on singularity potential coupling provided in this embodiment lies in the fact that this mechanism does not require learning an image from scratch, but rather learns the high-frequency sharpening residual relative to the original low-frequency potential. This greatly reduces the optimization difficulty and ensures the smoothness of the main ground features and the sharpness of the edges.
[0078] This embodiment overcomes the challenge of recovering high-frequency spatial details from low-frequency semantic potential energy through a manifold inverse mapping decoding architecture based on singularity potential energy coupling. Addressing the shortcomings of traditional decoders, such as edge jaggedness and energy dissipation caused by simple upsampling, this mechanism models the decoding process as a dynamic evolution in a high-dimensional hyperspectral space. By coupling the low-frequency potential energy field with the high-frequency singularity field, singularity gating guides anisotropic gradient flow to sculpt the potential energy surface. This point-to-surface evolution strategy allows the model to automatically grow sharp ground object edges while restoring resolution, solving the detail reconstruction problem caused by the high semantic density and low spatial density of traditional large-scale visual model backbone networks.
[0079] Finally, an inference-end-domain adaptation strategy is adopted, which uses full probability tensor accumulation to adaptively classify unlabeled local remote sensing images of arbitrary size, achieving the final land cover migration classification and recognition. The specific steps include: Linear classification is performed on the decoded features of each local image to obtain the predicted probability tensor of each local image; According to the coordinate region of the remote sensing image, each predicted probability tensor is added together and counted to obtain the cumulative probability tensor and the cumulative count. Based on the accumulated probability tensor and the accumulated count, a pixel-by-pixel probability averaging calculation is performed, and a full-map pixel-level land feature classification mask is generated according to the maximum probability principle. Based on the pixel-level land feature classification mask of the entire image, the land feature classification and recognition results are obtained.
[0080] This embodiment employs block-based prediction, inputting the decoded features corresponding to each local image into a linear classifier head. The linear classifier head then maps the decoded features into a final pixel-level class prediction probability tensor. In the formula, This represents the prediction probability tensor. Represents a linear classification head. This represents the activation function.
[0081] Then, a smooth stitching and final decision based on probability accumulation is adopted. Specifically, in order to eliminate the gaps at the sliding window stitching and suppress local prediction noise caused by domain offset, this embodiment first constructs two blank tensors with the same spatial size as the original remote sensing image: the accumulated probability tensor. and cumulative count tensor Then, an incremental update is performed. For ease of description, assume that each sliding window position in the set of coordinates at the top left corner represents (a, b), and the sliding window size is... The cropped portion of the image is represented as The predicted probability tensor corresponding to this local image is represented as: Using the Bayesian averaging approach, for each local image, the predicted probability tensor is summed and added to the coordinate region of the accumulated probability tensor, while simultaneously increasing the accumulated count tensor. Since the sliding window has a step size of 50%, the pixels in the central region are predicted 4 times. This takes advantage of the high confidence in the central region to automatically smooth and correct the low confidence uncertainty in the edge region.
[0082] Finally, normalization and decision-making are performed. After traversing the entire image and accumulating the results, pixel-by-pixel probability averaging is calculated, and the final pixel-level land cover classification mask for the entire image is generated based on the maximum probability principle. : In the formula, Argmax(*) represents the maximum value index function. By using the full-map pixel-level land feature classification mask, the corresponding land feature classification and recognition results can be obtained.
[0083] Furthermore, this embodiment employs a robust training strategy based on differentiated learning rates for the aforementioned static network. Before model training, the dataset is prepared by dividing the source domain dataset (such as the Globe230k dataset) into training and validation sets at a 9:1 ratio. During the training phase, conventional geometric transformation strategies such as random horizontal and vertical flipping are applied to expand the data samples. Since the target domain (unlabeled local images) lacks ground truth labels, the model can only be trained under supervision on the source domain. If a uniform learning rate is directly used for full parameter fine-tuning, the powerful DINOv3 backbone network is prone to overfitting to the specific texture of the source domain, causing its general semantic representation ability to collapse. Therefore, this embodiment adopts a hierarchical decoupling optimization strategy, dividing all network parameters involved in the above steps into two groups and applying differentiated update rules to achieve the best balance between preserving general knowledge and learning specific tasks.
[0084] Specifically, to implement differentiated training, the entire model parameter set is logically decoupled into two parts: a general knowledge set and a task adaptation set. The general knowledge set corresponds to the DINOv3 backbone network parameters, which contains visual priors obtained through pre-training on massive amounts of general data. The task adaptation set corresponds to the spiral vector field convolution parameters, manifold inverse mapping parameters, and classification head parameters. This part is responsible for mapping general features to specific land cover segmentation spaces.
[0085] To balance pixel-level classification accuracy with the morphological integrity of small targets (such as narrow roads), this embodiment defines a composite loss function, which is a weighted combination of cross-entropy loss and Dice loss. The calculated total loss gradient is used to update the two sets of parameters through backpropagation. The core of this embodiment lies in applying different orders of magnitude of learning rates to the two sets of parameters; for the general knowledge set, an extremely low learning rate (e.g., ...) is used. To ensure that the weights of DINOv3 undergo only minor perturbations, maintaining their robustness to different domains (illumination, sensor distortion); for the task-adapted group, a higher base learning rate (e.g., To achieve fast convergence, we learn how to extract divergence-free skeletons from general features and perform manifold reconstruction.
[0086] This embodiment provides a remote sensing image feature migration recognition method. Through a sequence spatial inverse reshaping mechanism based on target domain manifold induction, a spatial induction bias of the target domain is injected into the source of the feature input. This effectively solves the problem of high misclassification rate of general large model features in remote sensing overhead view tasks, providing a semantically dense and topologically complete manifold basis for subsequent processing. By using a divergence- and curl-induced spiral vector field convolution and Helmholtz potential field decoupling mechanism, the model possesses natural immunity to source domain texture noise, achieving zero-sample, extremely pure extraction in the unlabeled target domain. Through a manifold inverse mapping decoding architecture based on singularity potential coupling, the model can automatically grow sharp feature edges while restoring resolution, thus achieving accurate detail reconstruction.
[0087] Please see Figure 3 Based on the same inventive concept, the second embodiment of the present invention proposes a remote sensing image land cover migration identification system, comprising: The image segmentation module 10 is used to acquire remote sensing images of the area to be tested and to divide the remote sensing images into several local images using an overlapping sliding window mechanism. The feature extraction module 20 is used to perform multi-level feature extraction on each local image using a feature extraction model, and to perform noise filtering and dimension reshaping on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The feature decoupling module 30 is used to take the deep features of the two-dimensional feature pyramid as a low-frequency potential energy field, and to perform spiral vector field induction and potential field decoupling enhancement on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The cascaded reconstruction module 40 is used to couple and embed the high-frequency singularity field and the low-frequency potential energy field into a hyperspectral manifold to obtain a high-dimensional hyperspectral manifold, and to perform singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image. The sliding window inference module 50 is used to perform linear classification on the decoded features of each local image, and to accumulate and normalize the predicted tensor probabilities obtained from the classification to obtain the ground feature classification and recognition results.
[0088] The technical features and effects of the remote sensing image ground feature migration identification system proposed in this embodiment are the same as those of the method proposed in this embodiment, and will not be repeated here. Each module in the above-mentioned remote sensing image ground feature migration identification system can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0089] Furthermore, embodiments of the present invention also propose a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0090] Please see Figure 4 The diagram illustrates the internal structure of a computer device in one embodiment. This computer device can specifically be a terminal or a server. The computer device includes a processor, memory, network interface, display, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a remote sensing image feature migration recognition method. The display screen of the computer device can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device casing, or an external keyboard, touchpad, or mouse.
[0091] Those skilled in the art will understand that Figure 4The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computing devices may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.
[0092] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0093] In summary, the present invention proposes a method, system, device, and storage medium for remote sensing image feature migration identification. The method acquires remote sensing images of the area to be measured and divides these images into several local images using an overlapping sliding window mechanism. A feature extraction model is used to extract multi-level features from each local image, and noise filtering and dimensionality reshaping are applied to the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The deep features of the two-dimensional feature pyramid are used as a low-frequency potential energy field. Spiral vector field induction and potential field decoupling enhancement are applied to each shallow feature layer of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The high-frequency singularity field and the low-frequency potential field are coupled and hyperspectral manifold embedded mapping to obtain a high-dimensional hyperspectral manifold. Singularity-driven anisotropic gradient flow evolution and inverse mapping decoding are then performed on the high-dimensional hyperspectral manifold to obtain the decoded features of each local image. The decoded features of each local image are linearly classified, and the predicted tensor probabilities obtained from the classification are accumulated and normalized to obtain the ground feature classification and recognition results. This invention, through manifold topological induction, vector field theory physical constraints, and dynamic evolution mechanisms, can achieve physical-level decoupling and high-fidelity reconstruction of the ground feature skeleton under zero-sample conditions, thereby achieving accurate and efficient remote sensing ground feature migration recognition.
[0094] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0095] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A method for identifying the migration of ground features in remote sensing images, characterized in that, include: Acquire remote sensing images of the area to be measured, and divide the remote sensing images into several local images using an overlapping sliding window mechanism; A feature extraction model is used to extract multi-level features from each local image, and noise filtering and dimension reshaping are performed on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The deep features of the two-dimensional feature pyramid are used as a low-frequency potential energy field, and spiral vector field induction and potential field decoupling enhancement are performed on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The high-frequency singularity field and the low-frequency potential energy field are coupled and hyperspectral manifold embedded mapping is performed to obtain a high-dimensional hyperspectral manifold. The high-dimensional hyperspectral manifold is then subjected to singularity-driven anisotropic gradient flow evolution and inverse mapping decoding to obtain the decoding features of each local image. Linear classification is performed on the decoded features of each local image, and the predicted tensor probabilities obtained from the classification are accumulated and normalized to obtain the ground feature classification and recognition results.
2. The remote sensing image feature migration identification method according to claim 1, characterized in that, The steps of using a feature extraction model to extract multi-level features from each local image, and then filtering noise and reshaping the extracted multi-level features to obtain a two-dimensional feature pyramid include: The standardized local image is input into a feature extraction model built on the DINOv3 network to perform multi-level feature extraction, resulting in multi-level features, which include one deep feature and several shallow features. A semantic standing wave resonance mechanism is adopted to calculate the resonance coefficient between each shallow feature and the deep feature, and to perform semantic modulation on the shallow feature based on the resonance coefficient to obtain the semantic resonance feature corresponding to each shallow feature. A layer-by-layer differential entropy screening mechanism is used to filter noise from each semantic resonance feature, resulting in multi-level purified features. For each layer of purification features in the multi-level purification features, target domain structure manifold embedding and physical dimension inverse mapping are performed to obtain an initial two-dimensional feature map; The initial two-dimensional feature map is semantically repaired based on the anisotropic topological healing mechanism to obtain a two-dimensional feature map. The two-dimensional feature map is then dimensionally projected to obtain the solidified features corresponding to each layer of purified features. The solidified features of each layer are then assembled to obtain a two-dimensional feature pyramid.
3. The remote sensing image feature migration identification method according to claim 1, characterized in that, The step of inducing a spiral vector field and decoupling the potential field to enhance each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field includes: Perform a normalized field mapping on each shallow feature of the two-dimensional feature pyramid to generate a vector potential; The vector potential is convolved with a discrete curl operator having an antisymmetric cross topology to generate a vortex vector field, and the vortex vector field is dynamically shaped with an anisotropic induced tensor to generate a spiral characteristic field. The divergence source density of the spiral characteristic field is calculated using a discrete divergence operator, and a divergence potential energy surface is constructed based on the divergence source density. Based on the divergence potential energy surface, a purity-gated mask is calculated, and noise is stripped from the spiral feature field using the purity-gated mask to obtain a purified feature field. The purified feature field is reconstructed using the Poisson equation to obtain the purified feature vector field, and then the purified feature vector field is scalar transformed to obtain the scalar response features. The scalar response features of each shallow feature are spliced and recombined to obtain a high-frequency singularity field.
4. The remote sensing image feature migration identification method according to claim 1, characterized in that, The step of coupling the high-frequency singularity field and the low-frequency potential energy field and embedding the hyperspectral manifold to obtain a high-dimensional hyperspectral manifold includes: According to the high-frequency singularity field, the low-frequency potential energy field is aligned with the spatial coordinate system to obtain the initial potential energy field, and the high-frequency singularity field and the initial potential energy field are coupled to obtain the coupled field tensor. Hyperspectral manifold dilation is performed on the coupled field tensor to obtain an initial high-dimensional hyperspectral manifold; The initial high-dimensional hyperspectral manifold is subjected to manifold statistical normalization and Gaussian error linear mapping to obtain the high-dimensional hyperspectral manifold.
5. The remote sensing image feature migration identification method according to claim 4, characterized in that, The steps of performing singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image include: The manifold spatial gradient field is calculated for the high-dimensional hyperspectral manifold to obtain the manifold gradient vector field; Based on the high-frequency singularity field, a singularity gate is generated, and based on the singularity gate, anisotropic impact evolution is performed on the manifold gradient vector field to obtain evolution characteristics. The evolutionary features are orthogonally projected and collapsed to obtain low-dimensional features, and the low-dimensional features are added to the initial potential field to obtain the decoded features.
6. The remote sensing image feature migration identification method according to claim 5, characterized in that, The step of performing anisotropic impact evolution on the manifold gradient vector field based on the singularity gating to obtain evolution characteristics includes: The singularity gate is nonlinearly combined with the manifold gradient vector field to obtain the evolution increment; The evolutionary features are obtained by adding the high-dimensional hyperspectral manifold to the evolutionary increment.
7. The method for identifying land cover migration in remote sensing images according to claim 1, characterized in that, The steps of performing linear classification on the decoded features of each local image, and then accumulating and normalizing the predicted tensor probabilities obtained from the classification to obtain the land cover classification and recognition results include: Linear classification is performed on the decoded features of each local image to obtain the predicted probability tensor of each local image; According to the coordinate region of the remote sensing image, each predicted probability tensor is added together and counted to obtain the cumulative probability tensor and the cumulative count tensor. Based on the accumulated probability tensor and the accumulated count tensor, a pixel-by-pixel probability averaging calculation is performed, and a full-map pixel-level land feature classification mask is generated according to the maximum probability principle. Based on the pixel-level land feature classification mask of the entire image, the land feature classification and recognition results are obtained.
8. A remote sensing image feature migration recognition system, characterized in that, include: The image segmentation module is used to acquire remote sensing images of the area to be tested and to divide the remote sensing images into several local images using an overlapping sliding window mechanism. The feature extraction module is used to perform multi-level feature extraction on each local image using a feature extraction model, and to perform noise filtering and dimension reshaping on the extracted multi-level features to obtain a two-dimensional feature pyramid. The feature extraction model is constructed based on a self-supervised visual model. The feature decoupling module is used to take the deep features of the two-dimensional feature pyramid as a low-frequency potential energy field, and to perform spiral vector field induction and potential field decoupling enhancement on each shallow feature of the two-dimensional feature pyramid to obtain a high-frequency singularity field. The cascaded reconstruction module is used to couple and embed the high-frequency singularity field and the low-frequency potential energy field into a hyperspectral manifold to obtain a high-dimensional hyperspectral manifold, and to perform singularity-driven anisotropic gradient flow evolution and inverse mapping decoding on the high-dimensional hyperspectral manifold to obtain the decoding features of each local image. The sliding window inference module is used to perform linear classification on the decoded features of each local image, and to accumulate and normalize the predicted tensor probabilities obtained from the classification to obtain the ground feature classification and recognition results.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.