Remote Sensing Image Target Detection Method and Electronic Device Based on Hyperbolic Space Mapping

By using hyperbolic spatial mapping technology in remote sensing image object detection, the spatial geometric measurement information of high-dimensional feature data is solved, and the generalization ability and target positioning accuracy of the detection model are improved.

CN119251679BActive Publication Date: 2025-06-13BEIJING SATELLITE INFORMATION ENG RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411380020.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-06-13
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

Existing remote sensing image object detection methods based on deep learning are difficult to effectively capture the geometric information of high-dimensional space between feature map information, resulting in poor object detection results.

Method used

Using a method based on hyperbolic space mapping, the channel information of multiple first multi-level feature maps is projected into the hyperbolic space, and the geometric measurement information of the feature information in the high-dimensional hyperbolic space is obtained, and this information is used to assist in prediction in the feature decoupling stage.

Benefits of technology

It enhances the information dimension of the predictive feature map, helps the model better understand the distribution of data, learns more abstract feature representations, and improves the generalization ability of object detection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251679B_ABST
    Figure CN119251679B_ABST
Patent Text Reader

Abstract

The object detection method for remote sensing images based on hyperbolic space mapping of the present invention includes: S1, extracting the features of the remote sensing image to obtain a hierarchical feature map; S2, fusing the hierarchical feature maps of the same type through a feature fusion network to obtain a first multi-level feature map; S3, using a hyperbolic space mapping network to project the channel information of the first multi-level feature map into the hyperbolic space to obtain a second multi-level feature map; S4, splicing the first multi-level feature map and the second multi-level feature map to obtain a fused feature map; S5, constructing a feature detection head to detect the fused feature map and calculate the classification loss and the position prediction loss; S6, repeating S1 to S5 to train the object detection model for remote sensing images; S7, using the object detection model for remote sensing images obtained in S6 to detect the remote sensing image. The present invention enhances the information dimension of the prediction feature map, helps the model understand the data distribution, learns more abstract feature representations, and thus improves the generalization ability for the object detection task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of remote sensing satellite data processing, and in particular, to a method for detecting remote sensing image targets based on hyperbolic space mapping and an electronic device. Background Art

[0002] With the continuous development of remote sensing technology and the continuous improvement of satellite resolution, the quantity and resolution of remote sensing images are also increasing continuously, and the number of obtained remote sensing images has increased exponentially. Remote sensing image target detection refers to the process of automatically detecting targets of interest from remote sensing images obtained from long distances such as satellites or aerial drones. This technology has been widely applied in many fields, such as agriculture, forestry, urban planning, natural resource management, environmental monitoring, etc. Remote sensing images have problems in detection, such as wide swath width, complex imaging background, and a large number of small targets. Therefore, how to efficiently detect targets from a large number of remote sensing images has become an important research direction.

[0003] Currently, the commonly used remote sensing image target detection technologies include traditional feature engineering-based methods and deep learning-based methods. Traditional feature engineering-based methods require manual extraction of features in the image and use machine learning algorithms to classify these features. This method requires manual design of feature extraction rules and is difficult to process complex remote sensing image data. While deep learning-based methods use deep convolutional neural networks (CNNs) to automatically learn features in the image and achieve end-to-end target detection. Therefore, they have high accuracy and robustness and do not require manual feature extraction, so they are widely used in the field of remote sensing image target detection.

[0004] With the continuous development of deep learning technology, deep learning-based methods are also continuously improved and optimized. For example, techniques such as attention mechanisms and multi-task learning are introduced to improve the detection accuracy. However, most current deep learning-based remote sensing image target detectors simply decouple the information inside the feature map based on the calculation method in Euclidean space. Such strategies cannot effectively capture the high-dimensional spatial geometric metric information between the information in the feature map, and the high-dimensional spatial geometric metric information between targets (such as the structural properties and relative position relationships of global features and local features, etc.) is very important for the detection task. Therefore, it is difficult to effectively utilize the differences and complementarities between feature data, resulting in poor performance of the target remote sensing image target detection model. The above problems all hinder the existing remote sensing image target detection tasks from achieving better performance. Summary of the Invention

[0005] To solve the above technical problems existing in the prior art, the purpose of the present invention is to provide a remote sensing image target detection method and an electronic device based on hyperbolic space mapping, which can help the target remote sensing image target detection model capture the geometric measurement information of high-dimensional feature data in the high-dimensional hyperbolic space, use this information to assist prediction in the feature decoupling stage, enhance the information dimension of the predicted feature map, help the model better understand the data distribution, learn more abstract feature representations, and improve the generalization ability of the target detection task.

[0006] To achieve the above invention purpose, the present invention provides a remote sensing image target detection method based on hyperbolic space mapping, including the following steps:

[0007] Step S1: Extract the features of the input remote sensing image through the feature extraction backbone network to obtain a plurality of hierarchical feature maps corresponding one by one to the convolutional layers of the feature extraction backbone network;

[0008] Step S2: Through the feature fusion network, according to the types of the hierarchical feature maps, fuse the hierarchical feature maps of the same type to obtain the first multi-level feature map of the remote sensing image target, and the types of the hierarchical feature maps at least include deep feature maps, middle feature maps, and shallow feature maps;

[0009] Step S3: Use the hyperbolic space mapping network to project the channel information of multiple first multi-level feature maps into the hyperbolic space to obtain multiple second multi-level feature maps, and obtain the geometric measurement information of the feature information in the high-dimensional hyperbolic space;

[0010] Step S4: Concatenate the first multi-level feature map and the second multi-level feature map to obtain multiple fused feature maps;

[0011] Step S5: Construct a feature detection head to detect the fused feature map, and calculate the classification loss and the position prediction loss;

[0012] Step S6: Repeat steps S1 to S5 to train the remote sensing image target detection model, and the remote sensing image target detection model includes the feature extraction backbone network, the feature fusion network, the hyperbolic space mapping network, and the feature detection head;

[0013] Step S7: Use the remote sensing image target detection model obtained in step S6 to detect the remote sensing image.

[0014] According to a technical solution of the present invention, in step S1, it specifically includes:

[0015] Step S11: Construct the feature extraction backbone network;

[0016] Step S12: Perform preprocessing operations on the input remote sensing image; the preprocessing operations include image cropping, image flipping, projection transformation, mosaic enhancement, and / or image filling;

[0017] Step S13: Perform feature extraction operations on the preprocessed remote sensing image through the feature extraction backbone network to obtain the hierarchical feature map.

[0018] According to a technical solution of the present invention, in the step S2, it specifically includes:

[0019] Step S21: Construct the feature fusion network;

[0020] Step S22: Receive the hierarchical feature map through the feature fusion network, and perform fusion operations on the hierarchical feature maps of the same type according to the type of the hierarchical feature map to obtain the first multi-level feature map; the fusion operations include upsampling, horizontal connection, and / or dimension splicing.

[0021] According to a technical solution of the present invention, in the step S22, it specifically includes:

[0022] Step S221: Traverse all the hierarchical feature maps and make the number of channels of the hierarchical feature maps consistent;

[0023] Step S222: Perform size upsampling operations on the hierarchical feature maps with consistent number of channels. After adjusting the size of the shallow feature map to the size of the middle feature map, then perform dimension splicing or channel addition to obtain the first multi-level feature map.

[0024] According to a technical solution of the present invention, in the step S3, the hyperbolic space mapping network maps the first multi-level feature map vector through the hyperbolic space mapping algorithm, performs pixel-level data traversal on the channel information of the first multi-level feature map, extracts feature vectors with 1*1 channel number according to pixel cycle, obtains the corresponding number of feature points, and then performs hyperbolic space mapping processing on the obtained feature points to obtain the feature vectors;

[0025] The hyperbolic space mapping algorithm is Möbius transformation or Poincaré disk modeling.

[0026] According to a technical solution of the present invention, in the step S3, using Möbius transformation to project the channel information of the feature map into the hyperbolic space specifically includes:

[0027] Step S31: Perform vector representation on the first multi-level feature map. For any first multi-level feature map, define F∈R H×W×D as the first multi-level feature map from the feature fusion network, with a size of H×W and a channel number of D;

[0028] F = [x 1 , x 2 , …, x i , …, x N

[0029] Among them, [x 1 , x 2 , …, xi , …, x N represents the sequence data obtained after stretching transformation of the first multi - level feature map F, N = H×W, x i ∈R N*D represents the two - dimensional vector corresponding to the feature point at the i - th position in the first multi - level feature map, i ∈ [1, 2, …, N];

[0030] Step S32: Use the Möbius transformation method to map the channel information of multiple said first multi - level feature maps in the Euclidean distance space to the Riemann sphere space to obtain the data structure property information in the curved surface space. The overall overview of the Möbius transformation method is as follows:

[0031] Define the Möbius transformation of the extended complex plane , where C is the complex plane and ∞ is the point at infinity. For the two - dimensional vector x i = [N i , D i corresponding to the feature point at the i - th position in the first multi - level feature map F, after the Möbius transformation τ(·), it is obtained:

[0032]

[0033] Among them, S is an invertible matrix, used to represent the projection transformation parameters a, b, c, d:

[0034]

[0035] For the two - dimensional vector x i corresponding to the feature point at the i - th position in the first multi - level feature map F, the general Möbius transformation T(x i ) is expressed as follows:

[0036]

[0037] The Möbius transformation is invariant under scaling by the said projection transformation parameters. The Möbius transformation has at most two fixed points, and the fixed point γ can be obtained by solving:

[0038]

[0039] After the Möbius transformation, the feature matrix F of the similar shape of the first multi - level feature map F can be obtained​Mobius Its internal data is characterized by the feature mapping of the feature data of the first multi-level feature map F in the Riemann sphere space, that is:

[0040] F Mobius =[x Mobius,1 , x Mobius,2 ,…, x, Mobius,i ,…, x Mobius,M

[0041] where F Mobius is the feature matrix of the second multi-level feature map obtained after the transformation of the first multi-level feature map F, [x, Mobius,i represents the transformation structure corresponding to [x i , and M represents the serial number of the feature data obtained by the transformation.

[0042] According to a technical solution of the present invention, in the step S4, it specifically includes:

[0043] Step S41: Perform feature reshaping on the multiple second multi-level feature maps obtained in the step S3;

[0044] Step S42: Perform feature normalization on the multiple second multi-level feature maps obtained after the feature reshaping in the step S41, and normalize them to the range of 0 to 255;

[0045] Step S43: Use the multiple normalized second multi-level feature maps to fuse with the multiple first multi-level feature maps in the step S2 to obtain multiple fused feature maps; wherein, the fusion method includes addition operation, feature splicing operation, and / or dual fusion operation;

[0046] Step S44: Screen the multiple fused feature maps obtained in the step S43 through an attention mechanism to obtain the optimal channel information representation;

[0047] Step S45: Use the fused feature maps screened in the step S44 for transmission and send them to the feature detection head.

[0048] According to a technical solution of the present invention, in the step S41, it specifically includes:

[0049] Process the feature matrix F of the second multi-level feature map obtained after step S33 through a shape reshaping operation to normalize it within the domain of the first multi-level feature map F∈R Mobius H×W×D domain.

[0050] ​​According to a technical solution of the present invention, in the step S6, the difference between the simulated prediction result of the remote sensing image target detection model and the true annotation of the input remote sensing image is used as a judgment basis to control the training.

[0051] According to an aspect of the present invention, there is provided an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the above-mentioned remote sensing image target detection method based on hyperbolic space mapping.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] The present invention proposes a remote sensing image target detection method based on hyperbolic space mapping. When training a remote sensing image target detection model based on a convolutional neural network, the features of the input remote sensing image are extracted through a feature extraction backbone network to obtain a hierarchical feature map, thereby enhancing the model's ability to classify and locate the target of interest in the remote sensing image. Then, a feature fusion network is used to obtain the fused shallow feature map, middle feature map, and deep feature map. The remote sensing image target detection method based on hyperbolic space mapping projects the shallow feature map, middle feature map, and deep feature map generated in the feature fusion network into the hyperbolic space to obtain the geometric measurement information of the feature information in the high-dimensional hyperbolic space, which helps the target remote sensing image target detection model capture the spatial geometric measurement information of the high-dimensional feature data in the feature fusion stage, enhances the information dimension of the original prediction feature map, helps the model better understand the data distribution, learn more abstract feature representations, enhances the information dimension of the original prediction feature map, and thus improves the generalization ability of the target detection task.

[0054] Furthermore, the remote sensing image target detection model can obtain multi-scale information of the target during the detection process, obtain richer and more accurate target information, thereby enhancing the model's ability to classify and locate the target of interest in the remote sensing image, improving the accuracy of remote sensing image target positioning, and being of great significance for the detection of rotated bounding boxes in high-resolution remote sensing images. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 Schematically showing the flowchart of the remote sensing image target detection method based on hyperbolic space mapping in an embodiment of the present invention;

[0057] Figure 2 Schematically showing the structural diagram of the remote sensing image target detection model in another embodiment of the present invention;

[0058] Figure 3 Schematically showing the overall flowchart of the feature fusion target detection algorithm based on hyperbolic space mapping in another embodiment of the present invention;

[0059] Figure 4 Schematically showing the model algorithm flowchart of the remote sensing image target detection method based on hyperbolic space mapping in another embodiment of the present invention. Detailed implementation manners

[0060] The description of the implementation manners in this specification should be combined with the corresponding drawings, and the drawings should be regarded as a part of the complete specification. In the drawings, the shape or thickness of the embodiments can be enlarged, and simplified or conveniently marked. Furthermore, each part of the structure in the drawings will be described separately. It should be noted that the elements not shown or not described in words in the drawings are in the forms known to those of ordinary skill in the art.

[0061] Any reference to directions and orientations in the description of the embodiments herein is for the convenience of description only and should not be construed as any limitation on the protection scope of the present invention. The following description of the preferred embodiments involves combinations of features, which may exist independently or in combination. The present invention is not particularly limited to the preferred embodiments. The scope of the present invention is defined by the claims.

[0062] The specific embodiments of the present invention will be described in detail below with reference to the drawings of the specification.

[0063] As Figures 1 to 3 shown, the remote sensing image target detection method based on hyperbolic space mapping in this embodiment includes the following steps:

[0064] Step S1: Extract the features of the input remote sensing image through the feature extraction backbone network to obtain a plurality of hierarchical feature maps corresponding one by one to the convolutional layers of the feature extraction backbone network;

[0065] Step S1 specifically includes:

[0066] Step S11: Construct the feature extraction backbone network;

[0067] Step S12: Perform preprocessing operations on the input remote sensing image; the preprocessing operations include image cropping, image flipping, projection transformation, mosaic enhancement, and / or image filling;

[0068] Step S13: Perform feature extraction operations on the preprocessed remote sensing image through a feature extraction backbone network to obtain a hierarchical feature map.

[0069] By performing operations such as cropping and flipping on the remote sensing image, it is beneficial to enhance the robustness, universality, and generalization ability of the remote sensing image target detection model algorithm.

[0070] During the training process of the remote sensing image target detection model, the training dataset used is one of the important factors affecting the model performance. By screening and cleaning the dataset manually and algorithmically to remove data that does not meet the requirements or contains incorrect information, the interference of invalid data can be reduced, and the accuracy and reliability of the model can be improved.

[0071] Step S2: Through a feature fusion network, according to the types of hierarchical feature maps, fuse the hierarchical feature maps of the same type to obtain the first multi-level feature map of the remote sensing image target. The types of hierarchical feature maps include at least deep feature maps, middle feature maps, and shallow feature maps;

[0072] In step S2, it specifically includes:

[0073] Step S21: Construct a feature fusion network;

[0074] Step S22: Receive the hierarchical feature maps through the feature fusion network, and according to the types of hierarchical feature maps, perform fusion operations on the hierarchical feature maps of the same type to obtain the first multi-level feature map; the fusion operations include upsampling, horizontal connection, and / or dimension concatenation.

[0075] When training a remote sensing image target detection model based on a convolutional neural network, constructing a feature fusion network through a feature pyramid enables the model to obtain multi-dimensional information of the target, thereby enhancing the model's ability to classify and locate the target of interest in the remote sensing image.

[0076] Step S22 retains the traditional feature fusion method, which is convenient for subsequent use of hyperbolic space mapping to assist the feature fusion process, enabling more valuable information to be retained or enhanced. Step S22 specifically includes:

[0077] Step S221: Traverse all hierarchical feature maps and use 1*1 convolution to make the number of channels of the hierarchical feature maps consistent;

[0078] Step S222: Perform an upsampling operation on the hierarchical feature maps after the channel number is made consistent. After adjusting the size of the shallow feature maps to the size of the middle-level feature maps, perform dimensional concatenation or channel addition to obtain the first multi-level feature maps.

[0079] Step S3: Use a hyperbolic space mapping network to project the channel information of multiple first multi-level feature maps into the hyperbolic space to obtain multiple second multi-level feature maps, and obtain the geometric metric information of the feature information in the high-dimensional hyperbolic space;

[0080] In step S3, the hyperbolic space mapping network maps the first multi-level feature map vectors through the hyperbolic space mapping algorithm, traverses the pixel-level data of the channel information of the first multi-level feature maps, extracts the feature vectors with 1*1 channel numbers according to the pixel cycle to obtain the corresponding number of feature points, and then performs hyperbolic space mapping processing on the obtained feature points to obtain the feature vectors;

[0081] The hyperbolic space mapping algorithm is the Möbius transformation or the Poincaré disk modeling.

[0082] In step S3, using the Möbius transformation to project the channel information of the feature maps into the hyperbolic space specifically includes:

[0083] Step S31: Perform vector representation on the first multi-level feature maps. For any one of the first multi-level feature maps, define F∈R H×W×D as the first multi-level feature map from the feature fusion network, with a size of H×W and a channel number of D;

[0084] F = [x 1 , x 2 , …, x i , …, x N

[0085] where, [x 1 , x 2 , …, x i , …, x N represents the sequence data obtained after the first multi-level feature map F undergoes a stretching transformation, N = H×W, and x i ∈R N*D represents the two-dimensional vector corresponding to the feature point at the i-th position in the first multi-level feature map, i∈[1, 2, …, N];

[0086] Step S32: Use the Möbius transformation method to map the channel information of multiple first multi-level feature maps in the Euclidean distance space to the Riemann sphere space to obtain the data structure property information in the curved surface space. The overall overview of the Möbius transformation method is as follows:

[0087] Define the extended complex plane The Möbius transformation, where С is the complex plane and ∞ is the point at infinity. For the two-dimensional vector x corresponding to the feature point at the i-th position in the first multi-level feature map F i =[N i , D i , after the Möbius transformation τ(·), we can get:

[0088]

[0089] where S is an invertible matrix used to represent the projective transformation parameters a, b, c, d:

[0090]

[0091] For the two-dimensional vector x corresponding to the feature point at the i-th position in the first multi-level feature map F i , the general Möbius transformation τ(x i ) can be expressed as follows:

[0092]

[0093] The Möbius transformation is invariant under scaling by projective transformation parameters. The Möbius transformation has at most two fixed points (i.e., the identity mapping case), but the identity mapping is not considered in practical applications, so there is only one. Therefore, the fixed point γ can be obtained by solving:

[0094]

[0095] After the Möbius transformation, we can obtain the feature matrix F of a similar shape to the first multi-level feature map F Mobius , and its internal data representation is the feature mapping of the feature data of the first multi-level feature map F in the Riemann sphere space, that is:

[0096] F Mobius =[x Mobius,1 , x Mobius,2 , …, x ,Mobius,i , …, x Mobius,M

[0097] where F Mobius is the feature matrix of the second multi-level feature map obtained after the transformation of the first multi-level feature map F, [x ,Mobius,i represents the transformation structure corresponding to [x i , and M represents the serial number of the feature data obtained by the transformation;

[0098] Thus, by using the above method to obtain multiple second multi-level feature maps after the Möbius transformation, they can be involved in the feature fusion process to use the geometric metric representation of feature information in the high-dimensional hyperbolic space to assist model prediction. The feature fusion methods include feature stitching, feature summation, and / or skip connection.​

[0099] In this embodiment, hyperbolic space mapping is used to perform channel information mapping processing on the shallow, middle, and deep feature maps generated in the feature fusion network to obtain the geometric measurement information of the feature information in the high-dimensional hyperbolic space. Then, the obtained multiple feature maps after hyperbolic space mapping are fused with the original feature maps, helping the target remote sensing image target detection model to capture the spatial geometric measurement information of the high-dimensional feature data in the feature fusion stage, and being able to utilize the spatial geometric measurement information of the high-dimensional feature data to assist prediction in the feature decoupling stage. This enhances the information dimension of the original prediction feature map, helps the model better understand the data distribution, learn more abstract feature representations, enhances the information dimension of the original prediction feature map, and thus improves the generalization ability for the target detection task.

[0100] Step S4: Concatenate the first multi-level feature map and the second multi-level feature map to obtain multiple fused feature maps;

[0101] Step S4 specifically includes:

[0102] Step S41: Reshape the features of the multiple second multi-level feature maps obtained in step S3 to meet the subsequent fusion requirements;

[0103] Process the feature matrix F of the second multi-level feature map obtained after step S33 through a shape reshaping operation Mobius to normalize it within the domain of the first multi-level feature map F ∈ R H×W×D to obtain a feature map size that meets the subsequent fusion requirements.

[0104] Step S42: Normalize the features of the multiple second multi-level feature maps obtained after feature reshaping in step S41 to obtain a data representation within a suitable channel range, that is, normalize it to the range of 0 to 255;

[0105] Step S43: Use the normalized multiple second multi-level feature maps to fuse with the multiple first multi-level feature maps in step S2 to obtain multiple fused feature maps; among them, the fusion methods include addition operations, feature concatenation operations, and / or dual fusion operations;

[0106] Step S44: Screen the multiple fused feature maps obtained in step S43 through an attention mechanism to obtain the optimal channel information representation;

[0107] Step S45: Transmit the fused feature maps screened in step S44 and send them to the feature detection head for subsequent feature decoupling.

[0108] Step S5: Construct a feature detection head to detect the fused feature maps and calculate the classification loss and the location prediction loss.

[0109] Step S6: Repeat steps S1 to S5 to train the remote sensing image target detection model, which includes a feature extraction backbone network, a feature fusion network, a hyperbolic space mapping network, and a feature detection head.

[0110] The overall process of training the remote sensing image target detection model in this embodiment is as Figure 3 shown. First, read the remote sensing image and its true annotation information, and perform preprocessing steps on the remote sensing image. After completing the construction of the backbone feature extraction network, this embodiment adopts a feature fusion network architecture based on hyperbolic space mapping enhancement, which is different from the traditional feature fusion network. This network first receives multi-layer feature maps, and then performs feature transformation processing based on hyperbolic space mapping. Then, perform feature reshaping to adapt to the subsequent fusion size, and perform normalization processing on the feature map data. After that, generate traditional fusion feature maps, and fuse the feature maps based on hyperbolic space mapping with the traditional feature maps. Finally, the feature detection head decouples the fused feature maps for classification and location regression prediction. After completing the loss calculation, the model training is completed.

[0111] In step S6, it also includes obtaining the simulated prediction results of the remote sensing image target detection model, using the gap between the simulated prediction results and the true annotation of the input remote sensing image as the judgment basis to control the training.

[0112] During the training of the remote sensing image target detection model, human intervention is used to timely discover and correct the errors of the model, quickly identify the error points and give correct suggestions and correction measures. Human intervention can also perform data annotation and guidance. With the assistance of humans, more accurate annotation information can be provided for the model to improve the reliability and generation effect of the model.

[0113] Step S7: Use the remote sensing image target detection model obtained in step S6 to detect the remote sensing image.

[0114] In this embodiment, when training the remote sensing image target detection model based on a convolutional neural network, the feature extraction backbone network extracts the features of the input remote sensing image to obtain hierarchical feature maps, thereby improving the model's ability to classify and locate the interesting targets in the remote sensing image. Then, through the remote sensing image target detection method based on hyperbolic space mapping, the shallow feature maps, middle feature maps, and deep feature maps generated in the feature fusion network are projected into the hyperbolic space to obtain the geometric measurement information of the feature information in the high-dimensional hyperbolic space, helping the target remote sensing image target detection model to utilize the spatial geometric measurement information of the high-dimensional feature data in the feature decoupling stage to assist prediction, enhancing the information dimension of the original prediction feature map, helping the model to better understand the data distribution, learn more abstract feature representations, enhancing the information dimension of the original prediction feature map, and thus improving the generalization ability of the target detection task.

[0115] During the detection process, the remote sensing image target detection model can obtain multi-scale information of the target, obtain richer and more accurate target information, thereby enhancing the model's ability to classify and locate the target of interest in the remote sensing image, improving the accuracy of remote sensing image target location, and being of great significance for the detection of rotated bounding boxes in high-resolution remote sensing images.

[0116] In another embodiment of the present invention, a remote sensing image target detection method based on hyperbolic space mapping is provided. The specific steps are as Figure 4 shown: Step S100, obtain remote sensing image data and perform appropriate preprocessing operations, including image cropping, image flipping, projection transformation, etc.; Step S200, build a deep neural network model and extract target features through the backbone feature extraction network; Step S300, build a feature fusion network to fuse the feature maps obtained by the backbone feature extraction network; Step S400, use the hyperbolic space feature mapping algorithm to project the channel information of the multi-layer feature maps obtained by the feature fusion network into the hyperbolic space to obtain the geometric metric information of the feature information in the high-dimensional hyperbolic space; Step S500, splice the feature map obtained by the hyperbolic space mapping algorithm with the normal feature map that has not been projected to use the geometric metric information of the hyperbolic space to assist the model in prediction; Step S600, decouple the final feature map obtained after fusing the feature detection head, and calculate the classification loss and the location prediction loss; Step S700, determine whether the training is over. If so, execute Step S800. If not, execute Steps S100 to S600 again; Step S800, use the detection model obtained in Step S700 to detect the remote sensing image.

[0117] According to one aspect of the present invention, an electronic device is provided, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the above-mentioned remote sensing image target detection method based on hyperbolic space mapping.

[0118] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to this process, method, article or terminal device. Without more limitations, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or terminal device including the said element.

[0119] Finally, it should be noted that the above description is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, several improvements and refinements can still be made without departing from the principle described in the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.

Claims

1. A remote sensing image target detection method based on hyperbolic space mapping, characterized in that: The following steps are involved: Step S1, extracting features of an input remote sensing image through a feature extraction backbone network, and obtaining a plurality of hierarchical feature maps corresponding to the convolutional levels of the feature extraction backbone network; Step S2: using a feature fusion network, according to the type of the hierarchical feature map, the hierarchical feature maps of the same type are fused to obtain a first plurality of hierarchical feature maps of the remote sensing image target, wherein the types of the hierarchical feature maps include at least a deep feature map, a middle feature map, and a shallow feature map; Step S3, using a hyperbolic space mapping network to project the channel information of the plurality of the first multi-level feature maps into a hyperbolic space, to obtain a plurality of second multi-level feature maps, and to obtain geometric metric information of the feature information in a high-dimensional hyperbolic space; Step S4, concatenating the first multi-level feature map and the second multi-level feature map to obtain a plurality of fused feature maps; Step S5: construct a feature detection head, detect the fused feature map, and calculate the classification loss and the position prediction loss; Step S6, repeating steps S1 to S5 to train a remote sensing image target detection model, wherein the remote sensing image target detection model includes the feature extraction backbone network, the feature fusion network, the hyperbolic space mapping network and the feature detection head; Step S7: Detect the remote sensing image using the remote sensing image target detection model obtained in step S6.

2. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 1, characterized in that: In the step S1, it specifically includes: Step S11, constructing the feature extraction backbone network; Step S12, performing a preprocessing operation on the input remote sensing image; the preprocessing operation includes image cropping, image flipping, projection transformation, mosaic enhancement and / or image filling; Step S13: performing feature extraction operation on the preprocessed remote sensing image through the feature extraction backbone network to obtain the hierarchical feature map.

3. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 2, characterized in that: In the step S2, it specifically includes: Step S21, constructing the feature fusion network; Step S22: receiving the hierarchical feature map through the feature fusion network, and performing a fusion operation on the hierarchical feature maps of the same type according to the type of the hierarchical feature map to obtain the first multi-layer feature map; the fusion operation includes upsampling, lateral connection and / or dimensional splicing.

4. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 3, characterized in that: In the step S22, it specifically includes: Step S221, traversing all the hierarchical feature maps, and performing channel number consistency processing on the hierarchical feature maps; Step S222: perform a size upsampling operation on the hierarchical feature map after the channel number consistency processing, adjust the size of the shallow feature map to the size of the middle feature map, and then perform dimensional splicing or channel addition to obtain the first multi-level feature map.

5. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 4, characterized in that: In step S3, the hyperbolic space mapping network maps the first multi-level feature map vector by a hyperbolic space mapping algorithm, performs pixel-level data traversal on the channel information of the first multi-level feature map, extracts feature vectors of 1*1 channels according to pixel cycles, obtains a corresponding number of feature points, and then performs hyperbolic space mapping processing on the obtained feature points to obtain feature vectors; The hyperbolic space mapping algorithm is modeled as a Möbius transform or a Poincare disk.

6. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 5, characterized in that: In step S3, the channel information of the feature map is projected into the hyperbolic space using the Mobius transform, which specifically includes: Step S31: Perform vector representation on the first multi-level feature map. For any first multi-level feature map, define F∈R H×W×D is the first multi-level feature map from the feature fusion network, with a size of H×W and a number of channels of D; F=[x1,x2,…,x i ,…,x N ] Among them, [x1,x2,…,x i ,…,x N ] represents the sequence data obtained after the first multi-level feature map F is stretched, N = H × W, x i ∈R N*D Represents the two-dimensional vector corresponding to the feature point at the i-th position in the first multi-level feature map, i∈[1,2,…,N]; Step S32: Map the channel information of the plurality of the first multi-level feature maps in the Euclidean distance space to the Riemann sphere space using the Mobius transformation method to obtain data structure property information in the surface space. The overall overview of the Mobius transformation method is as follows: Define the extended complex plane The Mobius transform of , where С is the complex plane, ∞ is the point at infinity, for the two-dimensional vector x corresponding to the feature point at the i-th position in the first multi-level feature map F i =[N i ,D i ], after the Möbius transformation τ(·), we get: Where S is a reversible matrix used to represent the projection transformation parameters a, b, c, d: For the two-dimensional vector x corresponding to the feature point at the i-th position in the first multi-level feature map F i , the universal Möbius transformation τ(x i ) is expressed as follows: The Möbius transform remains unchanged by the scaling of the projection transformation parameters, and its fixed point γ can be obtained by solving: After the Möbius transformation, the feature matrix F of the similar shape of the first multi-level feature map F is obtained. Mobius , its internal data is characterized by the feature mapping of the feature data of the first multi-level feature map F in the Riemann sphere space, that is: F Mobius =[x Mobius,1 ,x Mobius,2 ,…,x Mobius,i ,…,x Mobius,M ] Among them, F Mobius is the feature matrix of the second multi-level feature map obtained by transforming the first multi-level feature map F, [x Mobius,i ] means [x i ] corresponds to the transformation structure, M represents the sequence number of the feature data obtained by the transformation.

7. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 6, characterized in that: The step S4 specifically includes: Step S41, reshaping the plurality of second multi-level feature maps obtained in step S3; Step S42, performing feature normalization on the plurality of second multi-level feature maps obtained after feature reshaping in step S41, and normalizing them to a range of 0 to 255; Step S43, using the normalized second multi-level feature maps to fuse with the first multi-level feature maps in step S2 to obtain a plurality of fused feature maps; wherein the fusion method includes an addition operation, a feature concatenation operation and / or a dual fusion operation; Step S44, filtering the multiple fused feature maps obtained in step S43 through an attention mechanism to obtain an optimal channel information representation; Step S45, using the fused feature map screened out in step S44 to transmit and send to the feature detection head.

8. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 7, characterized in that: In the step S41, it specifically includes: The feature matrix F of the second multi-level feature map obtained after step S33 is reshaped by the reshape operation Mobius Process it and normalize it to the first multi-level feature map F∈R H×W×D within the domain.

9. The method for remote sensing image target detection based on hyperbolic space mapping according to claim 1, characterized in that: In step S6, the simulation prediction result of the remote sensing image target detection model is obtained, and the difference between the simulation prediction result and the real annotation of the input remote sensing image is used as a judgment basis to control the training.

10. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device performs the remote sensing image target detection method based on hyperbolic space mapping as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Ground penetrating radar target detection method based on attention mechanism and YOLOv5

    CN116385828A

  • Remote sensing image feature fusion method based on Gaussian mixture model and electronic equipment

    CN116563680A