Multi-source remote sensing information fusion target detection and identification method based on Cartesian product
By adopting Cartesian product processing and channel attention in multi-source remote sensing information fusion, problems such as insufficient utilization of multi-source data complementarity and large parameters in the prior art are solved, and more efficient object detection performance is achieved.
Patent Information
- Application Number
- CN202510116995.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-13
AI Technical Summary
The existing multi-source remote sensing information fusion target detection method is difficult to fully utilize the complementarity of multi-source data, resulting in a decline in detection performance, and problems such as large parameters and heavy training burden.
The multi-source remote sensing information fusion method based on Cartesian product is adopted, and the multi-source information fusion method of Cartesian product processing and channel attention is realized by the multi-source fusion method of multi-source fusion between different modal features and the weighted fusion between the same modal features.
It improves the accuracy and robustness of object detection, reduces training parameters and training burden, and achieves better fusion effect.
Smart Images

Figure CN119992364A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning and information fusion, and in particular to a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product. Background Art
[0002] With the rapid development of remote sensing technology, multi-source remote sensing information fusion technology has shown great application potential in the field of target detection. The diversification of remote sensing platforms and the widespread application of advanced sensors have made it possible to obtain image data from different sensors with different physical properties. These data provide rich spatial and temporal information, which is of great significance for improving the accuracy and reliability of target detection. However, in practical applications, effectively fusing these multi-source remote sensing information to give full play to its advantages in target detection has become a technical problem that needs to be solved urgently. Most traditional target detection methods are designed for single sensor data. Directly inputting multi-source data often cannot fully utilize the complementarity of each sensor data, and may even lead to a decrease in detection performance due to data differences.
[0003] In order to solve the problem of multi-source remote sensing information fusion target detection, many methods have been proposed. Among them, data fusion, feature fusion and decision fusion are the three main technical routes. The data fusion method chooses to fuse multi-source data before feature extraction, and then detects the fused data. Although the method is simple and direct, it is difficult to effectively fuse deep features and semantic information. The feature fusion method fuses the features of different information sources at one or more stages of feature extraction, often involving complex attention mechanisms or mid-level feature fusion architectures. These methods have improved the fusion effect to a certain extent, but still face problems such as large number of parameters and heavy training burden. The decision fusion method detects the inputs of different sensors separately and then fuses the detection results. Although it can take advantage of the advantages of different detectors, it is limited by the characteristics of each detector working independently, the information utilization is limited, and the fusion effect is often unsatisfactory.
[0004] Although the existing multi-source remote sensing information fusion target detection methods have solved some problems to a certain extent, there are still many shortcomings. That is, data fusion methods are difficult to process multi-source data with large differences in semantic information due to their shallow fusion methods. Although feature fusion methods can deeply fuse features, the complex attention mechanism and large number of parameter requirements make these methods face training difficulties and performance bottlenecks in practical applications. Decision fusion methods are difficult to achieve ideal fusion effects due to limited information utilization. Therefore, developing a new technology that can fully utilize the complementarity of multi-source remote sensing information and avoid the defects of existing methods has become a key issue that needs to be urgently solved in the field of multi-source remote sensing information fusion target detection. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product.
[0006] The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for target detection and recognition based on multi-source remote sensing information fusion based on Cartesian product, comprising:
[0008] Acquire remote sensing images of multiple modalities obtained by different sensors observing the same area;
[0009] Perform image registration on remote sensing images to obtain multiple remote sensing registered images;
[0010] Multiple remote sensing registered images are input into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
[0011] Optionally, the pre-trained remote sensing fusion information detection model is an improved YOLOX detector, and the improved YOLOX detector includes: a plurality of DarkNet branches equal to the number of remote sensing image modalities, a plurality of Fusion feature fusion modules, a Neck module, and a Head layer;
[0012] Multiple DarkNet branches are connected in series with the corresponding Fusion feature fusion modules according to the dimension; all Fusion feature fusion modules are connected in series with the Neck module and the Head layer in turn.
[0013] Optionally, the processing process of the Fusion feature fusion module includes:
[0014] Multiple parallel branch convolution modules are set in each Fusion feature fusion module. The branch convolution module is used to perform intra-modal feature fusion and channel feature transformation on the DarkNet branch features in the corresponding dimension to obtain the DarkNet branch features; the DarkNet branch features are the output features of the DarkNet branch in the preset dimension;
[0015] Each branch convolution module is connected in series with the global maximum pooling module and the global average pooling module, and the global maximum pooling module and the global average pooling module are connected in parallel. The global maximum pooling module and the global average pooling module are used to perform global maximum pooling and global average pooling on the DarkNet branch features, respectively, to obtain the branch global features and branch average features accordingly;
[0016] All global maximum pooling modules and all global average pooling modules are connected in series with the Cartesian product processing module, which is used to perform feature splicing on the branch global features and the branch average features on each branch to obtain branch splicing features; perform augmentation processing on the branch splicing features to obtain augmented splicing features corresponding to each mode; and perform Cartesian product processing on the augmented splicing features of multiple modes to obtain a fused Cartesian product matrix;
[0017] The Cartesian product processing module is connected in series with the channel attention weighting module. The channel attention weighting module is used to compress the fused Cartesian product matrix by convolution processing to obtain a compressed vector; the compressed vector is decoded to obtain decompressed features equal to the number of modalities of the remote sensing image; the decompressed features are subjected to channel attention weighting processing with the DarkNet branch features under the corresponding modality to obtain multiple weighted features; multiple weighted features are concatenated to obtain the Fusion feature result corresponding to a single Fusion feature fusion module;
[0018] Among them, the convolution kernel size corresponding to the branch convolution module is 1×1.
[0019] Alternatively, the Cartesian product process is expressed as:
[0020]
[0021] M all represents the fused Cartesian product matrix, M fusion represents the multimodal interaction matrix, which is the result of fusing branch concatenation features of different modalities; v RGB represents the first branch splicing feature, v IR represents the second branch concatenation feature, and T represents the transposition process.
[0022] Optionally, image registration is performed on the remote sensing images to obtain a plurality of remote sensing registered images, including:
[0023] The rotation invariant feature transform (RIFT) is used to register remote sensing images and obtain multiple remote sensing registered images.
[0024] Optionally, the training process of the pre-trained remote sensing fusion information detection model includes:
[0025] Obtain remote sensing sample images of multiple modes obtained by different sensors observing the same area;
[0026] Performing image registration on remote sensing sample images to obtain multiple remote sensing registered sample images;
[0027] Input multiple remote sensing registration sample images into the initial remote sensing fusion information detection model and perform training based on a preset loss function;
[0028] The initial remote sensing fusion information detection model that meets the preset stopping conditions is used as the pre-trained remote sensing fusion information detection model; the initial remote sensing fusion information detection model has the same structure as the pre-trained remote sensing fusion information detection model; the preset stopping conditions include: the number of iterations in the training process meets the preset iteration threshold or the value of the preset loss function is continuously less than the loss threshold.
[0029] Optionally, the preset loss function is expressed as:
[0030] L all =L obj +L reg +L cls ;
[0031] Among them, L all Represents the value of the preset loss function, L obj Represents the value of confidence loss, L cls Represents the value of classification loss, L reg Represents the value of the position loss function corresponding to the detection box;
[0032]
[0033] Where y represents the true confidence value, p is the confidence value output by the pre-trained remote sensing fusion information detection model, A is the real target box in the label, B is the prediction box output by the pre-trained remote sensing fusion information detection model, E is the total bounding box, |E\(A∪B)| is the area of the total bounding box not included in the union of the target box and the prediction box A∪B, p i is the prediction probability of the i-th type of target predicted by the pre-trained remote sensing fusion information detection model, C is the total number of target categories, A∩B is the intersection of the target box and the prediction box, and the total bounding box is the bounding box of A∪B.
[0034] In a second aspect, the present invention provides a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product, and the multi-source remote sensing information fusion target detection and recognition device based on Cartesian product includes: an acquisition unit, an image registration unit and a detection and recognition unit;
[0035] The acquisition unit is used to: acquire remote sensing images of multiple modes obtained by different sensors observing the same area;
[0036] The image registration unit is used to: perform image registration on the remote sensing image to obtain a plurality of remote sensing registered images;
[0037] The detection and recognition unit is used to: input multiple remote sensing registration images into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
[0038] In the third aspect, the present invention provides a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product, including: a processor, a storage medium and a bus, the storage medium stores machine-readable instructions executable by the processor, when the multi-source remote sensing information fusion target detection and recognition device based on Cartesian product is running, the processor and the storage medium communicate through the bus, and the processor executes the machine-readable instructions to perform the steps of the multi-source remote sensing information fusion target detection and recognition method based on Cartesian product as described in the first aspect above.
[0039] The present invention provides a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product, comprising: acquiring remote sensing images of multiple modes obtained by observing the same area by different sensors; performing image registration on the remote sensing images to obtain multiple remote sensing registered images; inputting the multiple remote sensing registered images into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; the Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features. In the present invention, a multi-source information fusion method of Cartesian product and channel attention is adopted. Since Cartesian product processing fuses the features of different modalities, and the use of Cartesian product processing can greatly reduce the neural network parameters used in the fusion stage, it can almost not increase the training burden of the pre-trained remote sensing fusion information detection model when the features are well interacted, that is, it not only improves the fusion effect but also reduces the training parameters and training burden; in addition, since the multi-source information fusion method of channel attention realizes the weighted fusion between the features of the same modality in different processing stages, the final features are more comprehensive. Through the combination of the above two processing methods, the overall accuracy and robustness of target recognition are improved.
[0040] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 A flowchart of a method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion provided by an embodiment of the present invention;
[0042] Figure 2 The schematic diagram of the structure of the pre-trained remote sensing fusion information detection model is shown exemplarily;
[0043] Figure 3 The structural diagram of the Fusion feature fusion module is shown exemplarily;
[0044] Figure 4 A schematic diagram of the structure of a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product provided by an embodiment of the present invention;
[0045] Figure 5 A structural schematic diagram of a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0047] In order to improve the accuracy and robustness of target recognition, an embodiment of the present invention provides a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product. Figure 1 The present invention provides a flowchart of a method for multi-source remote sensing information fusion target detection and recognition based on Cartesian product. Figure 1 As shown, the target detection and recognition method includes:
[0048] S101. Acquire remote sensing images of multiple modes obtained by observing the same area with different sensors.
[0049] It should be noted that different sensors may be, for example, remote sensing data collected by at least two of the following devices: synthetic aperture radar, interferometric synthetic aperture radar, multispectral camera, hyperspectral camera, thermal infrared sensor, microwave scatterometer, and lidar.
[0050] S102: Perform image registration on the remote sensing images to obtain a plurality of remote sensing registered images.
[0051] Optionally, image registration is performed on the remote sensing images to obtain a plurality of remote sensing registered images, including:
[0052] The rotation invariant feature transform (RIFT) is used to register remote sensing images and obtain multiple remote sensing registered images.
[0053] S103, inputting multiple remote sensing registration images into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results.
[0054] The pre-trained remote sensing fusion information detection model is constructed based on the multi-source information fusion method of Cartesian product processing and channel attention; Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
[0055] The embodiment of the present invention provides a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product, which adopts the multi-source information fusion method of Cartesian product and channel attention. Since the Cartesian product processing fuses the features of different modalities, and the use of Cartesian product processing can greatly reduce the neural network parameters used in the fusion stage, it can be possible to almost not increase the training burden of the pre-trained remote sensing fusion information detection model when the features are well interacted, that is, it not only improves the fusion effect but also reduces the training parameters and training burden; in addition, since the multi-source information fusion method of channel attention realizes the weighted fusion between the features of the same modality in different processing stages, the final features are more comprehensive. Through the combination of the above two processing methods, the overall accuracy and robustness of target recognition are improved.
[0056] Optionally, the pre-trained remote sensing fusion information detection model is an improved YOLOX detector, and the improved YOLOX detector includes: a plurality of DarkNet branches equal to the number of remote sensing image modalities, a plurality of Fusion feature fusion modules, a Neck module, and a Head layer;
[0057] Multiple DarkNet branches are connected in series with the corresponding Fusion feature fusion modules according to the dimension; all Fusion feature fusion modules are connected in series with the Neck module and the Head layer in turn.
[0058] In order to clarify the structure of the pre-trained remote sensing fusion information detection model, this embodiment is explained by taking the sensors involving visible light cameras and infrared cameras as examples. Figure 2 The schematic diagram of the structure of the pre-trained remote sensing fusion information detection model is shown as an example. Figure 2As shown, the pre-trained remote sensing fusion information detection model includes: 2 DarkNet branches (blue and red), 3 Fusion feature fusion modules, Neck module and Head layer. It should be noted that the number of DarkNet branches is equal to the number of modalities of the remote sensing image; the blue and red DarkNet branches are used to perform feature extraction processing on the remote sensing registration images of the visible light camera and the remote sensing registration images of the infrared camera respectively. The Fusion feature fusion module includes: Fusion1, Fusion2 and Fusion3. Fusion1 is used to perform feature fusion processing on the features of the DarkNet branch in the third dimension, Fusion2 is used to perform feature fusion processing on the features of the DarkNet branch in the fourth dimension, and Fusion3 is used to perform feature fusion processing on the features of the DarkNet branch in the fifth dimension. The Neck module and the Head layer are used for feature fusion and prediction result output respectively, and finally the target detection and recognition result is obtained. In this embodiment, the target detection and recognition result may include: detection frame position, confidence and target category.
[0059] Optionally, Figure 3 The structural diagram of the Fusion feature fusion module is shown as an example. Figure 3 As shown in the figure, the processing process of the Fusion feature fusion module includes:
[0060] In each Fusion feature fusion module, multiple parallel branch convolution modules are set (corresponding to the 1*1 convolution of the red branch and the 1*1 convolution of the blue branch). The branch convolution module is used to perform intra-modal feature fusion and channel feature transformation on the DarkNet branch features in the corresponding dimension to obtain the DarkNet branch features; the DarkNet branch features are the output features of the DarkNet branch in the preset dimension;
[0061] Each branch convolution module is connected in series with the global maximum pooling module (GMP) and the global average pooling module (GAP), and the global maximum pooling module and the global average pooling module are connected in parallel. The global maximum pooling module and the global average pooling module are used to perform global maximum pooling and global average pooling on the DarkNet branch features, respectively, to obtain the branch global features and branch average features;
[0062] All global maximum pooling modules and all global average pooling modules are connected in series with the Cartesian product processing module, which is used to perform feature splicing on the branch global features and the branch average features on each branch to obtain branch splicing features; perform augmentation processing on the branch splicing features to obtain augmented splicing features corresponding to each mode; and perform Cartesian product processing on the augmented splicing features of multiple modes to obtain a fused Cartesian product matrix z (corresponding to Figure 3z) formed by the combination of purple, white, blue and orange blocks;
[0063] The Cartesian product processing module is connected in series with the channel attention weighting module. The channel attention weighting module is used to compress the fused Cartesian product matrix by convolution processing to obtain a compressed vector; the compressed vector is decoded to obtain decompressed features equal to the number of modalities of the remote sensing image; the decompressed features are subjected to channel attention weighting processing with the DarkNet branch features under the corresponding modality to obtain multiple weighted features; multiple weighted features are concatenated to obtain the Fusion feature result corresponding to a single Fusion feature fusion module (corresponding to Figure 2 The black cross in ); the convolution kernel size corresponding to the branch convolution module is 1×1.
[0064] Alternatively, the Cartesian product process is expressed as:
[0065]
[0066] M all represents the fused Cartesian product matrix, M fusion represents the multimodal interaction matrix, which is the result of fusing branch concatenation features of different modalities; v RGB represents the first branch splicing feature, v IR represents the second branch concatenation feature, and T represents the transposition process.
[0067] Optionally, the training process of the pre-trained remote sensing fusion information detection model includes: obtaining remote sensing sample images of multiple modalities obtained by different sensors observing the same area; performing image registration on the remote sensing sample images to obtain multiple remote sensing registered sample images; inputting the multiple remote sensing registered sample images into the initial remote sensing fusion information detection model, and training based on a preset loss function;
[0068] The initial remote sensing fusion information detection model that meets the preset stopping conditions is used as the pre-trained remote sensing fusion information detection model; the initial remote sensing fusion information detection model has the same structure as the pre-trained remote sensing fusion information detection model; the preset stopping conditions include: the number of iterations in the training process meets the preset iteration threshold or the value of the preset loss function is continuously less than the loss threshold.
[0069] Optionally, the preset loss function is expressed as: L all =L obj +L reg +L cls ;
[0070] Among them, L all Represents the value of the preset loss function, L obj Represents the value of confidence loss, L clsRepresents the value of classification loss, L reg Represents the value of the position loss function corresponding to the detection box;
[0071] L obj =-[ylog(p)+(1-y)log(1-p)];
[0072]
[0073]
[0074] Where y represents the true confidence value, p is the confidence value output by the pre-trained remote sensing fusion information detection model, A is the real target box in the label, B is the prediction box output by the pre-trained remote sensing fusion information detection model, E is the total bounding box, |E\(A∪B)| is the area of the total bounding box not included in the union of the target box and the prediction box A∪B, p i is the prediction probability of the i-th type of target predicted by the pre-trained remote sensing fusion information detection model, C is the total number of target categories, A∩B is the intersection of the target box and the prediction box, and the total bounding box is the bounding box of A∪B.
[0075] In order to verify the effectiveness of a multi-source remote sensing information fusion target detection and recognition method based on Cartesian product provided by an embodiment of the present invention, a simulation experiment was also carried out. In the experiment, the data set used the VEDAI data set and was compared with other existing fusion methods. The VEDAI data set is a multimodal data set for remote sensing image analysis, which contains optical images and infrared images, and is specifically used for remote sensing target detection, multi-source fusion and other research. The target categories in the data set include vehicles such as cars, pickups, trucks, and boats.
[0076] The dataset is randomly divided into training set and test set, with the sample size of 9:1. In the training stage, the learning rate is set to 0.001, the batch size is 48, the optimizer is Adam, and two GTX2080Ti are used for training. In the comparative experiment, the detection of a single optical image (RGB) and a single infrared image (IR) by the YOLOX-s network is selected as the baseline. In addition, other fusion methods for comparison include: splicing at the data layer, channel attention weighting at the data layer, splicing at the feature layer, SE channel attention weighting at the feature layer, and using the comparison algorithms SuperYOLO and YOLOFusion on the basis of the YOLOX-s network. The experimental results are shown in Table 1. In addition, it should be noted that the YOLOX network is a series of target detection algorithms based on the improved YOLO (You Only Look Once) architecture, and YOLOX-s is a specific model in the YOLOX network series, which is a lightweight model.
[0077] Table 1 Comparison of the detection results of the method of the present invention and other existing fusion methods
[0078]
[0079]
[0080] Comparing the mAP (Mean Average Precision) of the method of the present invention with other methods, it can be seen that the mAP of the method of the present invention is greater than that of all other compared methods, which indicates that the method of the present invention can achieve better detection effect, proving the effectiveness of the proposed multi-source remote sensing information fusion target detection technology.
[0081] The method provided in the embodiment of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, an intelligent mobile terminal, a server, etc., which is not limited in the embodiment of the present invention.
[0082] Based on the same inventive concept, an embodiment of the present invention also provides a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product. Figure 4 The present invention provides a schematic diagram of a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product. Figure 4 As shown, it includes: an acquisition unit 401, an image registration unit 402 and a detection and recognition unit 403;
[0083] The acquisition unit 401 is used to: acquire remote sensing images of multiple modes observed by different sensors on the same area;
[0084] The image registration unit 402 is used to: perform image registration on the remote sensing images to obtain a plurality of remote sensing registered images;
[0085] The detection and recognition unit 403 is used to: input multiple remote sensing registration images into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
[0086] Figure 5A structural diagram of a multi-source remote sensing information fusion target detection and recognition device based on Cartesian product provided in an embodiment of the present invention includes: a processor 510, a storage medium 520 and a bus 530, wherein the storage medium 520 stores machine-readable instructions executable by the processor 510, and when the multi-source remote sensing information fusion target detection and recognition device based on Cartesian product is running, the processor 510 communicates with the storage medium 520 through the bus 530, and the processor 510 executes the machine-readable instructions to execute the steps of the above method embodiment. The specific implementation method and technical effect are similar and will not be repeated here.
[0087] The storage medium may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk storage. Optionally, the storage medium may also be at least one storage device located away from the aforementioned processor.
[0088] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0089] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention.
[0090] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification.
[0091] Although the present invention is described herein in conjunction with various embodiments, in the process of implementing the claimed invention, those skilled in the art can understand and implement other changes of the above disclosed embodiments by viewing the drawings and the disclosed content. In the description of the present invention, the term "comprising" does not exclude other components or steps, "one" or "an" does not exclude multiple situations, and the meaning of "multiple" is two or more, unless otherwise clearly and specifically limited. In addition, certain measures are recorded in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0092] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. A multi-source remote sensing information fusion target detection and recognition method based on Cartesian product, characterized in that: include: Acquire remote sensing images of multiple modalities obtained by different sensors observing the same area; Performing image registration on the remote sensing images to obtain a plurality of remote sensing registered images; The multiple remote sensing registered images are input into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; the Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
2. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 1 is characterized in that: The pre-trained remote sensing fusion information detection model is an improved YOLOX detector, and the improved YOLOX detector includes: a plurality of DarkNet branches, a plurality of Fusion feature fusion modules, a Neck module and a Head layer, which are equal to the number of the remote sensing image modalities; The multiple DarkNet branches are connected in series with the corresponding Fusion feature fusion modules according to the dimensions; all the Fusion feature fusion modules are connected in series with the Neck module and the Head layer in turn.
3. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 2 is characterized in that: The processing process of the Fusion feature fusion module includes: A plurality of parallel branch convolution modules are provided in each of the Fusion feature fusion modules, and the branch convolution modules are used to perform intra-modal feature fusion and channel feature transformation on the DarkNet branch features in the corresponding dimensions to obtain DarkNet branch features; the DarkNet branch features are output features of the DarkNet branches in the preset dimensions; Each branch convolution module is connected in series with a global maximum pooling module and a global average pooling module, and the global maximum pooling module and the global average pooling module are connected in parallel. The global maximum pooling module and the global average pooling module are used to perform global maximum pooling and global average pooling on the DarkNet branch features, respectively, to obtain branch global features and branch average features; All the global maximum pooling modules and all the global average pooling modules are connected in series with the Cartesian product processing module, and the Cartesian product processing module is used to perform feature splicing on the branch global feature and the branch average feature on each branch to obtain a branch splicing feature; perform augmentation processing on the branch splicing feature to obtain an augmented splicing feature corresponding to each modality; and perform Cartesian product processing on the augmented splicing features of multiple modalities to obtain a fused Cartesian product matrix; The Cartesian product processing module is connected in series with the channel attention weighting module, and the channel attention weighting module is used to perform matrix compression processing on the fused Cartesian product matrix by convolution processing to obtain a compressed vector; decode the compressed vector to obtain decompressed features equal to the number of modalities of the remote sensing image; perform channel attention weighting processing on the decompressed features and the DarkNet branch features under the corresponding modality to obtain multiple weighted features; perform feature splicing on the multiple weighted features to obtain a Fusion feature result corresponding to a single Fusion feature fusion module; Among them, the convolution kernel size corresponding to the branch convolution module is 1×1.
4. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 1 or 3, characterized in that: The Cartesian product process is expressed as: M all represents the fused Cartesian product matrix, M fusion represents a multimodal interaction matrix, which is a result of fusing branch concatenation features of different modalities; v RGB represents the first branch splicing feature, v IR represents the second branch concatenation feature, and T represents the transposition process.
5. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 1 is characterized in that: The performing image registration on the remote sensing images to obtain a plurality of remote sensing registered images comprises: The remote sensing images are registered using rotation invariant feature transform (RIFT) to obtain the multiple remote sensing registered images.
6. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 1 is characterized in that: The training process of the pre-trained remote sensing fusion information detection model includes: Obtain remote sensing sample images of multiple modes obtained by different sensors observing the same area; Performing image registration on the remote sensing sample images to obtain a plurality of remote sensing registered sample images; Inputting the plurality of remote sensing registration sample images into an initial remote sensing fusion information detection model, and performing training based on a preset loss function; The initial remote sensing fusion information detection model that meets the preset stopping conditions is used as the pre-trained remote sensing fusion information detection model; the initial remote sensing fusion information detection model has the same structure as the pre-trained remote sensing fusion information detection model; the preset stopping conditions include: the number of iterations in the training process meets the preset iteration threshold or the value of the preset loss function is continuously less than the loss threshold.
7. The method for target detection and recognition based on Cartesian product of multi-source remote sensing information fusion according to claim 6 is characterized in that: The preset loss function is expressed as: L all =L obj +L reg +L cls ; Among them, L all Represents the value of the preset loss function, L obj Represents the value of confidence loss, L cls Represents the value of classification loss, L reg Represents the value of the position loss function corresponding to the detection box; Where y represents the true confidence value, p is the confidence value output by the pre-trained remote sensing fusion information detection model, A is the real target box in the label, B is the prediction box output by the pre-trained remote sensing fusion information detection model, E is the total bounding box, |E\(A∪B)| is the area of the total bounding box not included in the union of the target box and the prediction box A∪B, p i is the prediction probability of the i-th type of target predicted by the pre-trained remote sensing fusion information detection model, C is the total number of target categories, A∩B is the intersection of the target box and the prediction box, and the total bounding box is the bounding box of A∪B.
8. A multi-source remote sensing information fusion target detection and recognition device based on Cartesian product, characterized in that: The multi-source remote sensing information fusion target detection and recognition device based on Cartesian product includes: an acquisition unit, an image registration unit and a detection and recognition unit; The acquisition unit is used to: acquire remote sensing images of multiple modes obtained by different sensors observing the same area; The image registration unit is used to: perform image registration on the remote sensing images to obtain a plurality of remote sensing registered images; The detection and recognition unit is used to: input the multiple remote sensing registration images into a pre-trained remote sensing fusion information detection model to obtain target detection and recognition results; the pre-trained remote sensing fusion information detection model is constructed based on a multi-source information fusion method of Cartesian product processing and channel attention; the Cartesian product processing is used to perform multi-source fusion processing between different modal features, and the multi-source information fusion method of channel attention is used to perform weighted fusion processing between multiple different processing stages and the same modal features.
9. A multi-source remote sensing information fusion target detection and recognition device based on Cartesian product, characterized in that: include: A processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor. When the multi-source remote sensing information fusion target detection and recognition device based on Cartesian product is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of the multi-source remote sensing information fusion target detection and recognition method based on Cartesian product as described in any one of claims 1 to 7.