Cross-modal visible light and infrared multi-source remote sensing image target identification method and system
Through a cross-modal remote sensing image target recognition method and the use of the Transformer interactive attention mechanism for feature fusion, the problem of low accuracy in remote sensing image target recognition is solved, and more efficient multi-source remote sensing image target recognition is achieved.
Patent Information
- Application Number
- CN202510374866.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-09-26
AI Technical Summary
Existing remote sensing image target recognition methods have problems such as low target recognition accuracy, susceptibility to weather interference, and complex background when processing multi-source remote sensing images, especially visible light and infrared remote sensing images, making it difficult to effectively fuse information from remote sensing images of different modalities.
A cross-modal visible light and infrared multi-source remote sensing image target recognition method is adopted. Through multi-scale feature extraction, cross-modal feature fusion, multi-scale fusion feature aggregation and rotating box Gaussian modeling, the Transformer interactive attention mechanism is used for feature fusion to enhance feature representation and target recognition accuracy.
It improves the accuracy of remote sensing image target recognition, reduces the interference of redundant information, enhances the perception of targets of different sizes, and improves the robustness and accuracy of the detection regression process.
Smart Images

Figure CN120707808A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image target recognition, and in particular relates to a cross-modal visible light and infrared multi-source remote sensing image target recognition method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] With the rapid development of satellite remote sensing technology, multi-source remote sensing systems have become an important means of obtaining information about the Earth's surface. Multi-source remote sensing images have broad application value in fields such as environmental monitoring, urban planning, and resource surveys. However, due to the varying imaging mechanisms of different sensors, remote sensing images acquired from different sources vary significantly in size, resolution, and field of view, and the scenes are complex and varied. Therefore, effectively fusing multi-source remote sensing images to improve the accuracy of intelligent target recognition has become a hot topic of research.
[0004] Currently, intelligent target recognition in remote sensing images is primarily based on single-source machine learning and deep learning methods. Machine learning methods rely on manually constructing complex features, which are then fed into a classifier for target recognition. However, manually designed features have poor generalization capabilities and are difficult to adapt to complex remote sensing imagery scenarios. Classic detectors include HOG, DPM, and Viola Jones. Deep learning object detection methods can generally be categorized as two-stage detectors and single-stage detectors. Two-stage detectors generate candidate regions before classifying and regressing targets. Representative algorithms include the RCNN and Fast RCNN series. Single-stage detectors, such as the YOLO and SSD series, do not generate candidate regions and achieve end-to-end localization and classification. These deep learning methods are all based on networks built with convolutional neural networks (CNNs). By automatically learning image feature representations, they have achieved good results in remote sensing image recognition. However, CNNs have limitations in processing global information and long-range dependencies, making it difficult to fully capture the global information in remote sensing images, which has a certain impact on effective target recognition.
[0005] Existing research is largely based on single-modal remote sensing imagery. Because remote sensing images acquired through different imaging mechanisms have varying strengths and weaknesses, the detection effectiveness of algorithms using single-modal remote sensing imagery is limited when dealing with complex scenes. For example, visible light remote sensing images offer the advantages of high resolution and clear color and texture features, but are susceptible to weather interference. Clouds and fog obstructions, as well as background interference, can lead to false or missed detections of targets. Compared to visible light remote sensing images, infrared remote sensing images have lower resolution and less clear target edge textures, but are less susceptible to interference and have strong cloud and fog penetration capabilities.
[0006] Remote sensing images differ from typical natural scene images in that they feature high-altitude aerial perspectives, wide coverage, numerous targets, large variations in target size, and a high background ratio. Currently, most deep learning methods for convolutional neural networks primarily extract features from local locations and lack contextual information, which can effectively improve target detection. Although the Transformer network can effectively model global features, it cannot be directly used for remote sensing image target detection. This is because the Transformer network uses window-based multi-head self-attention to provide contextual information. However, remote sensing images are typically large in size, and encoding them individually after slicing will lose global information. Furthermore, the large number of parameters in the Transformer encoding limits the algorithm's portability and deployment on spaceborne GPU platforms. Summary of the Invention
[0007] In view of the problem that the imaging principles of visible light and infrared sensors are different, and single-modal remote sensing images have poor target recognition effect in scenarios such as weather influence, cloud, fog, rain and snow obstruction, complex background, and insufficient lighting, the present invention provides a cross-modal visible light and infrared multi-source remote sensing image target recognition method and system. It adopts the Transformer interactive attention mechanism of multi-source remote sensing data to perform cross-modal feature fusion, mines the effective complementary information of multi-modal remote sensing data, accurately expresses the intrinsic characteristics of the target, suppresses the redundant information caused by different imaging principles, and can effectively realize the intelligent recognition of targets in visible light and infrared multi-source remote sensing images.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A first aspect of the present invention provides a cross-modal visible light and infrared multi-source remote sensing image target recognition method.
[0010] In one or more embodiments, a cross-modal visible light and infrared multi-source remote sensing image target recognition method is provided, including:
[0011] Obtain visible light and infrared remote sensing images to be identified;
[0012] The trained target recognition network is used to process the visible light and infrared remote sensing images to be identified to obtain the recognition results of the remote sensing targets; wherein the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module and a target recognition result module;
[0013] The multi-scale feature extraction module is used to extract different scale features of visible light and infrared remote sensing images to be identified in parallel;
[0014] The cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modal features of the visible light and infrared remote sensing images to be identified;
[0015] The multi-scale fusion feature aggregation module is used to aggregate the vertical features of each scale;
[0016] The rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features, the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset of the predicted rotating frame vertices and the corresponding midpoints of the circumscribed horizontal frame;
[0017] The target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
[0018] A second aspect of the present invention provides a cross-modal visible light and infrared multi-source remote sensing image target recognition system.
[0019] In one or more embodiments, a cross-modal visible light and infrared multi-source remote sensing image target recognition system includes:
[0020] An image acquisition unit, which is used to acquire visible light and infrared remote sensing images to be identified;
[0021] A target recognition unit, which is used to process the visible light and infrared remote sensing images to be identified using a trained target recognition network to obtain recognition results of remote sensing targets; wherein the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module and a target recognition result module;
[0022] The multi-scale feature extraction module is used to extract different scale features of visible light and infrared remote sensing images to be identified in parallel;
[0023] The cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modal features of the visible light and infrared remote sensing images to be identified;
[0024] The multi-scale fusion feature aggregation module is used to aggregate the vertical features of each scale;
[0025] The rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features, the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset of the predicted rotating frame vertices and the corresponding midpoints of the circumscribed horizontal frame;
[0026] The target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
[0027] A third aspect of the present invention provides a computer-readable storage medium.
[0028] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the cross-modal visible light and infrared multi-source remote sensing image target recognition method as described above.
[0029] A fourth aspect of the present invention provides an electronic device.
[0030] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the cross-modal visible light and infrared multi-source remote sensing image target recognition method as described above are implemented.
[0031] Compared with the prior art, the present invention has the following beneficial effects:
[0032] (1) The cross-modal visible light and infrared multi-source remote sensing image target recognition method of the present invention constructs a twin network to extract the respective features of optical and infrared multi-source remote sensing images in parallel, establishes a Transformer cross-modal interactive attention fusion scheme for cross-modal feature fusion, enhances the effective feature fusion representation of visible light and infrared remote sensing images, and reduces the interference of redundant information.
[0033] (2) The present invention uses a multi-scale feature aggregation scheme to improve the integrity of information expression, effectively characterize shallow information texture information and deep semantic information, enhance the ability to perceive targets of different sizes, and reduce the missed detection of multimodal remote sensing targets.
[0034] (3) The present invention uses Gaussian modeling to model the rotating frame and adopts Bhattacharyya distance to calculate the two-dimensional space between the true frame and the predicted frame. By performing BDLM rotating frame calibration on the external horizontal frame of the true frame predicted by the network, the robustness of the target detection anchor point frame regression subtask can be guaranteed, and the accuracy of the target detection regression process can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0036] Figure 1 1 is a flow chart of a cross-modal visible light and infrared multi-source remote sensing image target recognition method according to an embodiment of the present invention;
[0037] Figure 2 is the modal Transformer cross attention of an embodiment of the present invention;
[0038] Figure 3 is the feature aggregation FAM of an embodiment of the present invention;
[0039] Figure 4 This is the detection box Gaussian modeling BDLM process of an embodiment of the present invention;
[0040] Figure 5 It is a schematic structural diagram of a cross-modal visible light and infrared multi-source remote sensing image target recognition system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0041] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0042] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0043] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0044] Figure 1 is a flow chart of a cross-modal visible light and infrared multi-source remote sensing image target recognition method according to an embodiment of the present invention, such as Figure 1 The cross-modal visible light and infrared multi-source remote sensing image target recognition method in this embodiment may include:
[0045] S101, obtaining visible light and infrared remote sensing images to be identified;
[0046] S102: Using the trained target recognition network to process the visible light and infrared remote sensing images to be recognized, to obtain recognition results of the remote sensing targets.
[0047] In this embodiment, the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module and a target recognition result module.
[0048] Specifically, the multi-scale feature extraction module is used to extract features of different scales of visible light and infrared remote sensing images to be identified in parallel.
[0049] In this embodiment, the multi-scale feature extraction module is two parallel multi-scale pyramid backbone networks.
[0050] Each sub-branch in the dual branch can extract the feature tensors of multiple scales of each of the two modalities horizontally by constructing pyramid networks of different scales. It can be expressed as:
[0051]
[0052] in Represents the feature tensors extracted by the convolution of the visible light and infrared branches at the i-th layer respectively. W, H, C represent the size of the feature tensor, ψ backbone (·) represents the visible light and infrared feature extraction function, which contains the visible light and infrared learnable parameters θ R and and I T ∈R W×H×C Represents the input remote sensing imagery in visible and infrared.
[0053] Specifically, the cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modality features of the visible light and infrared remote sensing images to be identified.
[0054] The different scale feature tensors extracted by the dual-branch backbone network are analyzed by using Figure 2 The cross-modal Transformer cross-attention is shown to perform vertical feature fusion. Feature tensors of three different scales are selected for fusion from the visible light and infrared dual-branch backbone sub-networks. This can be expressed as:
[0055]
[0056] in represents the fusion feature of the i-th layer, Φ fusion (·) represents the visible light and infrared feature fusion function, which contains the learnable parameter θ f . Represent the feature tensors extracted by the convolution of the visible light and infrared branches in the i-th layer respectively.
[0057] First, enter The visible light and infrared feature maps are flattened, and a set of learnable position codes are embedded into the tensor as tokens, so that different spatial information in the remote sensing image is encoded by tokens, and a set of tokens with spatial information is obtained. The calculation formula is as follows:
[0058]
[0059] Where flatten(·) is a tensor flattening function, which is used to expand the two-dimensional feature map into a one-dimensional vector along the rows and columns. The one-dimensional vector is the input feature vector of the linear projection module, Lin_weights is the projection matrix, position is the position code, T R , T T ∈R HW×C Spatial information tokens representing encoded visible and infrared remote sensing images.
[0060] Among them, the expression of the cross-modal feature fusion module for cross-modal Transformer cross-attention longitudinal feature fusion is:
[0061]
[0062] Among them F CF-T (·) represents the cross-modal Transformer cross attention fusion function, Represents the features after vertical fusion; T R , T T Spatial information tokens representing encoded visible and infrared remote sensing images.
[0063] The cross-modal Transformer cross attention fusion function is as follows:
[0064] The first step is to use the infrared mode token T T Projection to two matrices V T ,K T ∈R HW×C , calculate a set of value and key values. R Projection to matrix Q R ∈R HW×C Calculate the query value.
[0065]
[0066] in Represents the projection matrix.
[0067] In the second step, the visible light and infrared cross-modal Transformer cross-attention matrix is established through matrix dot multiplication operation. Then, the softmax function is used to activate the visible light and infrared cross-modal Transformer cross-attention fusion features to obtain the tensor Z T The formula is:
[0068]
[0069] T′ T =α·Z T WO +β·T T
[0070] in Indicates the calculation dimension operation, W O ∈R C×C Represents the projection matrix.
[0071] In the third step, the tensor is fed into the FFN(·) sub-network composed of fully connected layers to refine the final global tensor of visible light and infrared modality fusion to enhance the generalization ability, and the final Such as the formula:
[0072]
[0073] The same steps are used to obtain the visible light modal branch
[0074] The fourth step is to concatenate the cross-attention fusion tensors obtained from the visible light and infrared modalities to obtain the final fusion tensor. The formula is:
[0075]
[0076] Specifically, the multi-scale fusion feature aggregation module is used to fuse and aggregate the longitudinal features of each scale.
[0077] Among them, such as Figure 3 The expression of the multi-scale fusion feature aggregation module shown is:
[0078]
[0079] Where FAM(·) represents the feature aggregation operation; is the aggregation feature; They are scale fusion features, medium scale fusion features, and large scale fusion features.
[0080] The goal of feature aggregation is to aggregate valuable multi-level features, fully integrating spatial details and contextual features to achieve a more complete feature representation. Because shallow networks contain a large amount of object edge contour information, which is very helpful for angular regression of rotated objects, the feature transfer subnetwork uses a pyramid-like form to transfer features.
[0081] First, FAM(·) cascades each feature map F in a feature transfer subnetwork. iThe feature layers are upsampled to match the size of the adjacent feature layers, and then the two feature layers are aggregated to ensure that the spatial dimensions of the feature layers at both ends of the connection are the same. Deep semantic information is formed by horizontally connecting the feature layers of different sizes.
[0082] Then, each aggregated feature layer is further aggregated and connected, so that the cascade framework can efficiently integrate each feature layer into more representative features. i Feature conversion is a fusion of multiple layers of features. FAM(·) incorporates a feature fusion adaptation mechanism to achieve feature alignment across multiple feature maps. Specifically, upsampling is used at each layer to adjust the size of the feature layer and its corresponding aggregated feature layer. Branch features are then concatenated, and the number of channels in the concatenated features is compressed using one-dimensional convolution, ensuring that the channel dimensions of the feature layers requiring secondary aggregation are the same.
[0083] Finally, the aggregation feature layers of different sizes are aggregated twice to fully integrate the information of the shallow network, laying the foundation for the next classification and regression tasks.
[0084] Specifically, the rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features and the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset between the predicted rotating frame vertex and the midpoint of the corresponding edge of its circumscribed horizontal frame.
[0085] Specifically, the target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
[0086] The parameters required for the rotation box are: (x, y, w, h, θ). Here, x and y represent the center of the circumscribed horizontal box, w and h represent the width and height of the horizontal box, respectively, and θ represents the angle of rotation. The coordinates are defined in a two-dimensional Cartesian coordinate system rotation matrix M(θ). Using orthogonal matrix multiplication to represent the transformation of the rotation box, we can obtain a two-dimensional Gaussian distribution N(u, ∑):
[0087]
[0088]
[0089] Here, ∧ represents the diagonal matrix of the rotation matrix eigenvalues, which represents the geometric stretching relationship of the rectangular box during the rotation process. By modeling the rotated box as a two-dimensional Gaussian distribution, it can solve the problem of large positioning errors of square objects under periodic rotation angles. In addition, the discontinuity problem of bounding boxes based on IoU calculation is also solved.
[0090] In statistics, the Bhattacharyya distance measures the similarity between two discrete probability distributions. The advantage of this method is that even if the two rotating boxes overlap a lot, the distance between the distributions will be accurately calculated based on the standard deviation between them. The definition of the Bhattacharyya distance is as follows:
[0091]
[0092] Where p(x) and q(x) represent two generalized distributions. Based on the negative logarithmic form of the above formula, BDLM defines two two-dimensional Gaussian distributions: P~N P (u p ,∑ p )、G~N G (u g ,∑ g ), the Bhattacharyya distance between P and G can be derived as:
[0093]
[0094] in det is the function for calculating the determinant of the matrix. As can be seen from the above formula, the distance calculation between the two distributions is symmetrical, indicating that BDLM can symmetrically and smoothly calculate the spatial relative angle information, which also meets the application requirements of actual target detection. According to the above formula, the parameters of the predicted bounding box (u p ,w1,h1,θ) to obtain the partial derivatives and the solution expression related to the variables. This shows that the Bhattacharyya distance can control the dynamic adjustment of the gradient by adjusting the aspect ratio and rotation angle of the bounding box, and can be effectively applied to the backpropagation and position parameter update of the subsequent rotation anchor box regression. This process does not exist in the classic horizontal box detector.
[0095] However, BD(·||·) represents the distance between two distributions and cannot be used as a similarity measure with angular information between two rotated boxes. Therefore, it is necessary to establish a nonlinear relationship between the distance function f(·) and BD(·||·). Specifically, by converting BD(·||·) into a fractional form and normalizing it, a new similarity measure is obtained, called the normalized Bhattacharyya measure, as shown below:
[0096]
[0097] in τ is the offset hyperparameter, and f(·) is the nonlinear expression of the distance function. Generally, f(·) has two forms, namely, taking the logarithm or the square of it.
[0098] like Figure 4 As shown, the process of determining the target frame by the rotating frame Gaussian modeling module is as follows:
[0099] Model the rotated bounding box and obtain an ellipse represented by a two-dimensional Gaussian distribution corresponding to the rotated bounding box;
[0100] Calculate the two-dimensional Bhattacharyya distance between the true box and the rotated bounding box based on the ellipse represented by the two-dimensional Gaussian distribution of both the true box and the rotated bounding box;
[0101] Adjust the center point position, side length and angle of the rotated bounding box so that the Gaussian distribution error between the real box and the rotated bounding box is within the set range, and then determine the current rotated bounding box as the target box.
[0102] Figure 5 This is a schematic diagram of the structure of a cross-modal visible light and infrared multi-source remote sensing image target recognition system in an embodiment of the present invention. Figure 1 The cross-modal visible light and infrared multi-source remote sensing image target recognition method corresponds to Figure 5 As shown, the cross-modal visible light and infrared multi-source remote sensing image target recognition system in this embodiment may include:
[0103] An image acquisition unit 501 is used to acquire visible light and infrared remote sensing images to be identified;
[0104] A target recognition unit 502 is configured to process visible light and infrared remote sensing images to be recognized using a trained target recognition network to obtain recognition results of remote sensing targets; wherein the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module, and a target recognition result module;
[0105] The multi-scale feature extraction module is used to extract different scale features of visible light and infrared remote sensing images to be identified in parallel;
[0106] The cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modal features of the visible light and infrared remote sensing images to be identified;
[0107] The multi-scale fusion feature aggregation module is used to aggregate the vertical features of each scale;
[0108] The rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features, the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset of the predicted rotating frame vertices and the corresponding midpoints of the circumscribed horizontal frame;
[0109] The target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
[0110] Among them, the expression of the cross-modal feature fusion module for cross-modal Transformer cross-attention longitudinal feature fusion is:
[0111]
[0112] Among them F CF-T (·) represents the cross-modal Transformer cross attention fusion function, Represents the features after vertical fusion; T R , T T Spatial information tokens representing encoded visible and infrared remote sensing images.
[0113] Among them, the expression of the multi-scale fusion feature aggregation module is:
[0114]
[0115] Where FAM(·) represents the feature aggregation operation; is the aggregation feature; They are scale fusion features, medium scale fusion features, and large scale fusion features.
[0116] It should be noted here that, Figure 5 The various modules in the cross-modal visible light and infrared multi-source remote sensing image target recognition system are Figure 1 The various steps in the cross-modal visible light and infrared multi-source remote sensing image target recognition method correspond one to one, and the specific implementation process is the same, which will not be repeated here.
[0117] In one or more embodiments, an electronic device is provided that includes a central processing unit (CPU) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) or a program loaded from a storage portion into a random access memory (RAM). Various programs and data required for system operation are also stored in the RAM. The central processing unit, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.
[0118] The following components are connected to the I / O interface: an input section including a keyboard, mouse, etc.; an output section including a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section including a hard disk; and a communication section including a network interface card such as a local area network (LAN) card and a modem. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the I / O interface as needed. Removable media such as magnetic disks, optical disks, magneto-optical disks, semiconductor memories, etc. are installed in the drive as needed, so that computer programs read from them can be installed in the storage section as needed.
[0119] When the central processing unit in the electronic device of this embodiment executes the program, the following is achieved: Figure 1 The steps in the cross-modal visible light and infrared multi-source remote sensing image target recognition method are shown.
[0120] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, the computer program including a computer program for executing Figure 1 In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion and / or installed from a removable medium. When the computer program is executed by the central processing unit, the various functions defined in the apparatus of the present application are performed.
[0121] in, Figure 1 The computer program instructions corresponding to the method shown can also be stored in a computer readable memory that can guide a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0122] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0123] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A cross-modal visible light and infrared multi-source remote sensing image target recognition method, characterized in that: include: Obtain visible light and infrared remote sensing images to be identified; The trained target recognition network is used to process the visible light and infrared remote sensing images to be identified to obtain the recognition results of the remote sensing targets; wherein the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module and a target recognition result module; The multi-scale feature extraction module is used to extract different scale features of visible light and infrared remote sensing images to be identified in parallel; The cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modal features of the visible light and infrared remote sensing images to be identified; The multi-scale fusion feature aggregation module is used to aggregate the vertical features of each scale; The rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features, the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset of the predicted rotating frame vertices and the corresponding midpoints of the circumscribed horizontal frame; The target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
2. The cross-modal visible light and infrared multi-source remote sensing image target recognition method according to claim 1, characterized in that: The multi-scale feature extraction module is composed of two parallel multi-scale pyramid backbone networks.
3. The cross-modal visible light and infrared multi-source remote sensing image target recognition method according to claim 1, characterized in that: The expression for the cross-modal feature fusion module to perform cross-modal Transformer cross-attention longitudinal feature fusion is: Among them F CF-T (·) represents the cross-modal Transformer cross attention fusion function, Represents the characteristics after vertical fusion; T R , T T Spatial information tokens representing encoded visible and infrared remote sensing images.
4. The cross-modal visible light and infrared multi-source remote sensing image target recognition method according to claim 1, characterized in that: The expression of the multi-scale fusion feature aggregation module is: Where FAM(·) represents the feature aggregation operation; is the aggregation feature; They are scale fusion features, medium scale fusion features, and large scale fusion features.
5. The cross-modal visible light and infrared multi-source remote sensing image target recognition method according to claim 1, characterized in that: The process of determining the target frame by the rotating frame Gaussian modeling module is as follows: Model the rotated bounding box and obtain an ellipse represented by a two-dimensional Gaussian distribution corresponding to the rotated bounding box; Calculate the two-dimensional Bhattacharyya distance between the true box and the rotated bounding box based on the ellipse represented by the two-dimensional Gaussian distribution of both the true box and the rotated bounding box; Adjust the center point position, side length and angle of the rotated bounding box so that the Gaussian distribution error between the real box and the rotated bounding box is within the set range, and then determine the current rotated bounding box as the target box.
6. A cross-modal visible light and infrared multi-source remote sensing image target recognition system, characterized by: include: An image acquisition unit, which is used to acquire visible light and infrared remote sensing images to be identified; A target recognition unit, which is used to process the visible light and infrared remote sensing images to be identified using a trained target recognition network to obtain recognition results of remote sensing targets; wherein the target recognition network includes a multi-scale feature extraction module, a cross-modal feature fusion module, a multi-scale fusion feature aggregation module, a rotating frame Gaussian modeling module and a target recognition result module; The multi-scale feature extraction module is used to extract different scale features of visible light and infrared remote sensing images to be identified in parallel; The cross-modal feature fusion module is used to perform cross-modal Transformer cross-attention longitudinal feature fusion on the same-scale different-modal features of the visible light and infrared remote sensing images to be identified; The multi-scale fusion feature aggregation module is used to aggregate the vertical features of each scale; The rotating frame Gaussian modeling module is used to obtain the position and category of the rotating frame based on the relationship between the aggregated features, the rotating frame and the category relationship, calibrate the rotating frame through Gaussian modeling and predict the circumscribed horizontal frame of the rotating frame, and determine the target frame based on the offset of the predicted rotating frame vertices and the corresponding midpoints of the circumscribed horizontal frame; The target recognition result module is used to output the classification result of the target frame and its corresponding remote sensing target.
7. The cross-modal visible light and infrared multi-source remote sensing image target recognition system according to claim 6, characterized in that: The expression for the cross-modal feature fusion module to perform cross-modal Transformer cross-attention longitudinal feature fusion is: Among them F CF-T (·) represents the cross-modal Transformer cross attention fusion function, Represents the characteristics after vertical fusion; T R , T T Spatial information tokens representing encoded visible and infrared remote sensing images.
8. The cross-modal visible light and infrared multi-source remote sensing image target recognition system according to claim 6, characterized in that: The expression of the multi-scale fusion feature aggregation module is: Where FAM(·) represents the feature aggregation operation; is the aggregation feature; They are scale fusion features, medium scale fusion features, and large scale fusion features.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the cross-modal visible light and infrared multi-source remote sensing image target recognition method according to any one of claims 1 to 5 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the cross-modal visible light and infrared multi-source remote sensing image target recognition method according to any one of claims 1 to 5 are implemented.