Border patrol method and device based on cross-modal dynamic guidance mamba feature fusion, equipment and medium
By dynamically guiding the Mamba feature fusion network across modalities and combining visible light and infrared modal images, high accuracy and robustness of UAV target detection in complex environments are achieved. This solves the problem of insufficient global modeling capability in single-modal processing and improves the detection effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN NORMAL UNIVERSITY
- Filing Date
- 2026-06-24
- Publication Date
- 2026-07-24
AI Technical Summary
In existing UAV target detection tasks, single-modal processing is difficult to balance detection accuracy and environmental robustness, especially in complex, crowded or severely occluded scenarios. It also suffers from insufficient global modeling capabilities, low cross-modal interaction depth, and easy loss of frequency domain structural information.
A method based on cross-modal dynamic guided Mamba feature fusion is adopted. Dual-modal images are simultaneously acquired by border patrol drones. Single-modal features are extracted using a pre-set border target detection model. Spatial and frequency domain features are enhanced by combining a cross-modal dynamic guided Mamba feature fusion network. Cross-modal feature fusion is achieved through a hybrid scanning fusion network. Finally, cross-scale consistency verification is performed to improve detection accuracy and environmental robustness.
While ensuring the feasibility of edge deployment, the internal semantic consistency of the model's detection features has been enhanced, improving the detection accuracy and environmental robustness in complex and crowded scenarios.
Smart Images

Figure CN122454465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and in particular to a border inspection method, apparatus, equipment and medium based on cross-modal dynamic guided Mamba feature fusion. Background Technology
[0002] Currently, the implementation of target detection tasks on mobile platforms such as UAVs mainly relies on single-mode processing. However, single-mode processing is limited by its own characteristics and makes it difficult to balance detection accuracy and environmental robustness.
[0003] However, existing solutions for multispectral target detection mainly fall into two categories: 1) Convolutional neural network-based solutions: These have limited ability to model long-distance contextual dependencies in images. When dealing with complex, crowded, or severely occluded scenes common in urban environments, they struggle to effectively associate fragmented information of occluded targets, resulting in insufficient global modeling capabilities; 2) Transformer-based solutions: Their computational complexity is proportional to the square of the input sequence length, easily leading to insufficient global modeling capabilities due to limitations in computational overhead and memory usage, and model training convergence is difficult. Furthermore, these solutions often suffer from low cross-modal interaction depth (feature-level concatenation or weighted averaging) and easy loss of frequency domain structural information, making them unsuitable for border inspection. Summary of the Invention
[0004] In view of this, the purpose of this invention is to provide a border inspection method, apparatus, device, and medium based on cross-modal dynamic guided Mamba feature fusion, which can solve the problems existing in existing related solutions, realize deep, bidirectional cross-modal interaction, thereby enhancing the internal semantic consistency of the features detected by the model while ensuring the feasibility of edge deployment, and improving the detection accuracy and environmental robustness in complex, crowded, or severely occluded scenarios. The specific solution is as follows: Firstly, this application provides a border inspection method based on cross-modal dynamic guided Mamba feature fusion, applied to a border UAV inspection system, including: Border patrol drones simultaneously acquire dual-modal border images to identify the border image to be detected; wherein, the dual-modal includes visible light mode and infrared mode; The border image to be detected is input into the backbone network of a preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results; the preset border target detection model is located at the end of the border patrol drone; The feature extraction results are input into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial and frequency domains and obtain feature enhancement results; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling. The feature enhancement results are input into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; The feature fusion results are subjected to cross-scale consistency verification based on the preset border target detection model, and the feature fusion results are updated using the corresponding spatial confidence map to obtain the feature update results; The feature update result is input into the neck network of the preset border target detection model to complete the target detection corresponding to the border image to be detected, and the target border inspection result is obtained.
[0005] Optionally, the step of inputting the border image to be detected into the backbone network of a preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results includes: Input the border image to be detected into the backbone network of the preset border target detection model; The convolutional module, fast spatial pyramid pooling module, and C2f module in the backbone network are used to extract single-modal features from the border images to be detected for each modality, so as to obtain the feature extraction results corresponding to each modality.
[0006] Optionally, inputting the feature extraction results into the cross-modal dynamically guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial and frequency domains includes: The feature extraction results are input into a four-way Mamba network in a cross-modal dynamically guided Mamba feature fusion network; The feature extraction results are mapped and expanded by the linear embedding layer in the four-way Mamba network to complete the feature enhancement operation in the spatial domain and obtain the initial feature enhancement results corresponding to each modality. The initial feature enhancement result is input into the preset split spectrum enhancement network for frequency domain feature enhancement to obtain the output feature enhancement result.
[0007] Optionally, the step of inputting the initial feature enhancement result into the preset split spectrum enhancement network for frequency domain feature enhancement to obtain the output feature enhancement result includes: The initial feature enhancement result is subjected to a fast Fourier transform using the preset split spectrum enhancement network to obtain the spectrum mapping result; The preset split-type spectrum enhancement network is used to decompose the spectrum mapping result to determine the amplitude spectrum and phase spectrum; wherein, the amplitude spectrum includes high-frequency detail information, and the phase spectrum includes local spatial information; The amplitude spectrum is enhanced by the amplitude spectrum processing branch in the preset split spectrum enhancement network to determine the first enhancement result; the amplitude spectrum processing branch includes two cascaded convolutional blocks and activation functions. The phase spectrum is subjected to dual-scale local feature extraction through the phase spectrum processing branch in the preset split spectrum enhancement network to determine the first local feature extraction result; The phase spectrum processing branch is used to perform feature cross-fusion at different scales on the first local feature extraction result to determine the cross-fusion result. The phase spectrum processing branch is used to perform dual-scale local feature extraction on the cross-fusion result to determine the second local feature extraction result; The second local feature extraction result is fused using the phase spectrum processing branch to determine the second enhancement result; The first enhancement result and the second enhancement result are reconstructed using a complex frequency domain tensor using the preset split spectrum enhancement network to determine the tensor reconstruction result; The preset split-spectrum enhancement network is used to perform a two-dimensional inverse fast Fourier transform on the tensor reconstruction result to determine the feature mapping result; The feature mapping result and the initial feature enhancement result are multiplied element-wise according to the adaptive gating mechanism in the preset split spectrum enhancement network to determine the processed feature map. The processed feature map and the initial feature enhancement result are added element-wise according to the residual connection structure in the preset split spectrum enhancement network to determine the feature enhancement result.
[0008] Optionally, the step of inputting each of the feature enhancement results into a preset hybrid scan fusion network in the cross-modal dynamically guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results includes: The feature enhancement results are mapped in a high-dimensional space according to the linear embedding layer in the preset hybrid scanning fusion network to determine the mapping result; The mapping result is segmented along the channel dimension to obtain a segmentation result; wherein, the segmentation result includes a first feature map sequence and a second feature map sequence corresponding to the visible light mode, and a third feature map sequence and a fourth feature map sequence corresponding to the infrared mode; The first feature map sequence is flattened along four spatial directions to determine the first flattening result; The third feature map sequence is flattened along four spatial directions to determine the second flattening result; The first flattening result and the second flattening result are respectively input into a preset cross-modal fusion network to complete bidirectional cross-modal information fusion and obtain the first fusion result and the second fusion result; Based on the first fusion result, the feature map is restored to determine the first fused feature map sequence; Based on the second fusion result, the feature map is restored to determine the second fused feature map sequence; The first fused feature map sequence is multiplied element-wise using the second feature map sequence to determine the first multiplication result; The second fused feature map sequence is multiplied element-wise using the fourth feature map sequence to determine the second multiplication result; The feature fusion result is determined by using the first multiplication result, the second multiplication result, and the preset feature scanning branch.
[0009] Optionally, determining the feature fusion result using the first multiplication result, the second multiplication result, and a preset feature scanning branch includes: The visible light feature map in the feature enhancement result is globally enhanced by a first preset feature scanning branch to determine the first global enhancement feature; wherein, the first preset feature scanning branch includes a deep convolutional network, a layer normalization network, a batch normalization network, a convolutional feedforward neural network, and a state space model based on a dynamic adaptive scanning strategy. The infrared feature map in the feature enhancement result is globally enhanced by a second preset feature scanning branch to determine the second global enhancement feature; wherein, the second preset feature scanning branch has the same structure as the first preset feature scanning branch; The first global enhancement feature is fused with the first multiplication result to determine the target visible light fusion feature; The second global enhancement feature is fused with the second multiplication result to determine the target infrared fusion feature; The feature fusion result is determined using the target convolutional block, the target visible light fusion feature, and the target infrared fusion feature.
[0010] Optionally, the step of performing cross-scale consistency verification on the feature fusion result according to the preset border target detection model, and updating the feature fusion result using the corresponding spatial confidence map, includes: The feature fusion results are divided according to the feature scale to obtain deep fusion features and shallow fusion features; The deep fusion features are input into the target channel attention layer of the preset border target detection model to obtain a global semantic descriptor; By performing an upsampling operation on the global semantic descriptor, the global semantic descriptor is broadcast to the spatial dimension corresponding to the shallow fusion feature; The spatial confidence map is determined by analyzing the element-wise similarity between the shallow fusion features and the global semantic descriptor, and based on the corresponding analysis results. The shallow fusion features are weighted and modulated based on the spatial confidence map to obtain the modulated shallow fusion features; The feature update result is determined based on the deep fusion features and the modulated shallow fusion features.
[0011] Secondly, this application provides a border inspection device based on cross-modal dynamic guided Mamba feature fusion, applied to a border drone inspection system, comprising: A border image acquisition module is used to simultaneously acquire dual-modal border images via a border patrol drone to identify the border image to be detected; wherein, the dual-modal mode includes a visible light mode and an infrared mode; The image feature extraction module is used to input the border image to be detected into the backbone network of the preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results; the preset border target detection model is located on the end side of the border inspection UAV. The dual-domain feature enhancement module is used to input the feature extraction results of each feature into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial domain and frequency domain, and obtain feature enhancement results; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling. The cross-modal fusion module is used to input the feature enhancement results of each feature into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; The feature verification module is used to perform cross-scale consistency verification on the feature fusion result according to the preset border target detection model, and update the feature fusion result using the corresponding spatial confidence map to obtain the feature update result; The inspection result determination module is used to input the feature update result into the neck network in the preset border target detection model to complete the target detection corresponding to the border image to be detected and obtain the target border inspection result.
[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor is used to execute the computer program to implement the steps of the aforementioned border inspection method based on cross-modal dynamic guided Mamba feature fusion.
[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the aforementioned border inspection method based on cross-modal dynamic guided Mamba feature fusion.
[0014] As can be seen, the border drone inspection system applied in this application includes: simultaneously acquiring dual-modal border images via a border inspection drone to determine the border image to be detected; wherein the dual-modal includes a visible light mode and an infrared mode; inputting the border image to be detected into the backbone network of a preset border target detection model to complete single-modal image feature extraction, and obtaining feature extraction results; the preset border target detection model is located at the end of the border inspection drone; inputting each of the feature extraction results into a cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete spatial and frequency domain feature enhancement, and obtaining feature enhancement results; wherein the cross-modal dynamic guided Mamba feature fusion network includes a preset discrete spectrum enhancement network. The preset split-spectrum enhancement network is used for amplitude and phase decoupling feature enhancement; the feature enhancement results are input into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; the feature fusion results are subjected to cross-scale consistency verification according to the preset border target detection model, and the feature fusion results are updated using the corresponding spatial confidence map to obtain feature update results; the feature update results are input into the neck network in the preset border target detection model to complete the target detection corresponding to the border image to be detected and obtain the target border inspection results. In other words, the border drone inspection system applied in this application firstly acquires border images in visible light and infrared modes simultaneously using a border inspection drone. The obtained border images to be detected are then input into the backbone network of a pre-defined border target detection model to obtain feature extraction results. Next, the feature extraction results are input into a cross-modal dynamically guided Mamba feature fusion network within the model to perform feature enhancement in the spatial domain and amplitude and phase decoupling feature enhancement in the frequency domain, resulting in feature enhancement results. Then, the feature enhancement results are input into a pre-defined hybrid scanning fusion network to obtain feature fusion results. Finally, cross-scale consistency verification is performed on the feature fusion results, and feature updates are triggered using the corresponding spatial confidence map to obtain feature update results. The feature update results are then input into the neck network of the pre-defined border target detection model to obtain the target border inspection results. This approach solves the problems existing in related solutions, achieving deep, bidirectional cross-modal interaction. This enhances the internal semantic consistency of the detected features while ensuring the feasibility of edge deployment, improving detection accuracy and environmental robustness in complex, crowded, or severely occluded scenarios. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0016] Figure 1 A flowchart of a border inspection method based on cross-modal dynamic guided Mamba feature fusion provided for this application; Figure 2 A schematic diagram of the architecture of a pre-defined border target detection model based on cross-modal dynamic guided Mamba feature fusion provided for this application; Figure 3 This application provides a schematic diagram of a feature enhancement process based on a pre-defined split spectrum enhancement network; Figure 4 This application provides a schematic diagram of a cross-modal feature fusion process based on a preset hybrid scanning fusion network; Figure 5 A schematic diagram of a border inspection device based on cross-modal dynamic guided Mamba feature fusion provided for this application; Figure 6 This application provides a structural diagram of an electronic device. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Currently, target detection tasks on mobile platforms such as UAVs are mainly implemented using single-modal processing. However, single-modal processing, due to its inherent characteristics, struggles to balance detection accuracy and environmental robustness. Existing solutions for multispectral target detection primarily fall into two categories: 1) Convolutional neural network-based solutions: These have limited ability to model long-distance contextual dependencies in images. When dealing with complex, crowded, or severely occluded scenes common in urban environments, they struggle to effectively correlate fragmented information of occluded targets, resulting in insufficient global modeling capabilities; 2) Transformer-based solutions: Their computational complexity is proportional to the square of the input sequence length, easily leading to insufficient global modeling capabilities due to computational overhead and memory consumption, and model training convergence is difficult. Furthermore, these solutions often suffer from low cross-modal interaction depth (feature-level concatenation or weighted averaging) and the easy loss of frequency domain structural information.
[0019] To address this, this application provides a border inspection scheme based on cross-modal dynamic guided Mamba feature fusion, which can solve the problems existing in the current related schemes, realize deep and bidirectional cross-modal interaction, thereby enhancing the internal semantic consistency of the features detected by the model while ensuring the feasibility of deployment on the edge, and improving the detection accuracy and environmental robustness in complex, crowded or severely occluded scenarios.
[0020] See Figure 1 As shown, this invention discloses a border inspection method based on cross-modal dynamic guided Mamba feature fusion, applied to a border UAV inspection system, including: Step S11: Simultaneously acquire dual-modal border images using a border patrol drone to determine the border image to be detected; wherein, the dual-modal includes visible light mode and infrared mode.
[0021] In this embodiment, the visible light mode and infrared mode images of a certain border area can be simultaneously acquired by a camera device installed on a drone used for border inspection.
[0022] Step S12: Input the border image to be detected into the backbone network of the preset border target detection model to complete the single-modal image feature extraction and obtain the feature extraction result; the preset border target detection model is located on the end side of the border inspection UAV.
[0023] In this embodiment, combined with Figure 2As shown, the border images of each modality are first synchronously input into the backbone network of the preset border target detection model on the UAV for feature extraction. That is, the border image to be detected is input into the backbone network of the preset border target detection model. The convolutional module, fast spatial pyramid pooling module, and C2f module in the backbone network are used to perform single-modal feature extraction on the border image to be detected for each modality, thereby obtaining the feature extraction results corresponding to each modality. The C2f module, also known as C2-fusion, is a feature extraction module. Figure 2 'Extraction' in the backbone network; Figure 2 The 'fusion' in the backbone network is SPPF (Spatial Pyramid Pooling-Fast). This means that different images are extracted using different branches of the backbone network.
[0024] Step S13: Input the feature extraction results of each feature into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete the feature enhancement in the spatial domain and frequency domain, and obtain the feature enhancement result; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling.
[0025] In this embodiment, combined with Figure 2 As shown, after completing the single-modal feature extraction, the obtained feature extraction results and the core module of the preset border target detection model—the Multi-directional Mamba Cross-Attention Feature Fusion Module (M)—will be utilized. 2 CA-FFM, or feature enhancement based on a cross-modal dynamically guided Mamba feature fusion network, involves: inputting the feature extraction results into a four-way Mamba network within the cross-modal dynamically guided Mamba feature fusion network; performing feature mapping and expansion on the feature extraction results through linear embedding layers in the four-way Mamba network to complete the spatial domain feature enhancement operation, obtaining initial feature enhancement results for each modality; and inputting the initial feature enhancement results into a preset discrete spectrum enhancement network for frequency domain feature enhancement to obtain the output feature enhancement results. Mamba is a neural network architecture derived from the Selective State Space Model (SSM).
[0026] It is important to understand that Mamba is dynamically guided across modalities to enhance features of a single modality, providing a feature benchmark that already contains rich spatial long-range dependency information. Its output serves as the input for subsequent modules. For the input feature map, the original input features are first mapped and expanded to twice their original size through a linear embedding layer. Then, the row-direction 1D (one-dimensional) feature map is restored to a 2D (two-dimensional) feature map. Crucially, considering the significant differences in frequency domain information between visible light and infrared images—amplitude spectrum primarily carries global energy and high-frequency detail intensity information, while phase spectrum encodes key spatial structures (such as edges and textures) that determine content integrity—this embodiment proposes a Separable Spectral Enhancement Module (SSEM) based on Fast Fourier Transform to fully exploit the feature value of different frequency bands. This module uses a pre-defined separable spectral enhancement network to decouple and asymmetrically enhance the features of each modality before fusion, preserving key structural information to improve the information processing capability of subsequent feature fusion in the frequency domain. In other words, this embodiment introduces a split spectrum enhancement module (SSEM) in the feature enhancement stage, and the extracted features are: ; in These represent the branches in the backbone network that are oriented towards the visible light modal boundary image or the infrared modal boundary image, respectively. Feature maps of layer (i=1, 2, 3). , and These represent the height, width, and number of channels of the feature map, respectively. It is represented as a backbone network for single-modal feature extraction.
[0027] Furthermore, by utilizing a four-way Mamba network, i.e. Figure 2The four-way Mamba module effectively captures the global spatial dependencies of the image by scanning in parallel along the horizontal, vertical, and flip directions to complete spatial domain feature enhancement. Then, when using the obtained 2D feature map and the frequency domain feature enhancement based on the split-spectrum enhancement module, the following steps occur: The initial feature enhancement result is subjected to a Fast Fourier Transform using the preset split-spectrum enhancement network to obtain a spectrum mapping result; the spectrum mapping result is decomposed using the preset split-spectrum enhancement network to determine the amplitude spectrum and phase spectrum; wherein the amplitude spectrum includes high-frequency detail information, and the phase spectrum includes local spatial information; the amplitude spectrum is feature-enhanced through the amplitude spectrum processing branch in the preset split-spectrum enhancement network to determine a first enhancement result; the amplitude spectrum processing branch includes two cascaded convolutional blocks and activation functions; the phase spectrum is extracted using a dual-scale local feature extraction method through the phase spectrum processing branch in the preset split-spectrum enhancement network to determine a first local feature extraction result; the phase spectrum processing branch... The first local feature extraction result is subjected to feature cross-fusion at different scales to determine the cross-fusion result; the phase spectrum processing branch is used to extract local features at two scales from the cross-fusion result to determine the second local feature extraction result; the phase spectrum processing branch is used to fuse the second local feature extraction result to determine the second enhancement result; the first enhancement result and the second enhancement result are reconstructed into a complex frequency domain tensor using the preset split spectrum enhancement network to determine the tensor reconstruction result; the tensor reconstruction result is subjected to a two-dimensional inverse fast Fourier transform using the preset split spectrum enhancement network to determine the feature mapping result; the feature mapping result and the initial feature enhancement result are multiplied element-wise according to the adaptive gating mechanism in the preset split spectrum enhancement network to determine the processed feature map; the processed feature map and the initial feature enhancement result are added element-wise according to the residual connection structure in the preset split spectrum enhancement network to determine the feature enhancement result.
[0028] It is important to understand that regarding frequency domain feature enhancement based on a separate spectrum enhancement module, combined with... Figure 3 As shown, the core of SSEM lies in performing parallel, asymmetric, and targeted enhancement processing on the amplitude and phase spectra. First, the input feature map... Perform a Fast Fourier Transform to map it to a spectral representation, and then decompose it into an amplitude spectrum. Phase spectrum Two parts. Among them, Contains high-frequency detailed information. It carries local spatial information. Considering the differentiated information characteristics of the amplitude spectrum and phase spectrum, SSEM employs a parallel structure to model the two types of spectral components separately. In the amplitude spectrum processing branch, taking into account the global distribution characteristics of the amplitude spectrum, the module adopts a two-stage cascaded approach. Convolution, without disrupting the frequency domain spatial structure, enables global information exchange and adaptive feature enhancement between channels. Each convolutional stage is followed by an LReLU activation function (Leaky Rectified Linear Unit), introducing a non-linear mapping. The calculation process is shown below: ; In the formula, and Two levels convolution, This is the enhanced amplitude spectrum feature.
[0029] The phase spectrum carries core information about an image, including spatial location, edges, texture, and detailed structure. It is crucial for determining the integrity of image content, and capturing the contextual relationships between adjacent phase points is essential to fully preserve and enhance the image's structural information. To address the spatial information contained in the phase spectrum, the module designs a dual-scale cross-fusion convolutional architecture. This architecture extracts local features at different scales using convolutional kernels with different receptive fields (3×3 and 5×5), and achieves feature interaction across these scales through cross-merging. The computation process of this branch is shown below: ; ; .
[0030] In the formula, For the first The core size is convolutional layers, (Rectified LinearUnit) is an activation function. This is a feature concatenation operation at the channel dimension. This is the enhanced phase spectrum feature. The structure fully exploits feature information at different scales in the phase spectrum through a cross-fusion mechanism.
[0031] After enhancing the spectral components, SSEM first reconstructs the optimized amplitude and phase spectra into complex frequency domain tensors. Then, it maps the features back to the original spatial domain using a two-dimensional inverse fast Fourier transform. Based on this, the module introduces an adaptive gating mechanism and a residual connection structure. It achieves adaptive weighted filtering of the original features through element-wise multiplication, and then completes residual fusion through element-wise summation, finally outputting the module's enhanced features. The calculation process is shown below: ; In the formula, For frequency domain enhancement features in the spatial domain, For element-wise multiplication, This is an element-wise summation operation. The introduction of residual connections not only enhances the feature fusion effect but also effectively alleviates the gradient vanishing problem during deep network training, ensuring the training stability of the modules.
[0032] In summary, SSEM, through its core mechanism of frequency domain separation and differential enhancement, achieves in-depth mining of global and local feature information, effectively enhancing the model's expressive power and providing a frequency domain dimension modeling capability supplement to the Mamba fusion architecture.
[0033] Step S14: Input the feature enhancement results of each feature into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism.
[0034] In this embodiment, combined with Figure 2As shown, after feature enhancement is completed, a symmetric dual-branch architecture is constructed through a preset hybrid scanning fusion network, namely the Mixed Scanning Fusion Module (MSFM). This allows the global features of one modality to dynamically generate the core selective parameters of the state space model (SSM) of another modality, thereby achieving deep, bidirectional, and content-adaptive cross-modal interaction. Specifically, the feature enhancement results are mapped in a high-dimensional space according to the linear embedding layer in the preset hybrid scanning fusion network to determine the mapping result; the mapping result is segmented along the channel dimension to obtain the segmentation result; wherein the segmentation result includes the first and second feature map sequences corresponding to the visible light modality, and the third and fourth feature map sequences corresponding to the infrared modality; the first feature map sequence is flattened along the four spatial directions to determine the first flattening result; the third ... Flatten along four spatial directions to determine a second flattening result; input the first flattening result and the second flattening result into a preset cross-modal fusion network to complete bidirectional cross-modal information fusion, obtaining a first fusion result and a second fusion result; perform feature map reconstruction based on the first fusion result to determine a first fused feature map sequence; perform feature map reconstruction based on the second fusion result to determine a second fused feature map sequence; multiply the first fused feature map sequence element-wise using the second feature map sequence to determine a first multiplication result; multiply the second fused feature map sequence element-wise using the fourth feature map sequence to determine a second multiplication result; determine the feature fusion result using the first multiplication result, the second multiplication result, and a preset feature scanning branch.
[0035] It is important to understand that, in combination Figure 4 As shown, this MSFM module contains two symmetrical SSM processing branches: when processing visible light features, the key selectivity parameters in its state-space model are not independently calculated from the visible light features themselves, but are dynamically generated from the infrared features through a lightweight guiding network; and vice versa. This mechanism transcends traditional feature-level fusion, establishing a bidirectional, content-aware deep coupling, as shown in the following calculation: ; in , These represent the branches in the backbone network obtained in the first step that are oriented towards the visible light modal boundary image or the infrared modal boundary image, respectively. Feature maps of layer (i=1,2,3) F value. Indicates the first Layer fusion characteristics. This embodiment proposes a cross-modal dynamic guided Mamba cross-attention feature fusion module.
[0036] It is understandable that this embodiment extends the architecture of the original SSM. Its core innovation lies in the design of a cross-modal dynamic parameter generation mechanism: it allows the feature map of one modality to dynamically generate key parameters for regulating the internal state evolution of another modality SSM. Its core task is to achieve deep fusion of dual-modal features through a symmetrical parallel dual-branch architecture, combined with multi-branch sequence scanning modeling, the above-mentioned cross-modal dynamic parameter generation mechanism and multi-level residual connections, thereby improving the feature expression ability and robustness of the target detection network in complex scenarios.
[0037] It should be understood that, regarding feature fusion based on the MSFM module, the following related steps exist in this embodiment: Two sets of feature maps are generated using a method similar to that used in single-modal feature enhancement, derived from the visible light feature map. F rgb Received And from infrared feature maps F ir Received : ; .
[0038] Then, and Flattened along the four spatial directions respectively The generated one-dimensional sequence is input into the Four-way Cross-modal Fusion Module (FCFM) for cross-modal information fusion. In the FCFM, for each pair of sequences ( Perform bidirectional cross-modal modeling: ; .
[0039] In FCFM, the preceding input serves as the core subject to be processed, carrying the task of transferring the core features of the modality; the following input is specifically used to dynamically generate the projection parameters and time scale parameters required for SSM, providing a control basis for modeling the subject sequence. The visible light and infrared modes are the preceding and following inputs to each other, resulting in a symmetrical fusion structure, which not only ensures the equal interactive status of the two modes, but also fully explores the bidirectional cross-modal correlation information.
[0040] Expand the outputs in the four spatial directions respectively. This is restored to a 2D spatial feature map, which is then obtained through element-wise weighted summation. and : .
[0041] To preserve key information from the original input and enhance feature correlation, two sets of calibration features were used. Each element will be multiplied with its corresponding preceding input feature map to obtain the results. and .
[0042] ; .
[0043] in, , Represents generation , of Convolutional layer; SiLU (Sigmoid Linear Unit) is an activation function.
[0044] Meanwhile, the dynamic adaptive scanning branch in the MSFM module works in parallel: A first preset feature scanning branch globally enhances the visible light feature map in the feature enhancement result to determine the first global enhancement feature; wherein the first preset feature scanning branch includes a deep convolutional network (i.e., depthwise convolution (DWConv)), a layer normalization network, a batch normalization network, a convolutional feedforward neural network, and a state space model based on a dynamic adaptive scanning strategy (i.e., DAS-SSM (Dynamic Adaptive Scan - State Space Model, an improved state space model for vision tasks)); a second preset feature scanning branch globally enhances the infrared feature map in the feature enhancement result to determine the second global enhancement feature; wherein the second preset feature scanning branch has the same structure as the first preset feature scanning branch; the first global enhancement feature is fused with the first multiplication result to determine the target visible light fusion feature; the second global enhancement feature is fused with the second multiplication result to determine the target infrared fusion feature; the feature fusion result is determined using the target convolutional block, the target visible light fusion feature, and the target infrared fusion feature. The first and second preset feature scanning branches are... Figure 4 Dynamic adaptive scanning branches on the left and right sides of the center.
[0045] Regarding the dynamic adaptive scanning branch, the global enhancement features output by the scan are fused with the local focused enhancement features output by the dynamic adaptive scanning branch through element-wise addition, resulting in the final sequence enhancement features after hybrid scanning modeling: ; .
[0046] The resulting visible light fusion output Combined output with infrared Then through The convolution operation yields the final fused output. : .
[0047] Step S15: Perform cross-scale consistency verification on the feature fusion result according to the preset border target detection model, and update the feature fusion result using the corresponding spatial confidence map to obtain the feature update result.
[0048] In this embodiment, before the final target detection, the cross-scale consistency of the fused features is checked and optimized. This is to address the issue of multi-scale fused features { , , The potential semantic inconsistencies between them can be addressed by utilizing deep features. Strong semantic information, for shallow features , A top-down verification and guidance process is performed. By generating a global semantic descriptor and calculating a spatial confidence map, the shallow fusion features are weighted and modulated to obtain optimized features. , , To enhance internal semantic consistency, the following steps are taken: The feature fusion result is divided according to the feature scale to obtain deep fusion features and shallow fusion features; the deep fusion features are input into the target channel attention layer of the preset border target detection model to obtain a global semantic descriptor; the global semantic descriptor is upsampled and broadcast to the spatial dimension corresponding to the shallow fusion features; the element-wise similarity between the shallow fusion features and the global semantic descriptor is analyzed, and a spatial confidence map is determined based on the analysis results; the shallow fusion features are weighted and modulated according to the spatial confidence map to obtain modulated shallow fusion features; and the feature update result is determined based on the deep fusion features and the modulated shallow fusion features.
[0049] It is important to understand that, to address the potential semantic inconsistencies between multi-scale fused features obtained after MSFM processing, this embodiment introduces a cross-scale consistency verification and optimization mechanism. This mechanism utilizes the strong semantic information contained in deep fused features to perform top-down verification and guidance on shallow fused features. Specifically, it first verifies and guides the deepest fused feature map... Its global semantic descriptor is generated through a lightweight channel attention module. G Subsequently, G Broadcast to via upsampling operation and The spatial dimension. For each shallow feature layer. (i=3, 4), calculate its relationship with the global descriptor after broadcast. G i Element-wise similarity is used to generate a spatial confidence map. C i This confidence plot reflects the degree of consistency between each location in the shallow features and the global semantics. Finally, the confidence plot is used... C i Original shallow fusion features Perform weighted modulation = C i ten ,in The first term represents element-wise multiplication, and the second term is a residual join to preserve the original information. The resulting multi-scale feature after this step is { , , When fed into the neck network, its internal semantic consistency is significantly enhanced.
[0050] In this way, a global semantic descriptor is generated using deep fusion features in a top-down manner, and the shallow fusion features are then spatially weighted and modulated accordingly, enhancing semantic alignment while preserving details. This mechanism is specifically tailored to the fusion output characteristics of this embodiment, effectively improving the overall consistency and detection robustness of the feature pyramid.
[0051] Step S16: Input the feature update result into the neck network of the preset border target detection model to complete the target detection corresponding to the border image to be detected and obtain the target border inspection result.
[0052] In this embodiment, combined with Figure 2 As shown, after completing all processing of the border image, the obtained feature update results are input into the neck network of the preset border target detection model for multi-scale feature aggregation and detection head output. That is, the aggregated features obtained after cross-scale consistency verification and optimization are... , , The data is fed into the neck network for multi-scale feature fusion, and then downstream detection heads are used to achieve dual-modal target detection. The specific calculation formula is shown below: .
[0053] In the formula, The parameter is The YOLOv8 detection head integrates regression and classification branch outputs to generate detection results. , , These represent the decoded coordinates of the predicted bounding box boundaries, the class probability, and the detection box confidence information, respectively. N represents the total number of valid detected targets output by the detection head; This indicates traversing from k=1 to N. It's understandable that the detection targets in border patrol tasks can be configured and updated according to actual needs; further examples will not be provided.
[0054] In summary, the proposed scheme in this embodiment provides a high-quality single-modal feature foundation with sufficient enhancement in the frequency domain for subsequent deep fusion. Based on this, the cross-modal dynamically guided SSM fusion method achieves deeper inter-modal information interaction and complementarity than traditional fusion strategies. Finally, the cross-scale consistency verification and optimization mechanism further ensures the internal semantic consistency of the multi-scale features after fusion, effectively improving the quality of the feature pyramid. The synergistic effect of these three elements systematically solves the problem of insufficient dual-modal feature fusion, significantly improves the model's robustness in complex scenarios, and effectively overcomes the limitation of existing related schemes that can only model within a single modality, providing a novel and efficient solution for dual-modal target detection.
[0055] Therefore, in the border drone inspection system of this application, the visible light and infrared modes of border images are first acquired simultaneously by the border inspection drone. The obtained border images to be detected are then input into the backbone network of a preset border target detection model to obtain feature extraction results. Next, the feature extraction results are input into the model's cross-modal dynamically guided Mamba feature fusion network to perform feature enhancement in the spatial domain and amplitude and phase decoupling feature enhancement in the frequency domain, resulting in feature enhancement results. Then, the feature enhancement results are input into a preset hybrid scanning fusion network to obtain feature fusion results. Finally, cross-scale consistency verification is performed on the feature fusion results, and feature updates are triggered using the corresponding spatial confidence map to obtain feature update results. The feature update results are then input into the neck network of the preset border target detection model to obtain the target border inspection results. This approach solves the problems existing in related solutions, achieving deep, bidirectional cross-modal interaction. While ensuring the feasibility of edge deployment, it enhances the internal semantic consistency of the detected features and improves detection accuracy and environmental robustness in complex, crowded, or severely occluded scenarios.
[0056] See Figure 5 As shown in the figure, this application also discloses a border inspection device based on cross-modal dynamic guided Mamba feature fusion, applied to a border drone inspection system, including: The border image acquisition module 11 is used to simultaneously acquire dual-modal border images via a border patrol drone to determine the border image to be detected; wherein, the dual-modal includes a visible light mode and an infrared mode; The image feature extraction module 12 is used to input the border image to be detected into the backbone network of the preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results; the preset border target detection model is located on the end side of the border patrol drone; The dual-domain feature enhancement module 13 is used to input the feature extraction results of each feature into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial domain and frequency domain, and obtain feature enhancement results; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling. The cross-modal fusion module 14 is used to input the feature enhancement results of each feature into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; The feature verification module 15 is used to perform cross-scale consistency verification on the feature fusion result according to the preset border target detection model, and update the feature fusion result using the corresponding spatial confidence map to obtain the feature update result; The inspection result determination module 16 is used to input the feature update result into the neck network in the preset border target detection model to complete the target detection corresponding to the border image to be detected and obtain the target border inspection result.
[0057] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0058] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the border inspection method based on cross-modal dynamic guided Mamba feature fusion disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0059] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0060] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0061] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the border inspection method based on cross-modal dynamic guided Mamba feature fusion, which is executed by the electronic device 20 according to any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0062] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned border inspection method based on cross-modal dynamic guided Mamba feature fusion. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0063] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0064] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0065] The steps of the methods or algorithms described in conjunction with the embodiments disclosed in this embodiment can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0066] Finally, it should be noted that in this embodiment, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0067] The technical solutions provided in this application have been described in detail above. Specific examples have been used in this embodiment to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A border inspection method based on cross-modal dynamic guided Mamba feature fusion, applied to a border UAV inspection system, characterized in that, include: Border patrol drones simultaneously acquire dual-modal border images to identify the border image to be detected; wherein, the dual-modal includes visible light mode and infrared mode; The border image to be detected is input into the backbone network of a preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results; the preset border target detection model is located at the end of the border patrol drone; The feature extraction results are input into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial and frequency domains and obtain feature enhancement results; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling. The feature enhancement results are input into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; The feature fusion results are subjected to cross-scale consistency verification based on the preset border target detection model, and the feature fusion results are updated using the corresponding spatial confidence map to obtain the feature update results; The feature update result is input into the neck network of the preset border target detection model to complete the target detection corresponding to the border image to be detected, and the target border inspection result is obtained.
2. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to claim 1, characterized in that, The step of inputting the border image to be detected into the backbone network of a preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results includes: Input the border image to be detected into the backbone network of the preset border target detection model; The convolutional module, fast spatial pyramid pooling module, and C2f module in the backbone network are used to extract single-modal features from the border images to be detected for each modality, so as to obtain the feature extraction results corresponding to each modality.
3. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to claim 1, characterized in that, The step of inputting the feature extraction results into the cross-modal dynamically guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial and frequency domains includes: The feature extraction results are input into a four-way Mamba network in a cross-modal dynamically guided Mamba feature fusion network; The feature extraction results are mapped and expanded by the linear embedding layer in the four-way Mamba network to complete the feature enhancement operation in the spatial domain and obtain the initial feature enhancement results corresponding to each modality. The initial feature enhancement result is input into the preset split spectrum enhancement network for frequency domain feature enhancement to obtain the output feature enhancement result.
4. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to claim 3, characterized in that, The step of inputting the initial feature enhancement result into the preset split spectrum enhancement network for frequency domain feature enhancement to obtain the output feature enhancement result includes: The initial feature enhancement result is subjected to a fast Fourier transform using the preset split spectrum enhancement network to obtain the spectrum mapping result; The preset split-type spectrum enhancement network is used to decompose the spectrum mapping result to determine the amplitude spectrum and phase spectrum; wherein, the amplitude spectrum includes high-frequency detail information, and the phase spectrum includes local spatial information; The amplitude spectrum is enhanced by the amplitude spectrum processing branch in the preset split spectrum enhancement network to determine the first enhancement result; the amplitude spectrum processing branch includes two cascaded convolutional blocks and activation functions. The phase spectrum is subjected to dual-scale local feature extraction through the phase spectrum processing branch in the preset split spectrum enhancement network to determine the first local feature extraction result; The phase spectrum processing branch is used to perform feature cross-fusion at different scales on the first local feature extraction result to determine the cross-fusion result. The phase spectrum processing branch is used to perform dual-scale local feature extraction on the cross-fusion result to determine the second local feature extraction result; The second local feature extraction result is fused using the phase spectrum processing branch to determine the second enhancement result; The first enhancement result and the second enhancement result are reconstructed using a complex frequency domain tensor using the preset split spectrum enhancement network to determine the tensor reconstruction result; The preset split-spectrum enhancement network is used to perform a two-dimensional inverse fast Fourier transform on the tensor reconstruction result to determine the feature mapping result; The feature mapping result and the initial feature enhancement result are multiplied element-wise according to the adaptive gating mechanism in the preset split spectrum enhancement network to determine the processed feature map. The processed feature map and the initial feature enhancement result are added element-wise according to the residual connection structure in the preset split spectrum enhancement network to determine the feature enhancement result.
5. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to claim 1, characterized in that, The step of inputting each of the feature enhancement results into the preset hybrid scanning fusion network in the cross-modal dynamically guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results includes: The feature enhancement results are mapped in a high-dimensional space according to the linear embedding layer in the preset hybrid scanning fusion network to determine the mapping result; The mapping result is segmented along the channel dimension to obtain a segmentation result; wherein, the segmentation result includes a first feature map sequence and a second feature map sequence corresponding to the visible light mode, and a third feature map sequence and a fourth feature map sequence corresponding to the infrared mode; The first feature map sequence is flattened along four spatial directions to determine the first flattening result; The third feature map sequence is flattened along four spatial directions to determine the second flattening result; The first flattening result and the second flattening result are respectively input into a preset cross-modal fusion network to complete bidirectional cross-modal information fusion and obtain the first fusion result and the second fusion result; Based on the first fusion result, the feature map is restored to determine the first fused feature map sequence; Based on the second fusion result, the feature map is restored to determine the second fused feature map sequence; The first fused feature map sequence is multiplied element-wise using the second feature map sequence to determine the first multiplication result; The second fused feature map sequence is multiplied element-wise using the fourth feature map sequence to determine the second multiplication result; The feature fusion result is determined by using the first multiplication result, the second multiplication result, and the preset feature scanning branch.
6. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to claim 5, characterized in that, The step of determining the feature fusion result using the first multiplication result, the second multiplication result, and a preset feature scanning branch includes: The visible light feature map in the feature enhancement result is globally enhanced by a first preset feature scanning branch to determine the first global enhancement feature; wherein, the first preset feature scanning branch includes a deep convolutional network, a layer normalization network, a batch normalization network, a convolutional feedforward neural network, and a state space model based on a dynamic adaptive scanning strategy. The infrared feature map in the feature enhancement result is globally enhanced by a second preset feature scanning branch to determine the second global enhancement feature; wherein, the second preset feature scanning branch has the same structure as the first preset feature scanning branch; The first global enhancement feature is fused with the first multiplication result to determine the target visible light fusion feature; The second global enhancement feature is fused with the second multiplication result to determine the target infrared fusion feature; The feature fusion result is determined using the target convolutional block, the target visible light fusion feature, and the target infrared fusion feature.
7. The border inspection method based on cross-modal dynamic guided Mamba feature fusion according to any one of claims 1 to 6, characterized in that, The step of performing cross-scale consistency verification on the feature fusion result based on the preset border target detection model, and updating the feature fusion result using the corresponding spatial confidence map, includes: The feature fusion results are divided according to the feature scale to obtain deep fusion features and shallow fusion features; The deep fusion features are input into the target channel attention layer of the preset border target detection model to obtain a global semantic descriptor; By performing an upsampling operation on the global semantic descriptor, the global semantic descriptor is broadcast to the spatial dimension corresponding to the shallow fusion feature; The spatial confidence map is determined by analyzing the element-wise similarity between the shallow fusion features and the global semantic descriptor, and based on the corresponding analysis results. The shallow fusion features are weighted and modulated based on the spatial confidence map to obtain the modulated shallow fusion features; The feature update result is determined based on the deep fusion features and the modulated shallow fusion features.
8. A border inspection device based on cross-modal dynamic guided Mamba feature fusion, applied to a border UAV inspection system, characterized in that, include: A border image acquisition module is used to simultaneously acquire dual-modal border images via a border patrol drone to identify the border image to be detected; wherein, the dual-modal mode includes a visible light mode and an infrared mode; The image feature extraction module is used to input the border image to be detected into the backbone network of the preset border target detection model to complete single-modal image feature extraction and obtain feature extraction results; the preset border target detection model is located on the end side of the border inspection UAV. The dual-domain feature enhancement module is used to input the feature extraction results of each feature into the cross-modal dynamic guided Mamba feature fusion network in the preset border target detection model to complete feature enhancement in the spatial domain and frequency domain, and obtain feature enhancement results; wherein, the cross-modal dynamic guided Mamba feature fusion network includes a preset split spectrum enhancement network, which is used to perform feature enhancement with amplitude and phase decoupling. The cross-modal fusion module is used to input the feature enhancement results of each feature into the preset hybrid scan fusion network in the cross-modal dynamic guided Mamba feature fusion network to complete cross-modal feature fusion and obtain feature fusion results; wherein, the preset hybrid scan fusion network includes a preset feature scan branch and a preset cross-modal dynamic parameter generation mechanism; The feature verification module is used to perform cross-scale consistency verification on the feature fusion result according to the preset border target detection model, and update the feature fusion result using the corresponding spatial confidence map to obtain the feature update result; The inspection result determination module is used to input the feature update result into the neck network in the preset border target detection model to complete the target detection corresponding to the border image to be detected and obtain the target border inspection result.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the border inspection method based on cross-modal dynamic guided Mamba feature fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the border inspection method based on cross-modal dynamic guided Mamba feature fusion as described in any one of claims 1 to 7.