Essential feature adaptive visible light infrared fusion detection and recognition method and system
Through the X-type multi-layer feature interactive fusion module, the dynamic attention fusion module of conditional convolution and splicing operations, and the recomputation-bidirectional jump pyramid structure, the visible light and infrared image features are adaptively extracted and fused, solving the problem of insufficient fusion and improving the recognition accuracy and environmental adaptability.
Patent Information
- Application Number
- CN202410687752.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2044-05-30
AI Technical Summary
The existing visible light infrared fusion detection and recognition technology is difficult to adaptively extract multimodal features in complex and changing application scenarios, resulting in insufficient fusion and affecting the recognition effect.
The multi-scale features of visible light and infrared images are adaptively extracted and fused multi-scale features of visible light and infrared images using X-type multi-layer feature interactive fusion module, dynamic attention fusion module based on conditional convolution and splicing operations, and recomputation-quantitative bidirectional jump pyramid structure.
It improves the accuracy of visible infrared fusion recognition, adapts to object detection and recognition under different environmental conditions, and solves the problem of a single optical sensor being affected by environmental factors.
Smart Images

Figure CN118658028B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a method and system for detecting and identifying visible light and infrared fusion with adaptive essential features. Background Art
[0002] With the continuous advancement of image sensor technology, image sensors of different wavelengths have been widely used across various industries. Consequently, multi-band image fusion technology has also attracted increasing attention. Multi-band image fusion combines information from images captured by different image sensors within the same scene to obtain more comprehensive and accurate fused information. It has important applications in remote sensing, detection, military reconnaissance, security monitoring, healthcare, industrial production, and other fields. Among them, the acquisition and application of visible and infrared imagery are the most common. Infrared images can highlight the characteristics of infrared thermal target areas, but they often lack detailed information and have low contrast. Visible images, on the other hand, can reveal information such as texture, color, and contours of the imaged area. By fusing visible and infrared images, we can preserve both the thermal radiation information in the infrared image and the clear appearance and texture information in the visible image, thus obtaining a comprehensive and accurate description of the current scene. Combining visible and infrared image fusion with specific scene tasks, such as target detection and recognition, target tracking, and target localization, can achieve superior performance compared to single-source image sensors. In the typical task of target detection and recognition, visible light infrared fusion processing can reduce the impact of external factors on the detection of multiple targets. It is suitable for target detection and recognition tasks in all weather and at all times, and has received widespread attention and research.
[0003] Visible-light and infrared fusion detection and recognition technologies can be categorized into two main groups: traditional and intelligent methods. Traditional methods can generally be divided into two categories: transform-domain-based and spatial-domain-based fusion methods. Transform-domain-based methods commonly perform multiscale decomposition of the source image to be processed, then design fusion rules to fuse the decomposed images, and use multiscale inverse transforms to obtain the final fused image. Common multiscale decomposition methods include Laplace transform, Curvelet transform, and wavelet transform, as well as improved multiscale low-rank decomposition, adaptive sparse representation, and particle swarm optimization algorithms. Transform-domain-based methods improve fusion performance by refining and optimizing multiscale decomposition methods and fusion rules. However, in complex and diverse application scenarios, manual design of decomposition and fusion methods is often required, making them difficult to meet practical application requirements. Traditional spatial-domain-based fusion methods divide the source image into several sub-images and then combine them to obtain the final fused image. Alternatively, the target fusion region and background region of the source image can be superimposed on the background image to obtain the final fused image. The traditional fusion method based on the spatial domain has a good fusion effect for local images, but the fused image is prone to produce obvious edge effects at the edges of the segmented blocks, which affects the further processing of the image.
[0004] With the rapid development of artificial intelligence technology, deep learning methods based on neural networks have rapidly developed in visible-light and infrared image fusion processing. Researchers use pre-trained convolutional neural network models to provide deep learning feature maps, which are then fused. This method significantly improves performance compared to traditional fusion methods, but the fusion process is not end-to-end and relies on the design of multi-scale decomposition and fusion rules. Other researchers have proposed a fusion method based on generative adversarial networks. This method first constructs loss functions for the generator and discriminator to iteratively optimize network parameters. A dual discriminator is then designed to enhance image fusion. The loss function incorporates image content and edge information to ensure that the fused image retains both visible light image detail and infrared image edge information. Furthermore, researchers define a loss function called Local Binary Patterns (LBPs) to ensure that the generated fused image has rich edge information. The advantage of this method is that different loss functions can be assigned to the generator and discriminator to meet different fusion requirements. However, its disadvantages are that the network training process requires two stages: generation and discrimination. Furthermore, the trained network model depends on the distribution of the dataset data, resulting in poor robustness and generalization of the neural network. To address this problem, the researchers proposed an end-to-end deep convolutional neural network to perform local Gaussian blur construction, and then achieved image fusion by learning an end-to-end convolutional neural network. This method transforms unsupervised image fusion into a supervised image classification problem and is therefore more suitable for multi-source image fusion tasks. The researchers further designed an end-to-end fusion network for the fusion of visible and infrared images, using target structure similarity and pixel difference as loss functions. The learned shallow and deep features are combined in the feature extraction model, making it suitable for the fusion of infrared, visible, and other multimodal images.
[0005] However, visible light and infrared are two modalities with significantly different features, making them prone to inadequate fusion. Research on inadequate modal fusion focuses on the fusion stage, which can be categorized into image-level fusion, feature-level fusion, and decision-level fusion. Feature-level fusion is the most popular, with numerous researchers improving the feature fusion architecture of deep neural networks. Choi et al. (2016) proposed balancing feature recognition capabilities and resolution between different convolutional layers by fusing multi-scale feature maps. Zheng et al. (2019) designed a cross-modal detector based on dual SSD (single-shot detectors), utilizing the six different-scale feature maps in the SSD model to detect objects of varying scales. To leverage the complementarity between the two modalities, Zhang et al. (2019b) first applied the attention mechanism to cross-modal pedestrian detection, proposing a cross-modal interactive attention network. Yang et al. (2022) proposed a bidirectional adaptive attention gate (BAA-Gate) fusion module based on the attention mechanism. The attention mechanism is used to suppress noise in both modalities while simultaneously selecting valid information between them. Kim et al. (2022a) proposed a cross-modal pedestrian detection framework based on region of interest (RoI) uncertainty and prediction uncertainty.
[0006] However, all of the above methods are based on artificially designed feature extraction operators and structures. Once training is completed, they are unable to dynamically and adaptively extract the features of the input modality, resulting in insufficient fusion of features from different modalities. Compared with single-modal recognition, the other modality may even have a negative effect, further affecting the visible light infrared fusion recognition effect. Summary of the Invention
[0007] Single-source detection networks can only use data from a single-source sensor for target detection and are unable to simultaneously use data from multiple sources for fusion detection and recognition. In order to solve the above problems, the present invention proposes a method and system for visible light infrared fusion detection and recognition with adaptive essential features. By making full use of the rich colors and texture characteristics of visible light, the high image resolution, and the characteristics of infrared images that can effectively form images in poor light conditions and at night, the target is fused for detection and recognition. The present invention designs an X-shaped multi-layer feature interaction and fusion structure, builds a dynamic attention fusion module based on conditional convolution to extract essential fusion features, and uses a heavy-computation bidirectional jump pyramid structure to further retain high-precision multi-scale visible light infrared fusion features, thereby improving the accuracy of fusion recognition.
[0008] The technical solution adopted in the present invention is as follows:
[0009] On the one hand, the present invention proposes a method for detecting and identifying visible light and infrared fusion with adaptive essential features, comprising:
[0010] The features of images of different modalities at different scales are extracted through a preset X-shaped multi-layer feature interactive fusion module, and features with differences greater than a threshold are subjected to X-shaped interactive fusion to form a comprehensive feature sensitive to the target feature; the modalities include visible light and infrared;
[0011] The dynamic attention fusion module based on conditional convolution and splicing operations amplifies the feature differences between the modalities, calculates the compensation and compatible features between the modalities, and fuses them with the original features of the corresponding modalities to obtain fused features that are consistent with the target features.
[0012] The modal fusion feature analysis module based on the heavy-computation bidirectional jump pyramid connects the fusion features of different scales in sequence and cross-layer hybrid connection to perform target category classification and position regression, thereby obtaining fusion detection and recognition results.
[0013] Furthermore, the features of images of different modalities at different scales are extracted through a preset X-shaped multi-layer feature interactive fusion module, including: using two residual networks composed of residual blocks to respectively extract features of visible light and infrared images at different scales.
[0014] Furthermore, a dynamic attention fusion module based on conditional convolution and splicing operations is embedded in each layer of the residual network of the X-type multi-layer feature interaction fusion module, and the similarities and differences of the features of different modal data in each layer are learned.
[0015] Furthermore, in the dynamic attention fusion module based on conditional convolution and splicing operations, the conditional convolution calculates a weighted convolution kernel on the input sample before performing the convolution calculation, and each convolution kernel is calculated once and acts on the positions of different data.
[0016] Furthermore, in the modal fusion feature analysis module based on the heavy-computation bidirectional skip-connection pyramid, the original bidirectional pyramid structure is expanded from the original U-shaped feature transfer loop to N U-shaped feature transfer loops, deepening the connection between the backbone network and the detection head, and realizing the learning and extraction of multi-scale features; the layer-by-layer connection of features in the original bidirectional pyramid structure is optimized into a sequential and skip-level mixed connection, expanding the combination types of multi-scale features.
[0017] On the other hand, the present invention proposes a visible light infrared fusion detection and recognition system with adaptive essential features, comprising:
[0018] An X-shaped multi-layer feature interaction fusion module is configured to extract features of images of different modalities at different scales and perform X-shaped interaction fusion on features with differences greater than a threshold to form a comprehensive feature sensitive to the target feature; the modalities include visible light and infrared;
[0019] The dynamic attention fusion module based on conditional convolution and splicing operations is configured to amplify the feature differences between the modalities, calculate the compensation and compatibility features between the modalities, and fuse them with the original features of the corresponding modality to obtain fused features that are consistent with the target features;
[0020] The modal fusion feature analysis module based on the heavy-computation bidirectional jump pyramid is configured to connect the fusion features of different scales in sequence and cross-layer hybrid connection to perform target category classification and position regression, thereby obtaining a fusion detection and recognition result.
[0021] Furthermore, the features of images of different modalities at different scales are extracted through a preset X-shaped multi-layer feature interactive fusion module, including: using two residual networks composed of residual blocks to respectively extract features of visible light and infrared images at different scales.
[0022] Furthermore, a dynamic attention fusion module based on conditional convolution and splicing operations is embedded in each layer of the residual network of the X-type multi-layer feature interaction fusion module, and the similarities and differences of the features of different modal data in each layer are learned.
[0023] Furthermore, in the dynamic attention fusion module based on conditional convolution and splicing operations, the conditional convolution calculates a weighted convolution kernel on the input sample before performing the convolution calculation, and each convolution kernel is calculated once and acts on the positions of different data.
[0024] Furthermore, in the modal fusion feature analysis module based on the heavy-computation bidirectional skip-connection pyramid, the original bidirectional pyramid structure is expanded from the original U-shaped feature transfer loop to N U-shaped feature transfer loops, deepening the connection between the backbone network and the detection head, and realizing the learning and extraction of multi-scale features; the layer-by-layer connection of features in the original bidirectional pyramid structure is optimized into a sequential and skip-level mixed connection, expanding the combination types of multi-scale features.
[0025] The beneficial effects of the present invention are:
[0026] 1. This paper designs an X-shaped multi-layer feature interaction and fusion structure, builds a dynamic attention fusion module based on conditional convolution to extract essential fusion features, and uses a heavy-computation bidirectional pyramid structure to further retain high-precision multi-scale visible light infrared fusion features, thereby improving the positioning and recognition accuracy based on multimodal data and enhancing the accuracy of visible light infrared fusion recognition.
[0027] 2. This invention fully utilizes the color, texture, and high-resolution characteristics of visible light, as well as the ability of infrared images to effectively image in low-light conditions and at night. It adaptively learns the multi-scale compatible fusion features of visible and infrared images, automatically adjusts feature weights based on the specific conditions of the visible and infrared images, and performs multi-scale detection and recognition, improving the accuracy of fusion recognition. This invention effectively addresses issues such as poor single-source target detection and recognition due to deteriorating environmental conditions such as lighting and visibility, and insufficient target feature fusion caused by the inability of trained models to dynamically adapt to input images of varying quality.
[0028] 3. The present invention is applicable to target detection and recognition task scenarios in airborne and vehicle-mounted platforms equipped with both visible light and infrared optical sensors. It can solve the problem of low recognition accuracy caused by single optical sensors being easily affected by environmental factors such as weather, lighting, and distance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is a schematic diagram of an essential feature adaptive visible light infrared fusion detection and recognition system in Example 1 of the present invention;
[0030] Figure 2 Schematic diagram of an X-type multi-layer feature interactive fusion module in Example 1 of the present invention;
[0031] Figure 3 Schematic diagram of a dynamic attention fusion module based on conditional convolution and splicing operations in Example 1 of the present invention;
[0032] Figure 4 Schematic diagram of a modal fusion feature analysis module based on a recalculated bidirectional jump pyramid in Example 1 of the present invention. DETAILED DESCRIPTION
[0033] In order to have a clearer understanding of the technical features, purposes and effects of the present invention, the specific embodiments of the present invention are now described. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. That is, the embodiments described are only part of the embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0034] Example 1
[0035] This embodiment provides a visible-infrared fusion detection and recognition system with adaptive intrinsic features. It utilizes an X-shaped multi-layer feature interaction fusion module to fully fuse image data from different modalities. A dynamic attention fusion module based on conditional convolution and splicing operations is then used to obtain compatible fusion features from these different modal images. Finally, target classification and position regression are performed through the sequencing of fused features at multiple scales and cross-layer hybrid connections, thereby obtaining fused detection and recognition results. This solution adaptively learns multi-scale compatible fusion features for visible and infrared images based on their characteristics and advantages, adaptively adjusts feature weights based on the specific conditions of the visible and infrared images, and performs multi-scale detection and recognition, thereby improving fusion effectiveness.
[0036] like Figure 1 As shown, the visible light infrared fusion detection and recognition system of this embodiment includes: an X-type multi-layer feature interaction fusion module, a dynamic attention fusion module based on conditional convolution and splicing operations, and a modal fusion feature analysis module based on a heavy-computation bidirectional jump pyramid, which are specifically described as follows.
[0037] (1) X-type multi-layer feature interaction fusion module
[0038] Since the features of visible light and infrared images are quite different, this embodiment uses two backbone networks composed of residual blocks to extract features of the two types of images at different scales. In the second half of the backbone network, since the features of visible light and infrared images are already highly abstracted, structures such as dynamic convolutional attention are used to perform X-shaped interactive fusion of the features with large differences to form comprehensive features that are sensitive to the target features, and further pass them backward to obtain homogeneous features. The X-shaped multi-layer feature interactive fusion module is as follows: Figure 2 shown.
[0039] (2) Dynamic attention fusion module based on conditional convolution and splicing operations
[0040] Convolution has achieved unprecedented success in many computer vision tasks, but its performance gains stem primarily from increases in model size and capacity, as well as larger datasets. This increase in model size further increases computational complexity, making it more difficult to deploy high-quality models. Existing convolution algorithms use the same convolution parameters for all training examples. This results in multimodal tasks, such as visible and infrared target detection and recognition, requiring increased model parameters, depth, and number of channels to extract features from multiple modal data and improve model capacity. This further increases computational complexity and makes deployment more difficult.
[0041] To address the above issues, this embodiment designs a dynamic attention fusion module based on conditional convolution and splicing operations to achieve the fusion of visible light and infrared images. Conditional convolution breaks the traditional static convolution characteristics by inputting the calculation of convolution kernel parameters. The convolution kernel is parameterized as a linear combination of multiple expert knowledge: ,in is a weighting coefficient learned via gradient descent. To more effectively improve the capacity of multimodal models, such as those for visible light and infrared, the number of experts can be increased during network design. This is more efficient than increasing the size of the convolution kernel. Furthermore, expert knowledge only needs to be combined once, which increases model capacity while maintaining efficient inference.
[0042] In specific computations, conditional convolution uses multiple experts to calculate weighted convolution kernels on the input samples before performing the convolution calculation. Each convolution kernel only needs to be calculated once and applied to different data locations. This increases the capacity of multimodal visible and infrared networks by increasing the amount of expert data. The code consumes minimal inference time: each additional parameter requires only a single multiplication and addition. Dynamic convolution is equivalent to a linear combination of multiple static convolutions. Therefore, it has the same capacity as n experts, but with more efficient data computation for multimodal tasks.
[0043] A dynamic attention fusion module based on conditional convolution and concatenation is embedded in each layer of the residual network to learn the similarities and differences between the features of the different modal data at each layer. By amplifying the feature differences between the modalities, compensatory and compatible features are calculated and fused with the original features of the corresponding modalities to achieve a complementary effect of the different modal features, ultimately obtaining a fusion feature of visible light and infrared with consistent target features.
[0044] Dynamic attention fusion module based on conditional convolution and splicing operations such as Figure 3 shown.
[0045] (3) Modal fusion feature analysis module based on heavy-computation bidirectional jump pyramid
[0046] The X-type multi-layer feature interaction fusion module and the dynamic attention fusion module based on conditional convolution and splicing operations extract the fusion features of the visible light and infrared modalities, so as to obtain favorable information for target recognition and positioning that can complement each other's advantages in the two modalities. Before obtaining the target type location information in an end-to-end manner, in order to further fully fuse the visible light infrared fusion features, an improved modal fusion feature analysis module based on a heavy-duty bidirectional jump pyramid is used behind the network backbone to perform high-precision position prediction and classification of the visible light and infrared fusion features in multi-scale space. The modal fusion feature analysis module based on a heavy-duty bidirectional jump pyramid is as follows: Figure 4 shown.
[0047] The differences between the computationally intensive bidirectional skip-connection pyramid of this embodiment and the original bidirectional pyramid structure are: 1) the original bidirectional pyramid structure is expanded from the original U-shaped feature transfer loop to N U-shaped feature transfer loops, deepening the connection between the intelligent detection model backbone and the detection head, which facilitates the learning and extraction of multi-scale features; 2) the layer-by-layer connection of features in the original bidirectional pyramid structure is optimized to a mixed sequential and skip-level connection mode, enriching the scale combination types of visible light and infrared features and improving the multi-scale feature extraction capability.
[0048] Example 2
[0049] This embodiment is based on embodiment 1:
[0050] This embodiment provides a method for detecting and recognizing visible light and infrared fusion based on adaptive essential features. The method mainly includes two parts: model training and model inference. The specific implementation steps are as follows:
[0051] S1: Use visible light sensors and infrared sensors to obtain a large number of pixel-level aligned image pairs, which are divided into training sets and test sets;
[0052] S2: Build a visible-light and infrared fusion detection and recognition network with adaptive intrinsic features. The innovation of this network lies in the design of a dynamic attention fusion module based on conditional convolution and a computationally efficient bidirectional jump-connection pyramid modal fusion feature analysis module in the multimodal target intelligent detection and recognition model.
[0053] S3: Use the training set to train the intrinsic feature-adaptive visible-light and infrared fusion detection and recognition network, and test the trained model using the test data. During training and testing, input a visible light image to the visible light branch of the network backbone, and an infrared image to the infrared branch. Use maximum suppression post-processing to filter the output of the fusion detection and recognition network to obtain the final comprehensive detection and recognition results.
[0054] Preferably, the workflow of the dynamic attention fusion module based on conditional convolution is:
[0055] 1) The extracted visible light image features and infrared image features are subjected to spatial attention feature analysis respectively. Spatial attention is composed of operators such as global average pooling (GAP), fully connected (FC), and ReLU activation function;
[0056] 2) Use full connectivity and Sigmoid activation functions to map the attention features to the weighted weights of the convolution operator group and update the parameters of the conditional convolution. The spatial attention features of the visible light image features are represented as Z, and the spatial attention features of the infrared image features are represented as X. The mapping process is:
[0057] a) Input the spatial attention feature Z of the visible light image feature into the mapping function composed of FC and Sigmoid to obtain the mapping weight vector ;
[0058] b) Input the spatial attention feature X of the infrared image feature into another set of mapping functions composed of FC and Sigmoid to obtain the mapping weight vector ;
[0059] c) Map the weight vector and the mapping weight vector Combine them using dot products to get weighted weights: .
[0060] d) Combine the weighted weights and the convolution parameters of the convolution operator group through a dot product operation, and update the convolution parameters to: ,in is the weighted weight vector A single value in .
[0061] Example 3
[0062] This embodiment is based on embodiment 2:
[0063] This embodiment provides a method for detecting and identifying visible light and infrared fusion based on adaptive essential features, which specifically includes the following steps:
[0064] First, pixel-level aligned visible light and infrared image pairs are obtained, and the locations of the objects of interest in the images are marked with rectangular boxes and annotated with categories for training a visible light and infrared fusion model with adaptive essential features.
[0065] Build a visible light infrared fusion model that is adaptive to essential features. First, prepare the environment required for model building and model training. This embodiment uses the coding and running environment of Docker+VScode+PyTorch to implement a visible light infrared fusion model that is adaptive to essential features. The model is based on an end-to-end target detection and recognition network structure, and is mainly composed of a backbone network, a connection Neck, a detection head and other parts. In visible light infrared fusion detection and recognition, a backbone network with a similar structure is first used to perform single-source extraction of visible light and infrared image features respectively. The backbone network of this embodiment adopts the public Darknet53 backbone, which specifically includes modules such as residuals and spatial pooling pyramids.
[0066] Then, a conditional convolution-based dynamic attention fusion module is used to fuse the visible light or infrared features of a single source together using an X-shaped multi-layer feature interaction connection. This conditional convolution-based dynamic attention fusion module consists of operators such as full connectivity, sigmoid activation function, and convolution. In this embodiment, the convolution kernel size is 3*3, with a total of four groups, each containing one convolution. After constructing this conditional convolution-based dynamic attention fusion module, it is embedded in different positions between the visible light and infrared single-source trunks. The visible light and infrared features are fused and re-input into the trunk of the other source, forming an X-shaped feature interaction fusion structure.
[0067] The connection neck uses an improved U-shaped recalculation bidirectional jump pyramid structure. This differs from the original recalculation pyramid structure in that the original U-shaped feature transfer loop is expanded to N U-shaped loops, and features are connected in a hybrid sequential and jump-level manner, rather than layer-by-layer, to better integrate visible and infrared features. In this embodiment, N = 3.
[0068] Paired labeled visible and infrared images are fed into the intrinsic feature-adaptive visible-infrared fusion model for training. During training, the commonly used training parameters in YOLOv5 are used, combined with maximum suppression as post-processing. The trained model is saved and tested on test image pairs to obtain the final comprehensive object detection and recognition results.
[0069] This adaptive visible-light infrared fusion model leverages the color and texture characteristics of visible light, its high image resolution, and the ability of infrared images to effectively capture images in low-light conditions and at night. It adaptively learns multi-scale compatible fusion features from visible and infrared images, automatically adjusts feature weights based on their specific conditions, and performs multi-scale detection and recognition, improving fusion recognition accuracy. This model effectively addresses issues such as poor single-source target detection and recognition due to deteriorating environmental conditions such as lighting and visibility, and insufficient target feature fusion caused by the trained model's inability to dynamically adapt to input images of varying quality.
[0070] Example 4
[0071] This embodiment is based on embodiment 2:
[0072] This embodiment provides a computer device including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the adaptive visible light and infrared fusion detection and recognition method according to the essential feature of Embodiment 1. The computer program may be in source code form, object code form, an executable file, or some intermediate form.
[0073] Example 5
[0074] This embodiment is based on embodiment 2:
[0075] This embodiment provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the adaptive visible light infrared fusion detection and recognition method, which is an essential feature of Example 1. The computer program may be in source code form, object code form, executable file, or some intermediate form. The storage medium includes: any entity or device capable of carrying computer program code, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in the storage medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the storage medium does not include electric carrier signal and telecommunication signal.
[0076] It should be noted that, for the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
Claims
1. A method for detecting and identifying visible light and infrared fusion with adaptive essential features, characterized in that: include: The features of images of different modalities at different scales are extracted through a preset X-shaped multi-layer feature interactive fusion module, and features with differences greater than a threshold are subjected to X-shaped interactive fusion to form a comprehensive feature sensitive to the target feature; the modalities include visible light and infrared; The dynamic attention fusion module based on conditional convolution and splicing operations amplifies the feature differences between the modalities, calculates the compensation and compatible features between the modalities, and fuses them with the original features of the corresponding modalities to obtain fused features that are consistent with the target features. A modal fusion feature analysis module based on a heavy-computation bidirectional jump pyramid connects the fusion features of different scales in sequence and cross-layer hybrid connection to perform target category classification and position regression, thereby obtaining fusion detection and recognition results; In the modal fusion feature analysis module based on the heavy-computation bidirectional skip-connection pyramid, the original bidirectional pyramid structure is expanded from the original U-shaped feature transfer loop to N U-shaped feature transfer loops, the connection between the backbone network and the detection head is deepened, and the learning and extraction of multi-scale features are realized; the layer-by-layer connection of features in the original bidirectional pyramid structure is optimized to a mixed connection of sequential and skip levels, expanding the combination types of multi-scale features.
2. The essential feature adaptive visible light infrared fusion detection and recognition method according to claim 1 is characterized in that: The method of extracting features of images of different modalities at different scales through a preset X-shaped multi-layer feature interactive fusion module includes: using two residual networks composed of residual blocks to respectively extract features of visible light and infrared images at different scales.
3. The essential feature adaptive visible light infrared fusion detection and recognition method according to claim 2 is characterized in that: A dynamic attention fusion module based on conditional convolution and splicing operations is embedded in each layer of the residual network of the X-type multi-layer feature interaction fusion module, and the similarities and differences of the features of different modal data in each layer are learned.
4. The essential feature adaptive visible light infrared fusion detection and recognition method according to claim 1 is characterized in that: In the dynamic attention fusion module based on conditional convolution and splicing operations, the conditional convolution calculates the weighted convolution kernel for the input sample before performing the convolution calculation. Each convolution kernel is calculated once and acts on the positions of different data.
5. A self-adaptive visible light and infrared fusion detection and recognition system, characterized by: include: An X-shaped multi-layer feature interaction fusion module is configured to extract features of images of different modalities at different scales and perform X-shaped interaction fusion on features with differences greater than a threshold to form a comprehensive feature sensitive to the target feature; the modalities include visible light and infrared; The dynamic attention fusion module based on conditional convolution and splicing operations is configured to amplify the feature differences between the modalities, calculate the compensation and compatibility features between the modalities, and fuse them with the original features of the corresponding modality to obtain fused features that are consistent with the target features; A modal fusion feature analysis module based on a heavy-computation bidirectional jump pyramid is configured to sequentially and cross-layer hybridly connect the fusion features of different scales to perform target category classification and position regression, thereby obtaining a fusion detection and recognition result; In the modal fusion feature analysis module based on the heavy-computation bidirectional skip-connection pyramid, the original bidirectional pyramid structure is expanded from the original U-shaped feature transfer loop to N U-shaped feature transfer loops, the connection between the backbone network and the detection head is deepened, and the learning and extraction of multi-scale features are realized; the layer-by-layer connection of features in the original bidirectional pyramid structure is optimized to a mixed connection of sequential and skip levels, expanding the combination types of multi-scale features.
6. The essential feature adaptive visible light infrared fusion detection and recognition system according to claim 5 is characterized in that: The X-type multi-layer feature interactive fusion module extracts features of images of different modalities at different scales, including: using two residual networks composed of residual blocks to respectively extract features of visible light and infrared images at different scales.
7. The essential feature adaptive visible light infrared fusion detection and recognition system according to claim 6 is characterized in that: A dynamic attention fusion module based on conditional convolution and splicing operations is embedded in each layer of the residual network of the X-type multi-layer feature interaction fusion module, and the similarities and differences of the features of different modal data in each layer are learned.
8. The essential feature adaptive visible light infrared fusion detection and recognition system according to claim 5 is characterized in that: In the dynamic attention fusion module based on conditional convolution and splicing operations, the conditional convolution calculates the weighted convolution kernel for the input sample before performing the convolution calculation. Each convolution kernel is calculated once and acts on the positions of different data.
Citation Information
Patent Citations
Voice emotion recognition method and system based on cross-layer cross fusion
CN114898775A
Dynamic feature assisted visible light fire detection and identification method and device and medium
CN115841642A