Guiding registration method, system and device for RGB and infrared image fusion and storage medium
By combining a multi-scale feature extraction network and a semantic segmentation enhancement module with a cross-modal feature matching engine, the problems of high computational cost and mismatch in infrared and visible light image registration are solved, achieving high-precision image alignment in complex scenes.
Patent Information
- Application Number
- CN202511308389.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-15
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-15
AI Technical Summary
Existing infrared and visible light image registration methods are computationally intensive and ineffective in image fusion with large intensity differences. They are also prone to errors during feature matching, making it difficult to achieve accurate registration.
A multi-scale feature extraction network and a semantic segmentation enhancement module are employed. Spatial attention weights are generated through object detection boxes for pixel-level segmentation. Combined with a cross-modal feature matching engine, coarse-grained and fine-grained features are extracted. A global matcher and distortion thinning are used to generate a homography matrix for image registration.
It significantly improves the accuracy and efficiency of image registration, reduces background interference, and is suitable for image alignment tasks in complex scenes, achieving accurate alignment of infrared and visible light images.
Smart Images

Figure CN121095299A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of image enhancement based on computer vision, and particularly relates to a guided registration method, system, device and storage medium for RGB and infrared image fusion. BACKGROUND
[0002] Under the condition of thick fog and low illumination, infrared cameras can capture the thermal radiation of objects for imaging to effectively detect the target. However, the quality of infrared images is not as good as that of visible light images, and visible light images contain more rich scene information but are easily affected by the light condition. Infrared and visible light fusion technology combines the advantages of the two together, and fuses the information in the two images into one image so that the infrared and visible light information can be obtained from a single image, which promotes the development of computer vision tasks. However, if the infrared and visible light images are not accurately registered, the fusion effect may not be satisfactory, so it is very crucial to find a technology that can accurately register infrared and visible light images.
[0003] The region-based method and the feature-based method are the existing two image registration methods. The region-based registration algorithm is mainly based on the reference image, and the correlation of the two images is found in the reference image by various similarity measures in the to-be-registered image. This method has been widely used in the registration of multi-modal images with intensity differences. However, when these methods are faced with infrared and visible light images with large modal differences, local extreme values may occur, and a large amount of calculation is required. The feature-based method can continuously iterate by finding the optimal similarity of the two images, and the feature information of the image is pixel-level expanded to obtain the corner points, edge segments, feature regions, etc. of the image, and then the feature points of the two images are matched by the matching algorithm.
[0004] For the problem of feature matching, many scholars have carried out research on extracting the same or similar features from two different images. Among them, the classic ones are scale-invariant feature transform (SIFT) and faster but less accurate speeded up robust features (SURF). However, feature matching is to extract the stronger features or the features with obvious differences in the image, and the feature points extracted from the images with unobvious intensity differences are very few or almost none, which leads to few matching point pairs and even error matching when performing feature matching, so the final matching effect is not ideal. SUMMARY
[0005] In view of the above problems, the present application provides a guided registration method for RGB and infrared image fusion in the first aspect, which comprises the following steps: Step 1, real-time receiving RGB visible light image r and corresponding infrared image t; Step 2, based on the input of visible light image r and infrared image t, identifying and detecting the main target object in the image, and generating a preliminary target detection box; Step 3, taking the target detection box as spatial prior information, converting the target box into spatial attention weight through prompt coding mechanism, guiding focusing on specific area for pixel-level segmentation, and generating RGB image segmentation result sr and infrared image segmentation result st; Step 4, performing feature matching on the segmented sr and st, first extracting coarse-grained and fine-grained features through two parallel encoders, coarse-grained feature encoder capturing global structural information of the image, and fine-grained feature encoder focusing on local texture details, coarse-grained features passing through global matcher to obtain coarse-grained mapping and confidence from the overall scene structure, fine-grained features and coarse-grained mapping and confidence obtained by global matcher being further filtered through distortion refinement to obtain dense mapping and its confidence, then selecting a group of reliable matches through matching sampling to generate homography matrix, and finally performing homography transformation on the obtained homography matrix and original input image t to obtain the final registration image ft.
[0006] Preferably, the step 2 adopts a multi-scale feature extraction network, which extracts features from the input image through a series of convolutional layers and attention modules, and outputs the final target coordinate information through three detection heads; the main part of the network is composed of multiple convolutional layers and C3k2 modules stacked alternately, forming a deep feature extraction network; and the feature pyramid structure realizes the fusion of features of different scales through upsampling and feature connection operations, and finally processes feature maps of different scales through three detection heads respectively, and outputs the target positioning result, i.e. the target detection box.
[0007] Preferably, the multi-scale feature extraction network specifically comprises: The input image r and t are received, and preliminary features are extracted through two convolutional layers; then, through a backbone network, the feature maps are gradually down-sampled to different scales to extract multi-level features, and the backbone network is composed of four groups of "Conv+C3k2" modules; the network end includes an SPPF module to expand the receptive field through multi-scale pooling, and a C2PSA module to optimize deep features by applying an attention mechanism; in the feature pyramid part, the deep features are first up-sampled and fused with the middle-level features, and the fused features are again up-sampled and fused with the shallow-level features, and after each fusion, the C3k2 module is used to process and enhance the feature expression; three parallel detection heads are also included, which process features of different scales: Head1 is responsible for deep features for detecting large targets, Head2 processes middle-level features for medium targets, and Head3 processes shallow-level features for small target detection; finally, the prediction results of the three detection heads are integrated through the Coordinates module, and the NMS algorithm is applied to eliminate repeated detection, and the position coordinates, size and category information of the target are output.
[0008] Preferably, step 3 uses a semantic segmentation enhancement module to obtain the target detection box, extract rich visual features through an image encoder, then process the user-provided interaction information through a mask decoder and a prompt decoder, and finally obtain an accurate segmentation mask; specifically including: Rich visual features are extracted through a multi-layer image encoder structure to generate multi-scale image embeddings, which are dense feature maps containing local and global semantic information of the image; at the same time, user interaction prompts are received from the left, and the prompts are first processed by the prompt decoder to convert them into standardized prompt embeddings; the prompt embeddings are then sent to the mask decoder, which associates the prompt information with the image features through an attention mechanism to generate preliminary mask features; the mask features and the mask input from the top are processed by the convolution layer conv, and then fused with the image embeddings to produce enhanced mask representations; based on the fused features, the mask decoder generates the final segmentation mask through multi-head attention and a feedforward network to identify the image region corresponding to the user prompt.
[0009] Preferably, the global matcher in step 4 is specifically: Based on the coarse-grained features obtained by the encoder, key points are detected in two images and descriptors are generated, the similarity matrix between the descriptors of the two images is calculated to obtain the point in the other image that has the highest similarity to each feature point in one image, and finally the coarse-grained mapping and confidence are obtained.
[0010] Preferably, the warping refinement in step 4 is based on the fine-grained features from the encoder and the coarse-grained mapping and confidence from the global matcher to estimate a global geometric transformation, and then for each matched point in one image, the position of the corresponding point in the other image is fine-tuned, and the global geometric transformation and the matched position are iteratively transformed until convergence to obtain a dense mapping and its confidence.
[0011] Preferably, the matching sampling in step 4 is based on the dense mapping and its confidence from the warping refinement to remove the matches with low confidence, divide the image into grids, and randomly keep one match in each grid to avoid local clustering, and finally output the final matched point pairs and calculate the homography matrix from the final matched point pairs.
[0012] The second aspect of the present application provides a guided registration system for RGB and infrared image fusion, comprising a receiving module, a multi-scale feature extraction network module, a semantic segmentation enhancement module and a cross-modal feature matching engine network module. The receiving module is used to receive an RGB visible light image r and a corresponding infrared image t. The multi-scale feature extraction network module, based on the input of the visible light image r and the infrared image t, identifies and detects the main target objects in the image to generate a preliminary target detection box. The semantic segmentation enhancement module takes the target detection box as spatial prior information, converts the target box into spatial attention weights through a prompt encoding mechanism, guides focusing on a specific area for pixel-level segmentation, and generates an RGB image segmentation result sr and an infrared image segmentation result st. The cross-modal feature matching engine network module performs feature matching on the segmented sr and st, first extracts coarse-grained and fine-grained features through two parallel encoders, the coarse-grained feature encoder captures the global structural information of the image, while the fine-grained feature encoder focuses on local texture details, the coarse-grained features pass through a global matcher to obtain coarse-grained mapping and confidence from the overall scene structure, and the fine-grained features and the coarse-grained mapping and confidence obtained by the global matcher are further screened through warping refinement to obtain dense mapping and its confidence, then a group of reliable matches are selected through matching sampling to generate a homography matrix, and finally the obtained homography matrix is subjected to homographic transformation with the original input image t to obtain a final registration image ft.
[0013] The third aspect of the present application also provides a guided registration device for RGB and infrared image fusion, the device comprising at least one processor and at least one memory, the processor and the memory being coupled; the memory stores a computer execution program; when the processor executes the computer execution program stored in the memory, the processor executes the guided registration method for RGB and infrared image fusion as described in the first aspect.
[0014] The fourth aspect of the present application also provides a computer readable storage medium, wherein a computer execution program is stored in the computer readable storage medium, and the computer execution program, when executed by a processor, causes the processor to execute the guided registration method for RGB and infrared image fusion according to the first aspect.
[0015] Compared with the prior art, the present application has the following beneficial effects: The semantic segmentation enhancement module effectively removes or fades the background information irrelevant to the target, significantly enhances the shape features of the target to be registered, and the process not only improves the efficiency of feature extraction, but also significantly increases the number of extractable feature points, so that the image matching is more clean and accurate.
[0016] The newly proposed cross-modal feature matching engine strategy of the present application extracts coarse-grained and fine-grained features through two parallel encoders, the coarse-grained feature encoder can capture the global structure information of the image, and the fine-grained feature encoder focuses on local texture details, comprehensively extracts image features, and the two types of features are subjected to global matcher, distortion refinement and matching sampling to obtain a homography matrix, and based on the input original infrared image and the homography matrix, a final registration image is obtained through homography transformation, which is superior to the prior art in registration accuracy and stability, and provides a robust solution for the registration of infrared and visible light images.
[0017] The present application combines a multi-scale feature extraction network, a semantic segmentation enhancement module and a cross-modal feature matching engine network to realize accurate alignment of images in space, and through the cooperative work of multiple modules, the present application can effectively reduce background interference and improve the accuracy and efficiency of image registration, and is suitable for image alignment tasks in complex scenes. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1 The overall logic diagram of the guided registration method of the present application.
[0019] Figure 2 The multi-scale feature extraction network structure diagram of the present application.
[0020] Figure 3 The semantic segmentation enhancement module structure diagram of the present application.
[0021] Figure 4 The visualized experimental result diagram in the embodiment of the present application.
[0022] Figure 5 The structure diagram of the guided registration device of the present application. DETAILED DESCRIPTION
[0023] The application mainly provides a guided registration method and system structure for RGB and infrared image fusion, and the specific implementation logic is as shown in Figure 1 The method includes the following processes: Step 1, real-time receiving an RGB visible light image r and a corresponding infrared image t; Step 2, based on the input of the visible light image r and the infrared image t, identifying and detecting the main target object in the image to generate a preliminary target detection frame; Step 3, taking the target detection frame as spatial prior information, converting the target frame into spatial attention weight through a prompt coding mechanism, guiding focusing on a specific area for pixel-level segmentation to generate an RGB image segmentation result sr and an infrared image segmentation result st; Step 4, performing feature matching on the segmented sr and st, first extracting coarse-grained and fine-grained features through two parallel encoders, the coarse-grained feature encoder captures the global structure information of the image, while the fine-grained feature encoder focuses on local texture details, the coarse-grained features are subjected to a global matcher to obtain coarse-grained mapping and confidence from the overall scene structure, the fine-grained features and the coarse-grained mapping and confidence obtained by the global matcher are subjected to distortion and refinement to further screen dense mapping and its confidence, then a group of reliable matches is selected through matching sampling to generate a homography matrix, and finally the obtained homography matrix is subjected to homography transformation with the original input image t to obtain a final registration image ft.
[0024] To realize the above method, the application provides a registration system mainly including a multi-scale feature extraction network module, a semantic segmentation enhancement module and a cross-modal feature matching engine network module, through the cooperative work of the multi-modules, background interference can be effectively reduced, the accuracy and efficiency of image registration can be improved, and the registration system is suitable for image alignment tasks in complex scenes.
[0025] The application will be further described in combination with specific embodiments.
[0026] I. Multi-scale feature extraction network module The network module structure is as shown in Figure 2As shown, the module receives an RGB image of standard size 640x640x3. The input image first goes through a series of convolutional layers and down-sampling operations to form a backbone network for feature extraction. The backbone network adopts a CSPDarknet structure, which is composed of multiple convolutional layers and residual connections. The first convolutional layer uses 32 3x3 filters (with a stride of 2) to convert the input image into a feature map of 320x320x32. Subsequently, the image goes through a second convolutional layer (64 3x3 filters) to further extract features. Then, the feature map goes through a first down-sampling operation (labeled "Stride 2") with a stride of 2 to reduce the resolution to 160x160x128. The feature map continues to be processed by a third convolutional layer, and then goes through a second down-sampling operation with a stride of 2 to reduce the resolution to 80x80x256. Subsequently, the feature map goes through a fourth convolutional layer, and then goes through a third down-sampling operation with a stride of 2 to reduce the resolution to 40x40x512. The feature map continues to be processed by a fifth convolutional layer, and then goes through a fourth down-sampling operation with a stride of 2 to finally reduce the resolution to 20x20x1024.
[0027] At the end of the backbone network, the model applies a spatial pyramid pooling (labeled "Spatial Pyramid Pooling") module to expand the receptive field through multi-scale pooling operations and enhance the feature expression capability. Next, a self-attention mechanism is applied to the feature map to further optimize the feature representation by calculating the relationship between feature channels, outputting an enhanced feature map of 20x20x1024. The Pyramid Attention Network (PANet) part starts to process multi-scale feature fusion. First, the deepest layer of features (20x20x1024) is enlarged to 40x40 through an up-sampling operation (labeled "Up-sampling"), and is fused with the features (40x40x512) of the corresponding layer in the backbone network through a connection layer (labeled "Connection Layer"). The fused features go through a convolutional layer with a stride of 2 to generate 40x40x512 mid-level features. These mid-level features are again up-sampled to 80x80 and fused with the features (80x80x256) of the corresponding layer in the backbone network through another connection layer. The fused features go through a convolutional layer to generate 80x80x256 shallow-level features.
[0028] The detection head part contains three parallel branches. Detection head 1 processes deep features (20x20x1024), focusing on detecting large-sized objects; detection head 2 processes middle features (40x40x512), aiming at medium-sized objects; detection head 3 processes shallow features (80x80x256), responsible for detecting small-sized objects. Each detection head contains multiple convolutional layers and an output layer. Detection head 1 outputs a tensor of 20x20x(4+1+C), where 4 represents the bounding box coordinates (x, y, width, height), 1 represents the object confidence, and C represents the number of categories. Detection head 2 outputs a tensor of 40x40x(4+1+C), and detection head 3 outputs a tensor of 80x80x(4+1+C).
[0029] Finally, the outputs of the three detection heads are integrated and the Non-Maximum Suppression (NMS) algorithm (IoU threshold is usually 0.45) is applied to eliminate duplicate detections, generating the final detection results.
[0030] II. Semantic Segmentation Enhancement Module The module structure is shown in Figure 3 The processing flow starts from the right side of the image encoder. The module receives the input image (usually a 1024x1024x3 RGB image). The image encoder is based on the Vision Transformer (ViT) architecture, consisting of 12 Transformer blocks, each containing a multi-head self-attention mechanism (12 attention heads, each with 64 dimensions) and a feed-forward neural network. The decoder first divides the input image into 16x16 image blocks (patches), generating 64x64 patches, each of which is linearly projected into a 768-dimensional embedding space. These patch embeddings are then enhanced with positional encoding to add spatial position information. The enhanced embedding sequence is processed through the Transformer blocks to capture local and global contextual relationships in the image, and finally outputs a dense image embedding (Image Embedding) with a dimension of 64x64x256, represented as "Image Embedding" in the figure. At the same time, the model receives user interaction prompts from the left side, which can be clicks (positive and negative points), box selections, or text descriptions, collectively referred to as "images to be segmented". These prompts are first processed by the prompt decoder, which contains two layers of Transformer blocks to convert different types of prompts into standardized prompt embeddings with a dimension of Nx256, where N is the number of prompt points. The prompt embeddings are then fed into the mask decoder, which is a lightweight Transformer containing 2 self-attention layers and 2 cross-attention layers. The mask decoder associates the prompt information with the image features through cross-attention mechanisms to generate preliminary mask features.
[0031] Meanwhile, the mask input from the top (if any) is processed by a convolutional layer (Conv) to generate mask embeddings. The mask embeddings are fused with the image embeddings and the prompt embeddings through an additive operation (the "⊕" symbol in the figure) to produce enhanced mask representations. A mask decoder generates mask embeddings based on these fused features through a multi-head attention mechanism and a feedforward network, with a dimension of 256 x 256 x 1. These embeddings are converted into the final segmentation mask through upsampling (usually 4 times bilinear interpolation) with a dimension of 1024 x 1024 x 1, represented as the "segmented image" in the figure. Each pixel value in the mask ranges from 0 to 1, representing the probability that the pixel belongs to the target object.
[0032] III. Cross-modal feature matching engine network module Twist refinement processes features of different scales through two parallel paths. One is the fine-grained feature path, which extracts local texture details through a four-layer convolutional network (with channel numbers of 64, 128, 256, and 512) and outputs a 160 x 160 x 512 feature map. These features capture fine structures of the image, such as edges, corners, and texture changes; the other is the coarse-grained feature path, which processes features with larger receptive fields through a global matcher. The global matcher first enhances feature representation through a self-attention mechanism (8 attention heads, each with 64 dimensions), then calculates the phase correlation and confidence score across images. The phase correlation is represented as an 80 x 80 x 2 tensor, describing the pixel-level correspondence; the confidence score is an 80 x 80 x 1 tensor, used to identify reliable matching areas.
[0033] There is a bidirectional information flow between the two feature paths: fine-grained features guide the optimization of coarse-grained features, while the global consistency constraint of coarse-grained features also improves the matching accuracy of fine-grained features. This complementary mechanism is achieved through a feature fusion layer that integrates information of two scales using 1 x 1 convolution and residual connection. Two encoders process the optimized fine-grained and coarse-grained features, respectively, and each encoder contains three residual blocks, each composed of two 3 x 3 convolutional layers and a skip connection. The encoder outputs phase degree features, which encode pixel-level correspondence and deformation fields.
[0034] Where the global matcher processes the segmentation mask obtained from the semantic segmentation enhancement module, dense features (256-dimensional vector per pixel) are computed by the local feature extractor. The cross-correlation algorithm is used to compute the similarity matrix between feature points, and the soft-max function is used to generate the preliminary matching relationship. At the same time, this module also calculates the confidence score of each matching point (a scalar value between 0 and 1), which is used for subsequent filtering of unreliable matches. These estimated mappings and confidence are encoded as a 640x640x3 tensor, where the first two channels represent the mapping coordinates, and the third channel represents the confidence.
[0035] The matching sampling is based on dense mapping and confidence, and the global geometric transformation parameters are estimated by the weighted RANSAC algorithm. This module first selects the top N most reliable matching points (usually N=2000) based on the confidence, and then finds the best transformation model by iterative sampling (usually 5000 iterations). The calculated homography matrix is a 3x3 transformation matrix that describes the projection relationship from the source image to the target image.
[0036] The warp module uses the homography matrix to perform homographic transformation on the source image to obtain the final registration image. IV. Experimental results Table 1 and Table 2 show the results obtained by testing the present application on two different data sets, and the visualization of the experimental results is shown in Figure 4 The experimental results clearly show that the method proposed in the present application is always superior to other existing methods in various aspects. Specifically, the method of the present application shows more accurate performance in handling the registration task of infrared images and visible light images. It is worth noting that even under different shooting heights and angles, our method can still maintain high accuracy, ensuring accurate registration of images. This advantage enables our technology to be effectively applied in complex and variable environments, further proving its excellence and reliability in the field of image processing.
[0037] Table 1 Registration effect on public data set
[0038] Table 2 Registration effect on private data set
[0039] As Figure 5As shown, the application also provides a guided registration device for RGB and infrared image fusion, which comprises at least one processor and at least one memory, and further comprises a communication interface and an internal bus; the memory stores a computer-executable program; the memory stores a computer-executable program; when the processor executes the computer-executable program stored in the memory, the processor can execute a guided registration method for RGB and infrared image fusion. The internal bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component (PCI) bus, or an.Xtended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, the bus in the drawings of the present application does not limit to only one bus or one type of bus. The memory can include a high-speed RAM memory, and can also include a non-volatile storage NVM, for example, at least one disk memory, and can also be a U disk, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0040] The device can be provided as a terminal, a server or other forms of devices. In the exemplary embodiments, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic elements for executing the above-mentioned method.
[0041] The application also provides a computer-readable storage medium, which stores a computer-executable program, and when the computer-executable program is executed by a processor, the processor can execute a guided registration method for RGB and infrared image fusion.
[0042] Specifically, a system, device or equipment provided with a readable storage medium can be provided, and the readable storage medium stores software program codes for realizing the functions of any one of the above-mentioned embodiments, and the computer or processor of the system, device or equipment reads out and executes the instructions stored in the readable storage medium. In this case, the program codes read from the readable medium can realize the functions of any one of the above-mentioned embodiments, and therefore the machine-readable codes and the readable storage medium storing the machine-readable codes constitute a part of the application.
[0043] The above merely describes preferred embodiments of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, and the like made within the principle and technical scope of the present application should be included in the protection scope of the present application.
[0044] Although the specific embodiments of the present application are described above, the present application is not limited to the above. It should be understood by those skilled in the art that various modifications or changes can be made to the technical solutions of the present application without creative labor and still fall within the protection scope of the present application.
Claims
1. A guided registration method for RGB and infrared image fusion, characterized in that, Includes the following steps: Step 1: Receive and acquire the RGB visible light image r and the corresponding infrared image t in real time; Step 2: Based on the input of the visible light image r and the infrared image t, identify and detect the main target objects in the images, and generate preliminary target detection boxes; Step 3: Using the target detection bounding box as spatial prior information, the target bounding box is transformed into spatial attention weights through a cue coding mechanism, guiding the focus to a specific region for pixel-level segmentation, generating RGB image segmentation result sr and infrared image segmentation result st; Step 4: Perform feature matching on the segmented sr and st. First, extract coarse-grained and fine-grained features using two parallel encoders. The coarse-grained feature encoder captures the global structural information of the image, while the fine-grained feature encoder focuses on local texture details. The coarse-grained features are processed by a global matcher to obtain coarse-grained mappings and confidence scores from the overall scene structure. The coarse-grained mappings and confidence scores obtained from the fine-grained features and the global matcher are further filtered through distortion and thinning to obtain dense mappings and their confidence scores. Then, a set of reliable matches is selected through matching sampling to generate a homography matrix. Finally, the obtained homography matrix is homography transformed with the original input image t to obtain the final registered image ft.
2. The guided registration method for RGB and infrared image fusion as described in claim 1, characterized in that: Step 2 employs a multi-scale feature extraction network. Starting from the input image, features are extracted step by step through a series of convolutional layers and attention modules, and the final target coordinate information is output through three detection heads. The backbone of the network is composed of multiple convolutional layers and C3k2 modules stacked alternately to form a deep feature extraction network. The feature pyramid structure achieves the fusion of features at different scales through upsampling and feature connection operations. Finally, the feature maps at different scales are processed by three detection heads to output the target localization result, i.e., the target detection box.
3. The guided registration method for RGB and infrared image fusion as described in claim 2, characterized in that, The multi-scale feature extraction network is specifically as follows: The system receives input images r and t, extracts preliminary features through two convolutional layers, and then progressively downsamples the feature maps to different scales through a backbone network to extract multi-level features. The backbone network consists of four "Conv+C3k2" modules. The network ends include an SPPF module, which expands the receptive field through multi-scale pooling, and a C2PSA module, which applies an attention mechanism to optimize deep features. In the feature pyramid, deep features are first upsampled and fused with mid-level features, and the fused features are then upsampled again and fused with shallow features. After each fusion, the C3k2 module is used to enhance the feature representation. The system also includes three parallel detection heads, each processing features at different scales: Head1 handles deep features for detecting large targets, Head2 handles mid-level features for detecting medium-sized targets, and Head3 handles shallow features for detecting small targets. Finally, the Coordinates module integrates the prediction results of the three detection heads, applies the NMS algorithm to eliminate duplicate detections, and outputs the target's location coordinates, size, and category information.
4. The guided registration method for RGB and infrared image fusion as described in claim 1, characterized in that: Step 3 employs a semantic segmentation enhancement module. Based on the obtained target detection box, rich visual features are extracted by an image encoder. Subsequently, the interactive information provided by the user is processed by a mask decoder and a prompt decoder to finally obtain an accurate segmentation mask. Specifically, it includes: A multi-layer image encoder structure extracts rich visual features to generate multi-scale image embeddings. These image embeddings are dense feature maps containing both local and global semantic information of the image. Simultaneously, user interaction prompts are received from the left side. These prompts are first processed by a prompt decoder and converted into standardized prompt embeddings. The prompt embeddings are then fed into a mask decoder, which associates the prompt information with image features through an attention mechanism to generate preliminary mask features. These mask features, along with the mask input from the top, are processed through a convolutional layer (conv) and then fused with the image embeddings to produce an enhanced mask representation. Based on the fused features, the mask decoder generates the final segmentation mask through multi-head attention and a feedforward network, identifying the image region corresponding to the user prompt.
5. The guided registration method for RGB and infrared image fusion as described in claim 1, characterized in that: The global matcher in step 4 is specifically: Based on the coarse-grained features obtained from the encoder, key points are detected in two images and descriptors are generated. By calculating the similarity matrix between the descriptors of the two images, the point with the highest similarity in one image to the other image is obtained. Finally, the coarse-grained mapping and confidence are obtained.
6. The guided registration method for RGB and infrared image fusion as described in claim 1, characterized in that: The distortion thinning in step 4 is based on the fine-grained features obtained by the encoder and the coarse-grained mapping and confidence obtained by the global matcher to estimate a global geometric transformation. Then, for each matching point in one image, the position of the corresponding point is fine-tuned in another image. The global geometric transformation and matching position are iteratively transformed until the dense mapping and its confidence are converged.
7. The guided registration method for RGB and infrared image fusion as described in claim 1, characterized in that: The matching sampling in step 4 is based on the dense mapping and confidence obtained by distortion thinning to eliminate matches with low confidence. The image is divided into grids, and one match is randomly retained in each grid to avoid local clustering. Finally, the final matching point pairs are output, and the homography matrix is calculated from the final matching point pairs.
8. A guided registration system for RGB and infrared image fusion, characterized in that: It includes a receiving module, a multi-scale feature extraction network module, a semantic segmentation enhancement module, and a cross-modal feature matching engine network module; The receiving module is used to receive and acquire the RGB visible light image r and the corresponding infrared image t; The multi-scale feature extraction network module, based on the input of visible light image r and infrared image t, identifies and detects the main target objects in the image and generates a preliminary target detection box; The semantic segmentation enhancement module uses the target detection box as spatial prior information. Through a cue coding mechanism, the target box is transformed into spatial attention weights, which guide the focus on a specific region for pixel-level segmentation, generating RGB image segmentation result sr and infrared image segmentation result st. The cross-modal feature matching engine network module performs feature matching on the segmented sr and st. First, coarse-grained and fine-grained features are extracted by two parallel encoders. The coarse-grained feature encoder captures the global structural information of the image, while the fine-grained feature encoder focuses on local texture details. The coarse-grained features are processed by a global matcher to obtain coarse-grained mappings and confidence scores from the overall scene structure. The coarse-grained mappings and confidence scores obtained by the fine-grained features and the global matcher are further filtered through distortion and thinning to obtain dense mappings and their confidence scores. Then, a set of reliable matches is selected through matching sampling to generate a homography matrix. Finally, the obtained homography matrix is subjected to homography transformation with the original input image t to obtain the final registered image ft.
9. A guided registration device for RGB and infrared image fusion, characterized in that: The device includes at least one processor and at least one memory, the processor and the memory being coupled together; the memory stores a computer-executable program; when the processor executes the computer-executable program stored in the memory, the processor performs a guided registration method for RGB and infrared image fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer-executable program, which, when executed by a processor, causes the processor to perform a guided registration method for RGB and infrared image fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Lightweight image registration method, system and device based on global homography estimation
CN114565511A
Semantic segmentation method for power accessories of power transmission line
CN115376024A
Photovoltaic defect detection method based on infrared and visible light image fusion
CN115578378A
Night thermal infrared image semantic segmentation enhancement method based on improved ResNet
CN115601723A
Infrared and visible light image fusion method based on cross mode enhancement and multi-attention fusion strategy
CN118096554A
Cited By
Cross-modal semantic segmentation system and method
CN121482790A