Projectile drop point detection method, system and equipment based on dual-light image fusion and medium
Through the dual-light image fusion technology, visible light and thermal infrared imaging data are used for spatial and fusion of space-time registration and pre-detection, the problem of low positioning accuracy of projectile landing points in the prior art is solved, and efficient and accurate projectile landing points detection is achieved.
Patent Information
- Application Number
- CN202510407704.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
In the outdoor long-distance large-diameter projectile training scenario, the positioning accuracy of the landing point is low and the efficiency of manual target detection is low. The existing methods have problems such as complex, high cost, easy to be disturbed or false alarm rate.
Using a method based on dual-light image fusion, we collect visible light and thermal infrared imaging data, perform spatiotemporal registration and pre-detection fusion to generate high-quality fusion images, use typical target detection models to judge explosion phenomena, and use landing solution algorithm to determine landing locations.
It improves the accuracy and efficiency of projectile landing point detection, reduces complex background interference, optimizes the target detection results, and provides reliable landing point positioning.
Smart Images

Figure CN120339389A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of impact point detection, and in particular to a method, system, device and medium for detecting the impact point of a projectile based on dual-light image fusion. Background Art
[0002] In the field of shooting detection, especially in the outdoor long-distance large-caliber projectile training scenario, for the positioning of the impact point, there are currently methods such as manual target inspection and sighting target inspection, which have poor accuracy and low efficiency. Other target reporting methods that do not require human participation, such as radar, sound waves, and light curtains, each have their limitations. Radar technology is relatively complex, vulnerable to electromagnetic interference, and has a high cost; the sound wave detection method is greatly affected by factors such as terrain and noise, and the positioning accuracy is largely affected by the layout plan; the light curtain has problems such as complex layout and many interferences in a large field of view. The image-based method is based on the projectile explosion phenomenon and conforms to human intuition. The traditional binocular positioning method has high requirements for the site and complex layout; the image detection method of single visible light or thermal infrared is affected by conditions such as the size and shape of the projectile explosion, with a high false alarm rate or prone to missing detections, making it difficult to accurately detect or locate the impact point of the projectile. For example, the existing patent CN202210189027.4 discloses a method based on a visible light camera using traditional image processing methods, but this method is greatly affected by environmental factors and background complexity in actual applications; the prior art CN202311144943.7 discloses a method based on binocular vision to automatically detect the projectile explosion area and obtain the coordinates of the projectile impact point in the image, and calculate the coordinates and distance of the projectile impact point relative to the target; however, this prior art has a complex layout, and the selection of key parameters such as camera calibration, baseline length measurement, and angle measurement has a great impact on the results, and is greatly affected by environmental factors.
[0003] Therefore, it is urgent to solve the above problems. Summary of the Invention
[0004] Object of the Invention: The first object of the present invention is to provide a method for detecting the impact point of a projectile based on dual-light image fusion, which improves the detection accuracy and better completes the target reporting task.
[0005] The second object of the present invention is to provide a system for detecting the impact point of a projectile based on dual-light image fusion.
[0006] The third object of the present invention is to provide an electronic device.
[0007] The fourth object of the present invention is to provide a computer-readable storage medium.
[0008] Technical Solution: To achieve the above objects, the present invention discloses a method for detecting the impact point of a projectile based on dual-light image fusion, including the following steps:
[0009] S1. Collect visible light imaging data and thermal infrared imaging data of all target areas;
[0010] S2. Preprocess the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas;
[0011] S3. Project the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces;
[0012] S4. Perform pre-detection fusion on the registered visible light images and thermal infrared images to generate high-quality fusion images with both thermal saliency and visible light details, and then send the high-quality fusion images into a typical target detection model to judge frame by frame whether an explosion phenomenon is detected in the images, and obtain the bounding boxes corresponding to the explosion targets in the images;
[0013] S5. According to the explosion bounding boxes, use a landing point calculation algorithm to calculate and correct the position of the landing point, and determine the landing point image coordinates in the explosion area; according to the coordinate mapping relationship between the reference points on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface to obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface.
[0014] Optionally, the step S2 specifically includes the following steps: Configure the acquisition time sequence for the visible light imaging data and thermal infrared imaging data so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas.
[0015] Optionally, the step S3 specifically includes the following steps:
[0016] Pre-establish and store the mapping relationship of the position information between the ground target areas on the spatio-temporally registered visible light images and thermal infrared images and the reference projection target surface, and the position information is coordinate information;
[0017] Select a number of reference points in the target area in advance, the number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points show obvious features in both visible light and thermal infrared images, and the obvious features include but are not limited to morphological features, texture features, edge features or color features; or pre-manufacture a reference object for calibration, the surface of which has recognizable marking features, and select the corresponding feature points;
[0018] Record the relative position information of several reference points, and establish a reference projection target surface according to this information. Denote the position coordinates of the reference points on the reference projection target surface as the target position information; determine the image coordinate positions of several reference points in the target area image, denoted as the image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points between the target area image and the reference projection target surface.
[0019] Optionally, in step S4, a pre-detection fusion model is used for pre-detection fusion. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. Among them, the dual-branch feature extraction module uses a dual-branch structure to process the thermal infrared image and the visible light image respectively. The thermal infrared image extracts the shallow thermal infrared features through the first 2 convolutional layers, and then extracts the deep thermal infrared features through the last 3 convolutional layers. The visible light image extracts the shallow visible light image through the first 1 convolutional layer and the first 1 feature alignment module, and then extracts the deep visible light features through the last 1 feature alignment module and the last 2 convolutional layers. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the corresponding feature alignment modules to correct the spatial offset of the visible light image; each level of deep thermal infrared features and the corresponding deep visible light features are input into the feature interaction module together, and the output three groups of fused features are input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
[0020] Optionally, in the feature alignment module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input of the feature maps; then the offsets of the deformable convolution are dynamically generated through convolution operations; the obtained offsets are used as the sampling point offsets of the deformable convolution and applied to the visible light feature map to obtain the corrected visible light modality feature map.
[0021] Optionally, in the feature interaction module for the thermal infrared feature map and the visible light feature map, first splice them in the channel dimension to obtain a combined input of feature maps; introduce a set of random matrices with the same shape as the combined input feature maps, and perform Schmidt orthogonalization to obtain a set of orthogonalized filters; perform a pointwise multiplication operation between the orthogonalized filters and the combined input feature maps to obtain a one-dimensional orthogonalized weight vector; perform a one-dimensional convolution on the orthogonalized weight vector, and perform slicing and normalization operations on the convolution result to obtain attention weights for different modalities, denoted as the thermal infrared weight and the visible light weight respectively; multiply the original thermal infrared feature map by the thermal infrared weight to obtain a weighted thermal infrared feature, and add the weighted thermal infrared feature to the original thermal infrared feature map through a residual connection to obtain an orthogonalized thermal infrared feature map; multiply the original visible light feature map by the visible light weight to obtain a weighted visible light feature, and add the weighted visible light feature to the original visible light feature map through a residual connection to obtain an orthogonalized visible light feature map; finally, splice the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map and then perform a shared convolution operation to obtain a fused feature map.
[0022] Optionally, step S5 specifically includes the following steps: Obtain and store the bounding boxes and corresponding category information containing explosion events in consecutive frames, perform rate-of-change and stability analysis on the recorded bounding boxes and category information, calculate the change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames, and at the same time, by statistically analyzing the consecutive frame data, eliminate abnormal results caused by instantaneous noise or false detections; According to the bounding boxes and category information in the candidate frames, select the moment when a frame first satisfies the conditions of a stable bounding box and a clear category as the initial explosion moment, and extract the bounding box in this frame as the basis for impact point calculation; Use the determined bounding box to calculate the center point as the preliminary impact point image coordinates; With the help of additional information detected in multiple subsequent frames, supplement and correct the preliminary calculated impact point image coordinates, finally obtain the corrected impact point image coordinates, and output the corrected impact point image coordinates as the position basis corresponding to the initial explosion moment.
[0023] Based on the same inventive concept, the present invention discloses a projectile impact point detection system based on dual-light image fusion, including: a data acquisition module for acquiring visible light imaging data and thermal infrared imaging data of the entire target area;
[0024] A preprocessing module for preprocessing the visible light imaging data and the thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas;
[0025] A target area mapping module for projecting the ground target area on the spatio-temporally registered visible light image and thermal infrared image obtained onto the corresponding reference projection target surface;
[0026] The dual - wavelength fusion detection module is used to perform pre - detection fusion on the registered visible - light image and thermal infrared image to generate a high - quality fusion image with both thermal saliency and visible - light details, and then send the high - quality fusion image into a typical object detection model to judge frame by frame whether an explosion phenomenon is detected in the image, and obtain the bounding box corresponding to the explosion target in the image;
[0027] The impact point calculation module is used to calculate and correct the position of the impact point according to the explosion bounding box, and determine the image coordinates of the impact point in the explosion area; according to the coordinate mapping relationship between the reference points on the target area image and the reference projection target surface, map the image coordinates of the impact point to the reference projection target surface, obtain the position coordinate information of the impact point within the reference projection target surface, and record and output the relative position information of the impact point on the reference projection target surface.
[0028] Optionally, in the pre - processing module, the acquisition time sequence of the visible - light imaging data and the thermal infrared imaging data is configured so that each frame of visible - light image and thermal infrared image is synchronized in time, and then each frame of visible - light image and thermal infrared image is registered and aligned in space to obtain the visible - light image and thermal infrared image of the corresponding target area with spatio - temporal registration.
[0029] Optionally, in the target area mapping module, a mapping relationship of the position information between the ground target area on the spatio - temporally registered visible - light image and thermal infrared image and the reference projection target surface is established and stored in advance, and the position information is coordinate information;
[0030] Several reference points are selected in the target area in advance, the number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points show obvious features in both visible - light and thermal infrared images, and the obvious features include but are not limited to morphological features, texture features, edge features or color features; or a calibration reference object is made in advance, and it has recognizable marking features on its surface, and the corresponding feature points are selected;
[0031] Record the relative position information of several reference points, and establish a reference projection target surface according to this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points on the target area image and the reference projection target surface.
[0032] Optionally, a pre-detection fusion model is used for pre-detection fusion in the dual-spectrum fusion detection module. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. The dual-branch feature extraction module processes the thermal infrared image and the visible light image respectively using a dual-branch structure. The thermal infrared image passes through the first 2 convolutional layers to extract shallow thermal infrared features, and then passes through the subsequent 3 convolutional layers to extract deep thermal infrared features. The visible light image passes through the first 1 convolutional layer and the first 1 feature alignment module to extract the shallow visible light image, and then passes through the subsequent 1 feature alignment module and 2 convolutional layers to extract deep visible light features. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the corresponding feature alignment modules to correct the spatial offset of the visible light image. Each level of deep thermal infrared features and the corresponding deep visible light features are jointly input into the feature interaction module, and three groups of fused features are output and input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
[0033] Optionally, in the feature alignment module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input of the feature maps. Then, the offsets of the deformable convolution are dynamically generated through convolution operations. The obtained offsets are used as the sampling point offsets of the deformable convolution and applied to the visible light feature map to obtain the corrected visible light modality feature map.
[0034] Optionally, in the feature interaction module, for the thermal infrared feature map and the visible light feature map, they are first concatenated in the channel dimension to obtain a joint input of the feature maps. A set of random matrices with the same shape as the joint input feature maps is introduced and Schmidt orthogonalization is performed to obtain a set of orthogonalized filters. The orthogonalized filters are subjected to a pointwise multiplication operation with the joint input feature maps to obtain a one-dimensional orthogonalized weight vector. One-dimensional convolution is performed on the orthogonalized weight vector, and the convolution result is sliced and normalized to obtain attention weights for different modalities, denoted as the thermal infrared weight and the visible light weight respectively. The original thermal infrared feature map is multiplied by the thermal infrared weight to obtain a weighted thermal infrared feature, and the weighted thermal infrared feature and the original thermal infrared feature map are added together through a residual connection to obtain an orthogonalized thermal infrared feature map. The original visible light feature map is multiplied by the visible light weight to obtain a weighted visible light feature, and the weighted visible light feature and the original visible light feature map are added together through a residual connection to obtain an orthogonalized visible light feature map. Finally, the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map are concatenated and then a shared convolution operation is performed to obtain a fused feature map.
[0035] Optionally, in the landing point calculation module, the bounding boxes containing explosion events and corresponding category information in consecutive frames are acquired and stored, and the change rate and stability of the recorded bounding boxes and category information are analyzed. The change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames are calculated. At the same time, by statistically analyzing the consecutive frame data, abnormal results caused by instantaneous noise or misdetection are eliminated. According to the bounding boxes and category information in the candidate frames, the moment when a certain frame first meets the conditions of a stable bounding box and a clear category is selected as the initial explosion moment, and the bounding box in this frame is extracted as the basis for landing point calculation. Using the determined bounding box, the center point is calculated as the preliminary landing point image coordinates. With the help of additional information detected in multiple subsequent frames, the preliminary calculated landing point image coordinates are supplemented and corrected, and finally the corrected landing point image coordinates are obtained, and the corrected landing point image coordinates are output as the position basis corresponding to the initial explosion moment.
[0036] Based on the same inventive concept, the present invention discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a projectile landing point detection method based on dual - light image fusion as described above.
[0037] Based on the same inventive concept, the present invention discloses a computer - readable storage medium, on which a computer program is stored. The computer program is characterized in that it is executed by a processor to implement a projectile landing point detection method based on dual - light image fusion as described above.
[0038] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention has a simple layout and utilizes the complementary characteristics of two image modalities. The infrared thermal imaging modality is more sensitive to smoke and heat radiation, reducing false alarms caused by complex background interference. The visible light modality contains more detailed information, and the combination of the two improves the performance of landing point detection. In the present invention, based on the fusion processing of the registered visible - light image and thermal infrared image, the information from different sensors can be effectively integrated to optimize the target detection result. The pre - detection fusion method in the present invention pre - fuses the two - way data to generate a more representative fusion image for subsequent detection, providing a reliable and accurate input image for subsequent target detection through high - quality image - level fusion. Description of the Drawings
[0039] Figure 1 It is a flow chart of the present invention;
[0040] Figure 2 It is a reference schematic diagram of the actual layout of the target area in the present invention;
[0041] Figure 3 It is a reference schematic diagram of the projected target surface in the present invention.
[0042] Figure 4 It is a schematic diagram of the framework of the system in the present invention;
[0043] Figure 5 It is a schematic diagram of the framework of the pre-detection fusion model in the present invention;
[0044] Figure 6 It is a schematic diagram of the framework of the feature alignment module in the present invention;
[0045] Figure 7 It is a schematic diagram of the framework of the feature interaction module in the present invention. Detailed implementation manners
[0046] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] It should be understood that the present invention can be implemented in different forms and should not be construed as being limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. In the drawings, for clarity, the dimensions and relative dimensions of the components may be exaggerated. The same reference numerals denote the same components throughout.
[0048] Embodiment 1
[0049] As Figure 1 shown, the present invention discloses a projectile landing point detection method based on dual-light image fusion, including the following steps:
[0050] S1. Collect visible light imaging data and thermal infrared imaging data of the entire target area.
[0051] S2. Preprocess the visible light imaging data and the thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas; specifically: configure the acquisition time sequence of the visible light imaging data and the thermal infrared imaging data so that each frame of visible light image and thermal infrared image is synchronously aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas; the present invention uses a dual-light camera to collect data, and the visible light lens and the thermal infrared lens have different resolutions and field of view angles respectively. Spatio-temporal registration can achieve alignment in terms of resolution and field of view angle; however, due to the errors of imaging distortion and pseudo-coaxial error itself, there are still slight offsets in local areas of the images.
[0052] S3. Project the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces;
[0053] Specifically, it includes the following steps:
[0054] Pre - establish and store the mapping relationship of the position information between the ground target area on the visible - light image and the thermal - infrared image after spatio - temporal registration and the reference projection target surface. The position information is coordinate information. Select a number of reference points in the target area in advance. The number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area;
[0055] The reference points can be selected around the target area. The reference points present obvious features on both the visible - light image and the thermal - infrared image. The obvious features include but are not limited to morphological features, texture features, edge features or color features; alternatively, a reference object for calibration can be prepared in advance, with recognizable marking features on its surface, and corresponding feature points can be selected;
[0056] Record the relative position information of a number of reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of a number of reference points in the target - area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of a number of reference points between the target - area image and the reference projection target surface.
[0057] As Figure 2 and Figure 3 shown, the height of the dual - light camera is set to H. There is no specific height requirement for the height H. Considering the actual camera focal length, field - of - view angle and the distance to the target area, it is better that the field of view of the dual - light camera covers the target area. Theoretically, setting the installation position high and the camera angle downward at a certain inclination angle can obtain a better view of the target area; alternatively, it can be considered to deploy in combination with a tethered UAV, which can hover in the air for a long time. Set reference point 51, reference point 52, reference point 53 and reference point 54 around the ground ring target on the ground plane, and use the dual - light camera to collect images of the target area to form a corresponding reference projection target surface.
[0058] S4. Perform pre - detection fusion on the registered visible - light image and thermal - infrared image to generate a high - quality fusion image with both thermal saliency and visible - light details, and then send the high - quality fusion image into a typical target - detection model to judge frame by frame whether an explosion phenomenon is detected in the image, and obtain the bounding box of the explosion target corresponding in the image; based on the fusion processing of the registered visible - light image and thermal - infrared image, it can effectively integrate information from different sensors and optimize the target - detection result; the pre - detection fusion method is to pre - fuse the two - way data to generate a more representative fusion image and then perform detection, providing a reliable and accurate input image for subsequent target detection through high - quality image - level fusion.
[0059] As Figure 5As shown in the figure, pre-detection fusion is performed using a pre-detection fusion model. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. The dual-branch feature extraction module uses a dual-branch structure to process thermal infrared images and visible light images respectively. The thermal infrared image passes through the first 2 convolutional layers to extract shallow thermal infrared features, and then passes through the subsequent 3 convolutional layers to extract deep thermal infrared features. The visible light image passes through the first 1 convolutional layer and the first 1 feature alignment module to extract shallow visible light features, and then passes through the subsequent 1 feature alignment module and 2 convolutional layers to extract deep visible light features. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the feature alignment modules at the corresponding levels of the visible light branch to correct the spatial offset of the visible light image. Each level of deep thermal infrared features and the corresponding deep visible light features are input into the feature interaction module together, and three groups of fused features are output and input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
[0060] The dual-branch feature extraction module uses a dual-branch structure to process thermal infrared images and visible light images respectively. Each branch contains five convolutional layers. The first two convolutional layers are responsible for extracting shallow features of the input image, and combined with a feature alignment module based on offset learning, dynamically correct the spatial offset of the visible light modality feature map to ensure effective alignment with the thermal infrared modality feature map in terms of spatial distribution. The third, fourth, and fifth convolutional layers extract deep features. At the same time, a feature interaction module based on orthogonal attention is introduced to fully exploit the complementary information of the two modalities. After completing feature extraction and interaction, the multi-modal features output by the third, fourth, and fifth convolutional layers are fused to generate three groups of fused features, which are input into the image reconstruction module together. The image reconstruction module consists of three 3×3 convolutional layers, which are responsible for multi-scale reconstruction of the fused features, and finally generate a high-quality fused image with both thermal saliency and visible light details. The high-quality fused image is then used as input and fed into a typical object detection model for detection to determine whether an explosion phenomenon is detected in the image, and the bounding box corresponding to the explosion target in the image is obtained. The typical object detection model can adopt the YOLO series model.
[0061] As Figure 6 shown, the feature alignment module (OLDAM, Offset Learning-based Deformable Alignment Module) learns the dynamic offset of the visible light modality feature map to match it with the thermal infrared feature map in space; hierarchical progressive alignment can more fully combat modality differences and local deformations. The feature alignment module can adaptively process local spatial deformations, effectively reduce the artifact problem caused by spatial errors, and also provide a consistent and reliable spatial representation for subsequent feature fusion. Let F inf ∈R b×c×h×w be the thermal infrared feature map, and the visible light F vis ∈Rb×c×h×w are visible light feature maps, each having c channels and a spatial size of h×w. The aligned visible light modality feature maps are spatially aligned with the thermal infrared modality feature maps.
[0062] As Figure 7 shown, in the feature alignment module, the thermal infrared feature map F inf and the visible light feature map F vis are first concatenated in the channel dimension to obtain a joint input of the feature maps; then, the offsets of the deformable convolution are dynamically generated through a convolution operation to capture the spatial differences of the modalities. The size of the obtained offset tensor is The number of output channels is 2k 2 ; the obtained offsets are used as the sampling point offsets of the deformable convolution and applied to the visible light feature map F vis to obtain the corrected visible light modality feature map
[0063] In the feature interaction module, for the thermal infrared feature map F inf and the visible light feature map F vis , they are first concatenated in the channel dimension to obtain a joint input of the feature maps. Let the joint input feature map F joint ∈R b×2c×h×w , and for ease of representation, without considering the batch dimension, it is denoted as having a shape of (2c, h, w), where 2c is the number of channels; a set of random matrices with the same shape as the joint input feature map is introduced and Schmidt orthogonalization is performed to obtain a set of orthogonal filters, denoted as the orthogonal filter G ortho , also having a shape of (2C, H, W). The filter for each channel is a two-dimensional matrix (H, W). The orthogonalization operation ensures that the channels of the filters are orthogonal; the orthogonal filter and the joint input feature map are subjected to a pointwise multiplication operation to obtain a one-dimensional orthogonalized weight vector, which is used as the orthogonalized feature for each channel. One-dimensional convolution is performed on the orthogonalized weight vector, and the convolution result is sliced and normalized by sigmoid operation to obtain the attention weights for different modalities, denoted as the thermal infrared weight and the visible light weight respectively; the original thermal infrared feature map is multiplied by the thermal infrared weight to obtain the weighted thermal infrared feature, and the weighted thermal infrared feature and the original thermal infrared feature map are added together through a residual connection to obtain the orthogonalized thermal infrared feature map The original visible light feature map is multiplied by the visible light weight to obtain the weighted visible light feature, and the weighted visible light feature and the original visible light feature map are added together through a residual connection to obtain the orthogonalized visible light feature map The residual can help the network retain the original features while strengthening the gain brought by the orthogonal filter; finally, the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map Perform a shared convolution operation again after splicing to obtain the fused feature map F fuse 。
[0064] For the training of the pre-detection fusion model of the present invention, a total loss L composed of five sub-loss terms is designed total ,The total loss L total includes a pixel-based reconstruction loss L MSE 、a gradient-based edge loss L edge 、a structure similarity-based structure loss L SSIM 、a color loss L color and a total variation loss L TV ,
[0065] L total =L MSE +αL edge +βL SSIM +γL color +δL TV
[0066] α, β, γ and δ are weight coefficients corresponding to each loss
[0067] S5. Determine the landing point image coordinates through the explosion bounding box, and according to the coordinate mapping relationship between the landing point image coordinates and the predefined reference points on the target area image and the reference projection target surface, map the calculated landing point image coordinates to the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface
[0068] Adopt a landing point calculation algorithm to correct the position of the landing point based on the bounding box on the image corresponding to the explosion phenomenon provided by the fusion detection algorithm, and obtain the image coordinates of the landing point from the explosion area; according to the pre-stored coordinate mapping relationship between the reference points on the target area image and the reference projection target surface, map the image coordinates of the landing point to the reference projection target surface, and obtain the position coordinate information of the landing point in the target area, and then the specific coordinates of the landing point relative to the target area plane can be known
[0069] The landing point detection algorithm includes the following steps: Obtain and store the bounding boxes and corresponding class information of the candidate regions containing explosion phenomena in consecutive frames for subsequent judgment and statistics of the landing points; Analyze the rate of change and stability of the recorded bounding boxes and class information, calculate the change trends of the positions, sizes, and class confidence levels of the bounding boxes between frames, and at the same time, eliminate abnormal frames caused by instantaneous noise or false detections by statistically analyzing the consecutive frame data to ensure that only the detection results that truly reflect the initial moment of the explosion are retained; According to the bounding boxes and class information in the candidate frames, screen for the moment when a certain frame first meets the conditions of a stable bounding box and a clear class as the initial moment of the explosion, and extract the bounding box in this frame as the basis for landing point calculation; Use the determined bounding box to calculate the center point as the preliminary landing point image coordinates; With the help of additional information from multi-frame detections in subsequent frames, supplement and correct the coordinates obtained from the preliminary calculation, such as analyzing the change trend of the landing point coordinates in consecutive frames through linear regression, or performing an average operation on the coordinates of each frame, or learning and fitting through a machine learning model; Finally, obtain the corrected accurate image coordinates, and output the finally calculated landing point image coordinates as the position basis corresponding to the initial moment of the explosion.
[0070] Embodiment 2
[0071] As Figure 2 shown, the present invention discloses a projectile landing point detection system based on dual-light image fusion, including: A data acquisition module for acquiring visible light imaging data and thermal infrared imaging data of the entire target area. In the present invention, the visible light imaging data captured by the visible light acquisition module in the dual-light acquisition device and the thermal infrared imaging data captured by the thermal imaging acquisition module are connected to the detection and calculation device through the communication module of the auxiliary support device for communication. The dual-light acquisition device can be a dual-light camera, where dual-light refers to visible light and thermal infrared; it can also be two different cameras, one visible light camera and one thermal infrared imaging camera. All cameras mentioned in this embodiment include optical imaging devices such as cameras and video cameras.
[0072] A preprocessing module for preprocessing the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images that are spatio-temporally registered for the corresponding target area; specifically: Configure the acquisition timing of the visible light imaging data and thermal infrared imaging data so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images that are spatio-temporally registered for the corresponding target area; In the present invention, a dual-light camera is used to acquire data. The visible light lens and the thermal infrared lens have different resolutions and field of view angles for imaging. Spatio-temporal registration can achieve alignment in terms of resolution and field of view angle; However, due to the errors of imaging distortion and pseudo-coaxial error itself, there are still slight offsets in local areas of the images.
[0073] The target area mapping module is used to project the ground target areas on the visible light image and the thermal infrared image after spatio-temporal registration to the corresponding reference projection target surface.
[0074] In the target area mapping module, a mapping relationship of the position information between the ground target areas on the visible light image and the thermal infrared image after spatio-temporal registration and the reference projection target surface is established and stored in advance. The position information is coordinate information; several reference points are selected in advance in the target area, the number of reference points is greater than 3, and the connection area formed by the reference points should cover all areas of the target area; the reference points can be selected around the target area, and the reference points have obvious features on both the visible light and infrared thermal imaging images. The obvious features include but are not limited to morphological features, texture features, edge features or color features; or a reference object for calibration is made in advance, and it has marked features for easy identification, and corresponding feature points are selected.
[0075] Record the relative position information of several reference points, and establish a reference projection target surface according to this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the target position information and the image position information, obtain the coordinate mapping relationship of several reference points on the target area image and the reference projection target surface.
[0076] The dual-band fusion detection module is used to perform pre-detection fusion on the registered visible light image and thermal infrared image to generate a high-quality fusion image with both thermal saliency and visible light details, and then send the high-quality fusion image into a typical target detection model to judge frame by frame whether an explosion phenomenon is detected in the image, and obtain the bounding box corresponding to the explosion target in the image; based on the fusion processing of the registered visible light image and thermal infrared image, the information from different sensors can be effectively integrated to optimize the target detection result; the pre-detection fusion method is to pre-fuse the two-way data to generate a more representative fusion image for detection, and provide a reliable and accurate input image for subsequent target detection through high-quality image-level fusion.
[0077] Such as Figure 7As shown in the figure, pre-detection fusion is performed using a pre-detection fusion model. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. The dual-branch feature extraction module uses a dual-branch structure to process thermal infrared images and visible light images respectively. The thermal infrared image passes through the first 2 convolutional layers to extract shallow thermal infrared features, and then passes through the subsequent 3 convolutional layers to extract deep thermal infrared features. The visible light image passes through the first 1 convolutional layer and the first 1 feature alignment module to extract shallow visible light features, and then passes through the subsequent 1 feature alignment module and 2 convolutional layers to extract deep visible light features. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the feature alignment modules at the corresponding levels of the visible light branch to correct the spatial offset of the visible light image. Each level of deep thermal infrared features and the corresponding deep visible light features are input into the feature interaction module together, and three groups of fused features are output and input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
[0078] The dual-branch feature extraction module uses a dual-branch structure to process thermal infrared images and visible light images respectively. Each branch contains five convolutional layers. The first two convolutional layers are responsible for extracting shallow features of the input image and, combined with a feature alignment module based on offset learning, dynamically correct the spatial offset of the visible light modality feature map to ensure effective alignment with the thermal infrared modality feature map in terms of spatial distribution. The third, fourth, and fifth convolutional layers extract deep features. At the same time, a feature interaction module based on orthogonal attention is introduced to fully mine the complementary information of the two modalities. After completing feature extraction and interaction, the multi-modal features output by the third, fourth, and fifth convolutional layers are fused to generate three groups of fused features, which are input into the image reconstruction module together. The image reconstruction module consists of three 3×3 convolutional layers, which are responsible for multi-scale reconstruction of the fused features, and finally generate a high-quality fused image with both thermal saliency and visible light details. The high-quality fused image is then used as input and fed into a typical object detection model for detection to determine whether an explosion phenomenon is detected in the image, and the bounding box corresponding to the explosion target in the image is obtained. The typical object detection model can adopt the YOLO series model.
[0079] The feature alignment module (OLDAM, Offset Learning-based Deformable Alignment Module) learns the dynamic offset of the visible light modality feature map to achieve spatial matching with the thermal infrared feature map; hierarchical progressive alignment can more fully combat modality differences and local deformations. The feature alignment module can adaptively process local spatial deformations, effectively reduce the artifact problem caused by spatial errors, and also provide a consistent and reliable spatial representation for subsequent feature fusion. Let F inf ∈R b×c×h×w be the thermal infrared feature map, and the visible light F vis ∈R b×c×h×ware visible light feature maps, each having c channels and a spatial size of h×w. The aligned visible light modality feature maps are spatially aligned with the thermal infrared modality feature maps.
[0080] In the feature alignment module, the thermal infrared feature map F inf and the visible light feature map F vis are first concatenated in the channel dimension to obtain a joint input of the feature maps; then the offsets of the deformable convolution are dynamically generated through convolution operations to capture the spatial differences of the modalities. The size of the obtained offset tensor is with the number of output channels being 2k 2 ; the obtained offsets are used as the sampling point offsets of the deformable convolution and applied to the visible light feature map F vis to obtain the corrected visible light modality feature map
[0081] In the feature interaction module, the thermal infrared feature map F inf and the visible light feature map F vis are first concatenated in the channel dimension to obtain a joint input of the feature maps. Let the joint input feature map F joint ∈R b×2c×h×w . For ease of representation, without considering the batch dimension, it is denoted as having a shape of (2c, h, w), where 2c is the number of channels; a set of random matrices with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonal filters, denoted as the orthogonal filter G ortho , which also has a shape of (2C, H, W). The filter for each channel is a two-dimensional matrix (H, W). The orthogonalization operation ensures that the channels of the filters are orthogonal; the orthogonal filter and the joint input feature map are subjected to a pointwise multiplication operation to obtain a one-dimensional orthogonalized weight vector, which is used as the orthogonalized feature for each channel. One-dimensional convolution is performed on the orthogonalized weight vector, and the convolution result is sliced and normalized by sigmoid operation to obtain the attention weights for different modalities, denoted as the thermal infrared weight and the visible light weight respectively; the original thermal infrared feature map is multiplied by the thermal infrared weight to obtain the weighted thermal infrared feature, and the weighted thermal infrared feature and the original thermal infrared feature map are added together through a residual connection to obtain the orthogonalized thermal infrared feature map The original visible light feature map is multiplied by the visible light weight to obtain the weighted visible light feature, and the weighted visible light feature and the original visible light feature map are added together through a residual connection to obtain the orthogonalized visible light feature map The residual can help the network retain the original features while strengthening the gain brought by the orthogonal filter; finally, the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map are concatenated and then a shared convolution operation is performed again to obtain the fused feature map Ffuse 。
[0082] For the training of the pre-detection fusion model of the present invention, a total loss L composed of five sub-loss terms is designed total , and the total loss L total includes a pixel-based reconstruction loss L MSE , a gradient-based edge loss L edge , a structure similarity-based structure loss L SSIM , a color loss L color and a total variation loss L TV ,
[0083] L total = L MSE + αL edge + βL SSIM + γL color + δL TV
[0084] α, β, γ and δ are weight coefficients corresponding to each loss.
[0085] The landing point calculation module is used to calculate and correct the position of the landing point according to the explosion bounding box by using the landing point calculation algorithm, and determine the landing point image coordinates from the explosion area; according to the coordinate mapping relationship of the reference point on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface to obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface. In the landing point calculation module, the bounding boxes and corresponding category information of the candidate areas containing explosion phenomena in consecutive frames are obtained and stored, the change rate and stability of the recorded bounding boxes and category information are analyzed, the change trends of the positions, sizes and category confidence levels of the bounding boxes between frames are calculated, and at the same time, by statistically analyzing the consecutive frame data, the abnormal frames caused by instantaneous noise or misdetection are eliminated; according to the bounding boxes and category information in the candidate frames, select the moment when a certain frame first satisfies the conditions of a stable bounding box and a clear category as the explosion initial moment, and extract the bounding box in this frame as the basis for landing point calculation; use the determined bounding box to calculate the center point as the preliminary landing point image coordinates; with the help of the additional information detected in multiple frames in subsequent frames, supplement and correct the preliminary calculated landing point image coordinates, finally obtain the corrected landing point image coordinates, and output the corrected landing point image coordinates as the position basis corresponding to the explosion initial moment.
[0086] Embodiment 3
[0087] Another embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the above-mentioned method for detecting the landing point of a projectile based on dual-optical image fusion.
[0088] An electronic device may include: a processor, a memory, a bus, and a communication interface, where the processor, the communication interface, and the memory are connected through the bus; a computer program that can run on the processor is stored in the memory, and when the processor runs the computer program, it executes a method for detecting the projectile landing point based on dual - optical image fusion provided in any of the foregoing embodiments of the present invention.
[0089] Among them, the memory may include a high - speed random - access memory (RAM: Random Access Memory), and may also include a non - volatile memory, such as at least one disk memory. The communication connection between the device network element and at least one other network element is realized through at least one communication interface (which can be wired or wireless), and the Internet, wide - area network, local area network, metropolitan area network, etc. can be used.
[0090] The bus can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory is used to store the program, and after receiving the execution instruction, the processor executes the program. A method for detecting the projectile landing point based on dual - optical image fusion disclosed in any of the foregoing embodiments of the present invention can be applied to the processor or implemented by the processor.
[0091] The processor may be an integrated circuit chip with signal - processing capabilities. In the implementation process, each step of the above - mentioned method can be completed by the integrated logic circuit in the hardware of the processor or the instruction in the form of software. The above - mentioned processor can be a general - purpose processor, which may include a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general - purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art, such as a random - access memory, flash memory, read - only memory, programmable read - only memory, or electrically erasable programmable memory, register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above - mentioned method.
[0092] The electronic device provided by the embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.
[0093] Embodiment 4
[0094] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the method of any of the above embodiments. The computer-readable storage medium is an optical disc, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method provided by any of the foregoing embodiments.
[0095] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.
[0096] The computer-readable storage medium provided by the above embodiment of the present application and the method provided by the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by the application program stored on it.
Claims
1. A method for detecting the impact point of a projectile based on dual - light image fusion, characterized in that, It includes the following steps: S1. Collect visible light imaging data and thermal infrared imaging data of all target areas; S2. Preprocess the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas; S3. Project the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces; S4. Perform pre-detection fusion on the registered visible light images and thermal infrared images to generate high-quality fusion images with both thermal saliency and visible light details, and then send the high-quality fusion images into a typical target detection model to judge frame by frame whether an explosion phenomenon is detected in the images, and obtain the bounding boxes corresponding to the explosion targets in the images; S5. According to the explosion bounding boxes, use the landing point calculation algorithm to calculate and correct the position of the landing point, and determine the landing point image coordinates in the explosion area; according to the coordinate mapping relationship between the reference points on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface to obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface.
2. The method for detecting the projectile landing point based on dual - light image fusion according to claim 1, wherein, The specific steps of step S2 include the following: Configure the acquisition time sequence for the visible light imaging data and thermal infrared imaging data so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas.
3. The method for detecting the projectile impact point based on dual - light image fusion according to claim 1, wherein: The specific steps of step S3 include the following: Pre-establish and store the mapping relationship of the position information between the ground target areas on the spatio-temporally registered visible light images and thermal infrared images and the reference projection target surface, and the position information is coordinate information; Select several reference points in the target area in advance, the number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points show obvious features in both visible light and thermal infrared images, and the obvious features include but are not limited to morphological features, texture features, edge features or color features; or pre-fabricate a reference object for calibration, whose surface has recognizable marking features, and select the corresponding feature points; Record the relative position information of several reference points, and establish a reference projection target surface according to this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points on the target area image and the reference projection target surface.
4. The method for detecting the projectile impact point based on dual - light image fusion according to claim 1, wherein, In step S4, pre-detection fusion is performed using a pre-detection fusion model. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. The dual-branch feature extraction module uses a dual-branch structure to process the thermal infrared image and the visible light image respectively. The thermal infrared image extracts shallow thermal infrared features through the first two convolutional layers and then extracts deep thermal infrared features through the next three convolutional layers. The visible light image extracts shallow visible light features through the first convolutional layer and the first feature alignment module, and then extracts deep visible light features through the next feature alignment module and the next two convolutional layers. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the corresponding feature alignment modules to correct the spatial offset of the visible light image. Each level of deep thermal infrared features and the corresponding deep visible light features are input into the feature interaction module together, and the output three groups of fused features are input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
5. The method for detecting the projectile impact point based on dual - light image fusion according to claim 4, wherein, In the feature alignment module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input of the feature maps. Then, the offset of the deformable convolution is dynamically generated through convolution operations. The obtained offset is used as the sampling point offset of the deformable convolution and applied to the visible light feature map to obtain a corrected visible light modality feature map.
6. The method for detecting the projectile landing point based on dual - light image fusion according to claim 4, wherein, In the feature interaction module, for the thermal infrared feature map and the visible light feature map, they are first concatenated in the channel dimension to obtain a joint input of the feature maps. A set of random matrices with the same shape as the joint input feature map is introduced and Schmidt orthogonalization is performed to obtain a set of orthogonal filters. The orthogonal filters and the joint input feature map are subjected to a point-by-point multiplication operation to obtain a one-dimensional orthogonal weight vector. One-dimensional convolution is performed on the orthogonal weight vector, and the convolution result is sliced and normalized to obtain attention weights of different modalities, which are respectively denoted as the thermal infrared weight and the visible light weight. The original thermal infrared feature map is multiplied by the thermal infrared weight to obtain a weighted thermal infrared feature, and the weighted thermal infrared feature and the original thermal infrared feature map are added together through a residual connection to obtain an orthogonalized thermal infrared feature map. The original visible light feature map is multiplied by the visible light weight to obtain a weighted visible light feature, and the weighted visible light feature and the original visible light feature map are added together through a residual connection to obtain an orthogonalized visible light feature map. Finally, the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map are concatenated and then subjected to a shared convolution operation to obtain a fused feature map.
7. The bullet impact point detection method based on dual - light image fusion according to claim 1, characterized in that, The specific steps of step S5 are as follows: Obtain and store the bounding boxes containing explosion events and corresponding category information in consecutive frames, analyze the change rate and stability of the recorded bounding boxes and category information, calculate the change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames, and at the same time, by statistically analyzing the consecutive frame data, eliminate abnormal results caused by instantaneous noise or misdetection; According to the bounding boxes and category information in the candidate frames, screen for the moment when a certain frame first meets the conditions of stable bounding boxes and clear categories as the initial explosion moment, and extract the bounding box in this frame as the basis for impact point calculation; Use the determined bounding box to calculate the center point as the preliminary impact point image coordinates; With the help of additional information detected in multiple subsequent frames, supplement and correct the preliminary calculated impact point image coordinates, finally obtain the corrected impact point image coordinates, and output the corrected impact point image coordinates as the position basis corresponding to the initial explosion moment.
8. A projectile landing point detection system based on dual - light image fusion, characterized in that, Including: A data acquisition module for acquiring visible light imaging data and thermal infrared data of the entire target area; A preprocessing module for preprocessing the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target area; A target area mapping module for projecting the ground target area on the obtained spatio-temporally registered visible light image and thermal infrared image onto the corresponding reference projection target surface; A dual-light fusion detection module for performing pre-detection fusion on the registered visible light image and thermal infrared image to generate a high-quality fusion image with both thermal saliency and visible light details, and then sending the high-quality fusion image into a typical target detection model to judge frame by frame whether an explosion phenomenon is detected in the image, and obtaining the bounding box corresponding to the explosion target in the image; An impact point calculation module for calculating and correcting the position of the impact point using an impact point calculation algorithm according to the explosion bounding box, and determining the impact point image coordinates in the explosion area; According to the coordinate mapping relationship of the reference point on the target area image and the reference projection target surface, map the impact point image coordinates to the reference projection target surface to obtain the position coordinate information of the impact point within the reference projection target surface, and record and output the relative position information of the impact point on the reference projection target surface.
9. The bullet impact point detection system based on dual - light image fusion according to claim 8, characterized in that: In the preprocessing module, the acquisition time sequence of the visible light imaging data and thermal infrared imaging data is configured so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then each frame of visible light image and thermal infrared image is registered and aligned in space to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target area.
10. The bullet impact point detection system based on dual - light image fusion according to claim 8, characterized in that: In the target area mapping module, a mapping relationship of the position information (the position information is coordinate information) between the ground target area on the spatio-temporally registered visible light image and thermal infrared image and the reference projection target surface is established and stored in advance. Select a number of reference points in the target area in advance. The number of reference points is greater than 3, and the connected area formed by the reference points should cover the entire target area. The reference points can be selected around the target area, and the reference points have obvious features in both visible light and thermal infrared images. The obvious features include, but are not limited to, morphological features, texture features, edge features, or color features. Or prepare a reference object for calibration in advance, with marked features on its surface for easy identification, and select the corresponding feature points. Record the relative position information of a number of reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as the target position information. Determine the image coordinate positions of the reference points in the target area image, denoted as the image position information. According to the image position information and the target position information, obtain the coordinate mapping relationship of the reference points on the target area image and the reference projection target surface.
11. The bullet impact point detection system based on dual - light image fusion according to claim 8, wherein: In the dual-band fusion detection module, a pre-detection fusion model is used for pre-detection fusion. The pre-detection fusion model includes a dual-branch feature extraction module, a feature interaction module, and an image reconstruction module. Among them, the dual-branch feature extraction module uses a dual-branch structure to process the thermal infrared image and the visible light image respectively. The thermal infrared image extracts the shallow thermal infrared features through the first 2 convolutional layers, and then extracts the deep thermal infrared features through the subsequent 3 convolutional layers. The visible light image extracts the shallow visible light image through the first 1 convolutional layer and the first 1 feature alignment module, and then extracts the deep visible light features through the subsequent 1 feature alignment module and 2 convolutional layers. The second-level shallow thermal infrared features and the third-level deep thermal infrared features are respectively input into the corresponding feature alignment modules to correct the spatial offset of the visible light image. Each level of deep thermal infrared features and the corresponding deep visible light features are input into the feature interaction module together, and three groups of fused features are output and input into the image reconstruction module together to generate a high-quality fused image with both thermal saliency and visible light details.
12. The bullet impact point detection system based on dual - light image fusion according to claim 11, wherein: In the feature alignment module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain the joint input of the feature maps. Then, the offset of the deformable convolution is dynamically generated through convolution operations. The obtained offset is used as the sampling point offset of the deformable convolution and applied to the visible light feature map to obtain the corrected visible light modal feature map.
13. A bullet impact point detection system based on dual - light detection pre - fusion according to claim 11, characterized in that: In the feature interaction module, for the thermal infrared feature map and the visible light feature map, they are first concatenated in the channel dimension to obtain the joint input of the feature maps. Introduce a set of random matrices with the same shape as the joint input feature map, and perform Schmidt orthogonalization to obtain a set of orthogonalized filters. The orthogonalized filters perform point-by-point multiplication operations with the joint input feature map to obtain a one-dimensional orthogonalized weight vector. Perform one-dimensional convolution on the orthogonalized weight vectors, slice and normalize the convolution results to obtain the attention weights of different modalities, denoted as the thermal infrared weight and the visible light weight respectively; multiply the original thermal infrared feature map by the thermal infrared weight to obtain the weighted thermal infrared feature, and add the weighted thermal infrared feature to the original thermal infrared feature map through residual connection to obtain the orthogonalized thermal infrared feature map; multiply the original visible light feature map by the visible light weight to obtain the weighted visible light feature, and add the weighted visible light feature to the original visible light feature map through residual connection to obtain the orthogonalized visible light feature map; finally, splice the orthogonalized thermal infrared feature map and the orthogonalized visible light feature map and perform a shared convolution operation again to obtain the fused feature map.
14. The bullet impact point detection system based on dual - light image fusion according to claim 8, characterized in that: The landing point calculation module acquires and stores the bounding boxes containing explosion events and the corresponding category information in consecutive frames, analyzes the change rate and stability of the recorded bounding boxes and category information, calculates the change trends of the positions, sizes and category confidence levels of the bounding boxes between frames, and at the same time eliminates abnormal results caused by instantaneous noise or misdetection by statistically analyzing the consecutive frame data; according to the bounding boxes and category information in the candidate frames, select the moment when a frame first satisfies the conditions of stable bounding box and clear category as the initial explosion moment, and extract the bounding box in this frame as the basis for landing point calculation; use the determined bounding box to calculate the center point as the preliminary landing point image coordinates; with the help of the additional information detected in multiple subsequent frames, supplement and correct the preliminary calculated landing point image coordinates, finally obtain the corrected landing point image coordinates, and output the corrected landing point image coordinates as the position basis corresponding to the initial explosion moment.
15. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a projectile landing point detection method based on dual-light image fusion as described in any one of claims 1-7.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to implement a projectile landing point detection method based on dual-light image fusion as described in any one of claims 1-7.
Citation Information
Patent Citations
Detection method and device for shell explosion information and storage medium
CN114549498A
Geographic positioning method for shell drop point of double-point fixed camera
CN117392233A