Projectile drop point detection method, system and equipment based on dual-light end-to-end fusion and medium

Through the dual-light end-to-end fusion method, combined with visible light and thermal infrared imaging data, the multi-modal fusion detection model is used to perform spatial and temporal registration and explosion phenomenon detection in the target area, solving the problem of low landing accuracy of projectiles in the existing technology, and achieving high-precision and robust landing detection effect.

CN120339388APending Publication Date: 2025-07-18NANJING RES INST ON SIMULATION TECHN

Patent Information

Application Number
CN202510407703.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art has low landing point detection accuracy in outdoor long-distance large-diameter projectile training scenarios. The existing methods such as radar, sound waves, light curtains, etc. have high cost, easy interference or complex layout problems. In addition, a single visible light or thermal infrared image detection method has high false alarm rate or miss detection, making it difficult to achieve accurate positioning.

Method used

The method based on dual-light end-to-end fusion is adopted, combined with visible light and thermal infrared imaging data, and the target area spatiotemporal registration and explosion phenomenon detection model are used to detect target areas, and multi-scale, multi-modal feature interaction and gated fusion technology are used to improve detection accuracy.

Benefits of technology

It realizes high-precision projectile landing point detection, reduces complex background interference, improves detection performance, meets real-time and accuracy requirements, and enhances robustness in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339388A_ABST
    Figure CN120339388A_ABST
Patent Text Reader

Abstract

The invention discloses a projectile drop point detection method, system and equipment based on dual-light end-to-end fusion, and a medium. The method comprises the following steps: collecting visible light imaging data and thermal infrared imaging data of all target regions; preprocessing the data to obtain visible light and thermal infrared images of corresponding target region time-space registration; projecting a ground target area on the obtained image after space-time registration to a corresponding reference projection target surface; inputting the registered image into a multi-modal fusion detection model, and judging whether an explosion phenomenon is detected in the image frame by frame to obtain a bounding box corresponding to an explosion target in the image; according to the explosion bounding box, resolving and correcting the position of the drop point by adopting a drop point resolving algorithm, and determining the image coordinate of the drop point in the explosion area; and mapping the drop point to a reference projection target surface, obtaining position coordinate information of the drop point in the reference projection target surface, and recording and outputting relative position information of the drop point on the reference projection target surface. The detection precision is improved, and the target scoring task is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of impact point detection, and particularly to a method, system, device and medium for projectile impact point detection based on dual-light end-to-end fusion. Background Art

[0002] In the field of shooting detection, especially in the outdoor long-distance large-caliber projectile training scenario, for the positioning of the impact point, there are currently methods such as manual target inspection and sighting target inspection, which have poor accuracy and low efficiency. Other target reporting methods that do not require human participation, such as radar, sound waves, and light curtains, each have their limitations. Radar technology is relatively complex, vulnerable to electromagnetic interference, and has a high cost; the sound wave detection method is greatly affected by factors such as terrain and noise, and the positioning accuracy is largely affected by the layout scheme; the light curtain has problems such as complex layout and many interferences in a large field of view. The image-based method is based on the projectile explosion phenomenon and conforms to human intuition. The traditional binocular positioning method has high requirements for the site and complex layout; the image detection method of single visible light or thermal infrared is affected by conditions such as the size and shape of the projectile explosion, with a high false alarm rate or prone to missed detection, making it difficult to accurately detect or locate the projectile impact point. For example, the existing patent CN202210189027.4 discloses a method based on a visible light camera using traditional image processing methods, but this method is greatly affected by environmental factors and background complexity in actual applications; the existing technology CN202311144943.7 discloses a method based on binocular vision, which automatically detects the projectile explosion area and obtains the coordinates of the projectile impact point in the image, and calculates the coordinates and distance of the projectile impact point relative to the target; however, this existing technology has a complex layout, and the selection of key parameters such as camera calibration, baseline length measurement, and angle measurement has a great impact on the results, and is greatly affected by environmental factors.

[0003] Therefore, it is urgent to solve the above problems. Summary of the Invention

[0004] Object of the Invention: The first object of the present invention is to provide a method for projectile impact point detection based on dual-light end-to-end fusion, which improves the detection accuracy and better completes the target reporting task.

[0005] The second object of the present invention is to provide a system for projectile impact point detection based on dual-light end-to-end fusion.

[0006] The third object of the present invention is to provide an electronic device.

[0007] The fourth object of the present invention is to provide a computer-readable storage medium.

[0008] Technical Solution: To achieve the above objects, the present invention discloses a method for projectile impact point detection based on dual-light end-to-end fusion, including the following steps:

[0009] S1. Collect visible light imaging data and thermal infrared imaging data of all target areas;

[0010] S2. Preprocess the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas;

[0011] S3. Project the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces;

[0012] S4. Input the registered visible light images and thermal infrared images into a multi-modal fusion detection model, and frame by frame determine whether an explosion phenomenon is detected in the images to obtain the bounding boxes of the explosion targets corresponding in the images;

[0013] S5. According to the explosion bounding boxes, use a landing point calculation algorithm to calculate and correct the position of the landing point, and determine the landing point image coordinates in the explosion area; according to the coordinate mapping relationship of the reference points on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface to obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface.

[0014] Optionally, the step S2 specifically includes the following steps: Configure the acquisition time sequence for the visible light imaging data and thermal infrared imaging data, so that each frame of visible light image and thermal infrared image is synchronously aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas.

[0015] Optionally, the step S3 specifically includes the following steps:

[0016] Pre-establish and store the mapping relationship of the position information between the ground target areas on the spatio-temporally registered visible light images and thermal infrared images and the reference projection target surfaces, and the position information is coordinate information;

[0017] Select a number of reference points in advance in the target area, the number of reference points is greater than 3, and the connected area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points show obvious features in both visible light and thermal infrared images, and the obvious features include but are not limited to morphological features, texture features, edge features or color features; or pre-fabricate a reference object for calibration, whose surface has recognizable marking features, and select the corresponding feature points;

[0018] Record the relative position information of several reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; and obtain the coordinate mapping relationship of several reference points between the target area image and the reference projection target surface according to the image position information and the target position information.

[0019] Optionally, the multi-modal fusion detection model includes a dual-branch feature extraction module, a gated fusion module, and a detection head. The dual-branch feature extraction module uses a dual-branch structure to process visible light images and thermal infrared images respectively. The thermal infrared image and the visible light image respectively pass through a convolutional layer to extract first-level shallow thermal infrared features and first-level shallow visible light features. After being corrected by a recalibration module, the first-level shallow thermal infrared features and the first-level shallow visible light features are respectively input into a convolutional layer to extract second-level shallow thermal infrared features and second-level shallow visible light features. After being corrected by a recalibration module, the second-level shallow thermal infrared features and the second-level shallow visible light features are respectively input into a convolutional layer to extract third-level deep thermal infrared features and third-level deep visible light features. After cross-modal interaction by a cross-modal interaction module, the third-level deep thermal infrared features and the third-level deep visible light features are respectively input into a convolutional layer to extract fourth-level deep thermal infrared features and fourth-level deep visible light features. After cross-modal interaction by a cross-modal interaction module, the fourth-level deep thermal infrared features and the fourth-level deep visible light features are respectively input into a convolutional layer to extract fifth-level deep thermal infrared features and fifth-level deep visible light features; in the gated fusion module, the third-level, fourth-level, and fifth-level deep thermal infrared features and deep visible light features are respectively input into a same-scale fusion module for multi-scale fusion to obtain third-level, fourth-level, and fifth-level fusion features. The fifth-level fusion feature is obtained through a convolutional operation to get a fifth-level fusion intermediate feature. The fifth-level fusion intermediate feature is upsampled and input into a cross-scale recombination module together with the fourth-level fusion feature to obtain a fourth-level recombination intermediate feature. The fourth-level recombination intermediate feature is upsampled and input into a cross-scale recombination module together with the third-level fusion feature to obtain a third-level recombination feature. The third-level recombination feature is downsampled, the fifth-level fusion intermediate feature is upsampled, and input into a cross-scale recombination module together with the fourth-level recombination intermediate feature to obtain a fourth-level recombination feature. The fourth-level recombination feature is downsampled and input into a cross-scale recombination module together with the fifth-level fusion intermediate feature to obtain a fifth-level recombination feature; the third-level, fourth-level, and fifth-level recombination features are respectively input into the detection head, and then post-processed to output the category information and positioning information corresponding to the explosion target, which are drawn in the original image to obtain an image with an explosion target bounding box.

[0020] Optionally, in the recalibration module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input feature map; the joint input feature map is subjected to shared convolution to achieve feature dimension reduction to obtain a general representation; then the general representation is output as a thermal infrared offset vector and a visible light offset vector after depthwise separable convolution operation. The thermal infrared offset vector is used as the sampling point offset and applied to the thermal infrared feature map to obtain a corrected aligned thermal infrared feature map; the visible light offset vector is used as the sampling point offset and applied to the visible light feature map to obtain a corrected aligned visible light feature map.

[0021] Optionally, in the cross-modal interaction module, the thermal infrared feature map and the visible light feature map are first linearly mapped respectively to generate Q, K, and V vectors corresponding to the features of each modality; then through cross-attention operation, the Q vector of the thermal infrared feature map is subjected to attention operation with the K and V vectors of the visible light feature map to obtain a thermal infrared fusion feature map; the Q vector of the visible light feature map is subjected to attention operation with the K and V vectors of the thermal infrared feature map to obtain a visible light fusion feature map; the visible light fusion feature map and the visible light feature map are element-wise added and normalized to output a visible light fusion normalized feature map. The visible light fusion normalized feature map is processed by a feed-forward neural network and then element-wise added and normalized with the visible light fusion normalized feature map again, and finally an optimized and enhanced interactive visible light feature map is output; the thermal infrared fusion feature map and the thermal infrared feature map are element-wise added and normalized to output a thermal infrared fusion normalized feature map. The thermal infrared fusion normalized feature map is processed by a feed-forward neural network and then element-wise added and normalized with the thermal infrared fusion normalized feature map again, and finally an optimized and enhanced interactive thermal infrared feature map is output.

[0022] Optionally, in the same-scale fusion module, first, two feature maps with the same scale in the input are concatenated and convolutionally downsampled to obtain grouped features. Then, adaptive average pooling operations are performed on the grouped features in the width dimension and height dimension respectively to capture spatial context information from both horizontal and vertical directions. Next, the spatial context information is concatenated and feature fusion is performed through convolution. Then, a Split operation is performed to separate the fused features back into height and width components, and a Sigmoid activation function is applied to generate spatial attention weights. The spatial attention weights are applied to the grouped features to obtain enhanced features. At the same time, local spatial features are extracted from the grouped features through convolution. The enhanced features and local spatial features are respectively subjected to global average pooling operations, Reshape operations, and Softmax to calculate the channel attention weights of the enhanced features and the channel attention weights of the local spatial features. At the same time, after the enhanced features and local spatial features are subjected to Reshape operations, cross-attention operations are performed, that is, the reshaped enhanced features are multiplied by the channel attention weights of the local spatial features, and the reshaped local spatial features are multiplied by the channel attention weights of the enhanced features, and then added to obtain cross-attention. Then, the cross-attention is Reshaped to the original spatial size, and a Sigmoid activation function is applied to obtain the final weights. The final weights are applied to the grouped features and recombined in the channel dimension to obtain the final fused features.

[0023] Optionally, in the cross-scale recombination module, first, two feature maps with different scales in the input are concatenated in the channel dimension to obtain a joint input of the feature maps. After convolution downsampling, a joint input with the same shape as the original feature map is obtained. Then, a random matrix with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonalized filters. The orthogonalized filters are subjected to pointwise multiplication operations with the joint input feature map to obtain a one-dimensional orthogonalized weight vector. A Softmax operation is performed on the orthogonalized weight vector, and a weight matrix is adaptively learned based on the current feature relationship. The gated fusion mechanism assigns weights to different scales, performs weighted combination, and then performs residual connection, and adds them to the original two features respectively to obtain the final fused features.

[0024] Optionally, step S5 specifically includes the following steps: acquiring and storing the bounding boxes containing explosion events and the corresponding category information in consecutive frames, performing change rate and stability analysis on the recorded bounding boxes and category information, calculating the change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames, and at the same time, by statistically analyzing the consecutive frame data, eliminating abnormal results caused by instantaneous noise or misdetection; according to the bounding boxes and category information in the candidate frames, screening the moment when a frame first satisfies the conditions of a stable bounding box and a clear category as the initial explosion moment, and extracting the bounding box in this frame as the basis for impact point calculation; using the determined bounding box, calculating the center point as the preliminary impact point image coordinates; with the help of additional information detected in multiple subsequent frames, supplementing and correcting the preliminary calculated impact point image coordinates, finally obtaining the corrected impact point image coordinates, and outputting the corrected impact point image coordinates as the position basis corresponding to the initial explosion moment.

[0025] Based on the same inventive concept, the present invention discloses a projectile impact point detection system based on dual optical end-to-end fusion, including: a data acquisition module for acquiring visible light imaging data and thermal infrared imaging data of the entire target area;

[0026] a preprocessing module for preprocessing the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas;

[0027] a target area mapping module for projecting the ground target area on the spatio-temporally registered visible light image and thermal infrared image onto the corresponding reference projection target surface;

[0028] an end-to-end fusion detection module for inputting the registered visible light image and thermal infrared image into a multi-modal fusion detection model, and judging frame by frame whether an explosion phenomenon is detected in the image to obtain the bounding box corresponding to the explosion target in the image;

[0029] an impact point calculation module for calculating and correcting the position of the impact point according to the explosion bounding box, and determining the impact point image coordinates in the explosion area; according to the coordinate mapping relationship of the reference point on the target area image and the reference projection target surface, mapping the impact point image coordinates to the reference projection target surface to obtain the position coordinate information of the impact point within the reference projection target surface, and recording and outputting the relative position information of the impact point on the reference projection target surface.

[0030] Optionally, in the preprocessing module, the acquisition time sequence of the visible light imaging data and the thermal infrared imaging data is configured so that each frame of visible light image and thermal infrared image is synchronized in time, and then each frame of visible light image and thermal infrared image is registered and aligned in space to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas.

[0031] Optionally, a mapping relationship between the position information of the ground target area on the visible light image and the thermal infrared image after spatio-temporal registration and the reference projection target surface is established and stored in advance in the target area mapping module, and the position information is coordinate information;

[0032] Select several reference points in the target area in advance. The number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points have obvious features in both the visible light and thermal infrared images. The obvious features include but are not limited to morphological features, texture features, edge features or color features; or a reference object for calibration is made in advance, and it has identifiable marking features on its surface, and corresponding feature points are selected;

[0033] Record the relative position information of several reference points, and establish a reference projection target surface according to this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points on the target area image and the reference projection target surface.

[0034] Optionally, the multi-modal fusion detection model includes a dual-branch feature extraction module, a gated fusion module, and a detection head. The dual-branch feature extraction module uses a dual-branch structure to process visible light images and thermal infrared images respectively. The thermal infrared image and the visible light image pass through convolutional layers to extract the first-level shallow thermal infrared features and the first-level shallow visible light features. After being corrected by the recalibration module, the first-level shallow thermal infrared features and the first-level shallow visible light features are respectively input into convolutional layers to extract the second-level shallow thermal infrared features and the second-level shallow visible light features. After being corrected by the recalibration module, the second-level shallow thermal infrared features and the second-level shallow visible light features are respectively input into convolutional layers to extract the third-level deep thermal infrared features and the third-level deep visible light features. After cross-modal interaction by the cross-modal interaction module, the third-level deep thermal infrared features and the third-level deep visible light features are respectively input into convolutional layers to extract the fourth-level deep thermal infrared features and the fourth-level deep visible light features. After cross-modal interaction by the cross-modal interaction module, the fourth-level deep thermal infrared features and the fourth-level deep visible light features are respectively input into convolutional layers to extract the fifth-level deep thermal infrared features and the fifth-level deep visible light features. In the gated fusion module, the third-level, fourth-level, and fifth-level deep thermal infrared features and deep visible light features are respectively input into the same-scale fusion module for multi-scale fusion to obtain the third-level, fourth-level, and fifth-level fusion features. The fifth-level fusion feature is obtained through a convolutional operation to obtain the fifth-level fusion intermediate feature. The fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level fusion feature to obtain the fourth-level recombination intermediate feature. The fourth-level recombination intermediate feature is upsampled and input into the cross-scale recombination module together with the third-level fusion feature to obtain the third-level recombination feature. The third-level recombination feature is downsampled, and the fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level recombination intermediate feature to obtain the fourth-level recombination feature. The fourth-level recombination feature is downsampled and input into the cross-scale recombination module together with the fifth-level fusion intermediate feature to obtain the fifth-level recombination feature. The third-level, fourth-level, and fifth-level recombination features are respectively input into the detection head, and then post-processed to output the category information and localization information corresponding to the explosion target, which are drawn in the original image to obtain an image with the explosion target bounding box.

[0035] Optionally, in the recalibration module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input feature map. The joint input feature map undergoes shared convolution to achieve feature dimensionality reduction to obtain a general representation. Then, the general representation undergoes depthwise separable convolution operations to output a thermal infrared offset vector and a visible light offset vector. The thermal infrared offset vector is used as the sampling point offset and applied to the thermal infrared feature map to obtain the corrected aligned thermal infrared feature map. The visible light offset vector is used as the sampling point offset and applied to the visible light feature map to obtain the corrected aligned visible light feature map.

[0036] Optionally, in the cross-modal interaction module, the thermal infrared feature map and the visible light feature map are first linearly mapped to generate Q, K, and V vectors corresponding to the features of each modality. Then, through cross-attention operation, the Q vector of the thermal infrared feature map is subjected to attention operation with the K vector and V vector of the visible light feature map to obtain a thermal infrared fusion feature map; the Q vector of the visible light feature map is subjected to attention operation with the K vector and V vector of the thermal infrared feature map to obtain a visible light fusion feature map; the visible light fusion feature map and the visible light feature map are added element-wise and normalized to output a normalized visible light fusion feature map. The normalized visible light fusion feature map is processed by a feed-forward neural network and then added element-wise and normalized again with the normalized visible light fusion feature map to finally output an optimized and enhanced interactive visible light feature map; the thermal infrared fusion feature map and the thermal infrared feature map are added element-wise and normalized to output a normalized thermal infrared fusion feature map. The normalized thermal infrared fusion feature map is processed by a feed-forward neural network and then added element-wise and normalized again with the normalized thermal infrared fusion feature map to finally output an optimized and enhanced interactive thermal infrared feature map.

[0037] Optionally, in the same-scale fusion module, first, two feature maps input at the same scale are concatenated and convolutionally downsampled to obtain grouped features. Then, adaptive average pooling operations are performed on the grouped features in the width dimension and height dimension respectively to capture spatial context information from both horizontal and vertical directions. The spatial context information is concatenated and feature fusion is performed through convolution, and then a Split operation is applied to separate the fused features back into height and width components, and a Sigmoid activation function is applied to generate spatial attention weights. The spatial attention weights are applied to the grouped features to obtain enhanced features. At the same time, local spatial features are extracted from the grouped features through convolution. The enhanced features and local spatial features are respectively subjected to global average pooling operation, Reshape operation, and Softmax to calculate the channel attention weights of the enhanced features and the channel attention weights of the local spatial features. At the same time, after the enhanced features and local spatial features are reshaped through a Reshape operation, a cross-attention operation is performed, that is, the reshaped enhanced features are multiplied by the channel attention weights of the local spatial features, and the reshaped local spatial features are multiplied by the channel attention weights of the enhanced features, and added together to obtain cross-attention. Then, the cross-attention is reshaped to the original spatial size, and a Sigmoid activation function is applied to obtain the final weights. The final weights are applied to the grouped features and recombined in the channel dimension to obtain the final fused features.

[0038] Optionally, in the cross-scale recombination module, first, two feature maps of different scales in the input are concatenated in the channel dimension to obtain a combined input of the feature maps; after convolution for dimensionality reduction, a combined input with the same shape as the original feature map is obtained; then, a random matrix with the same shape as the combined input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonalized filters; the orthogonalized filters are subjected to a pointwise multiplication operation with the combined input feature map to obtain a one-dimensional orthogonalized weight vector; a Softmax operation is performed on the orthogonalized weight vector, and a weight matrix is adaptively learned based on the current feature relationship. The gating fusion mechanism allocates weights for different scales, performs weighted combination, and then performs a residual connection, adding them to the original two features respectively to obtain the final fused feature.

[0039] Optionally, in the landing point calculation module, the bounding boxes containing explosion events and the corresponding category information in consecutive frames are acquired and stored, and the change rate and stability analysis are performed on the recorded bounding boxes and category information. The change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames are calculated. At the same time, by statistically analyzing the consecutive frame data, the abnormal results caused by instantaneous noise or misdetection are eliminated; based on the bounding boxes and category information in the candidate frames, the moment when the bounding box is first stable and the category is clear is selected as the initial explosion moment, and the bounding box in this frame is extracted as the basis for landing point calculation; using the determined bounding box, the center point is calculated as the preliminary landing point image coordinates; with the help of the additional information detected in multiple subsequent frames, the preliminary calculated landing point image coordinates are supplemented and corrected, and finally the corrected landing point image coordinates are obtained, and the corrected landing point image coordinates are output as the position basis corresponding to the initial explosion moment.

[0040] Based on the same inventive concept, the present invention discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a projectile landing point detection method based on dual optical end-to-end fusion as described above.

[0041] Based on the same inventive concept, the present invention discloses a computer-readable storage medium, on which a computer program is stored. The computer program is characterized in that it is executed by a processor to implement a projectile landing point detection method based on dual optical end-to-end fusion as described above.

[0042] Advantages: Compared with the prior art, the present invention has the following remarkable advantages. The present invention is simply arranged and utilizes the complementary characteristics of two image modalities. The infrared thermal imaging modality is more sensitive to smoke and heat radiation, reducing false alarms caused by complex background interference. The visible light modality contains more detailed information, and the combination of the two improves the performance of impact point detection. The network input of the end-to-end multi-modal fusion detection model designed in the present invention is the registered visible light image and thermal infrared image. To fully exploit the complementary information of multi-modal features and meet the requirements of real-time performance and accuracy for the impact point detection task, an adaptive fusion design is adopted throughout the stages from the dual-branch feature extraction module to the gating fusion module and then to the loss function, realizing end-to-end multi-modal and multi-scale optimization of dynamic features, preserving the uniqueness of each modal feature and achieving information complementarity and deep fusion through cross-modal interaction. Description of the Drawings

[0043] Figure 1 It is a schematic flow diagram of the present invention;

[0044] Figure 2 It is a schematic reference diagram of the actual arranged target area in the present invention;

[0045] Figure 3 It is a schematic reference projection target surface diagram in the present invention.

[0046] Figure 4 It is a schematic framework diagram of the system in the present invention;

[0047] Figure 5 It is a schematic framework diagram of the multi-modal fusion detection model in the present invention;

[0048] Figure 6 It is a schematic framework diagram of the recalibration module in the present invention;

[0049] Figure 7 It is a schematic framework diagram of the cross-modal interaction module in the present invention;

[0050] Figure 8 It is a schematic framework diagram of the same-scale fusion module in the present invention;

[0051] Figure 9 It is a schematic framework diagram of the cross-scale recombination module in the present invention. Detailed Embodiments

[0052] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] It should be understood that the present invention can be implemented in different forms and should not be construed as limited to the embodiments presented herein. On the contrary, providing these embodiments will make the disclosure thorough and complete, and will fully convey the scope of the present invention to those skilled in the art. In the drawings, for clarity, the dimensions and relative dimensions of the components may be exaggerated. The same reference numerals throughout the drawings denote the same components.

[0054] Embodiment 1

[0055] As Figure 1 shown, the present invention discloses a bullet impact point detection method based on dual - optical end - to - end fusion, including the following steps:

[0056] S1. Collect visible - light imaging data and thermal - infrared imaging data of the entire target area.

[0057] S2. Pre - process the visible - light imaging data and thermal - infrared imaging data to obtain visible - light images and thermal - infrared images with spatio - temporal registration for the corresponding target areas; specifically: configure the acquisition time sequence for the visible - light imaging data and thermal - infrared imaging data so that each frame of visible - light image and thermal - infrared image is synchronized and aligned in time, and then register and align each frame of visible - light image and thermal - infrared image in space to obtain visible - light images and thermal - infrared images with spatio - temporal registration for the corresponding target areas; the present invention uses a dual - optical camera to collect data, and the visible - light lens and thermal - infrared lens have different resolutions and field - of - view angles respectively. Spatio - temporal registration can achieve alignment in terms of resolution and field - of - view angle; however, due to the errors of imaging distortion and pseudo - coaxial error itself, there are still slight offsets in local areas of the images.

[0058] S3. Project the ground target areas on the spatio - temporally registered visible - light images and thermal - infrared images onto the corresponding reference projection target surfaces;

[0059] Specifically, it includes the following steps:

[0060] Pre - establish and store the mapping relationship of the position information between the ground target areas on the spatio - temporally registered visible - light images and thermal - infrared images and the reference projection target surfaces, where the position information is coordinate information; select several reference points in advance in the target area, the number of reference points is greater than 3, and the connection area formed by the reference points should cover the entire area of the target area;

[0061] The reference points can be selected around the target area, and the reference points have obvious features on both the visible - light image and the thermal - infrared image. The obvious features include but are not limited to morphological features, texture features, edge features or color features; alternatively, a calibration reference object can be prepared in advance, whose surface has recognizable marking features, and select the corresponding feature points;

[0062] Record the relative position information of several reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points between the target area image and the reference projection target surface.

[0063] As Figure 2 and Figure 3 shown, the height of the dual - light camera is set to H, and there is no clear height requirement for H. Considering the actual camera focal length, field of view angle, and the distance to the target area, it is better that the field of view of the dual - light camera covers the target area. Theoretically, setting the installation position high and the camera angle tilted downward at a certain inclination can obtain a better view of the target area; it is also possible to consider deploying in combination with a tethered UAV, which can hover in the air for a long time. Set reference points 51, 52, 53, and 54 around the ground ring target on the ground plane, and use the dual - light camera to collect images of the target area to form a corresponding reference projection target surface.

[0064] S4. Input the registered visible - light image and thermal - infrared image into the multi - modal fusion detection model, and judge frame by frame whether an explosion phenomenon is detected in the image to obtain the bounding box corresponding to the explosion target in the image.

[0065] As Figure 5 shown, the multi - modal fusion detection model includes a dual - branch feature extraction module, a gated fusion module, and a detection head. Among them, the dual - branch feature extraction module uses a dual - branch structure to process the visible - light image and the thermal - infrared image respectively. The thermal - infrared image and the visible - light image respectively pass through the convolutional layer P1 to extract the first - level thermal - infrared shallow - layer feature and the first - level visible - light shallow - layer feature. After being corrected by the recalibration module RCM, the first - level thermal - infrared shallow - layer feature and the first - level visible - light shallow - layer feature are respectively input into the convolutional layer P2 to extract the second - level thermal - infrared shallow - layer feature and the second - level visible - light shallow - layer feature. After being corrected by the recalibration module RCM, the second - level thermal - infrared shallow - layer feature and the second - level visible - light shallow - layer feature are respectively input into the convolutional layer P3 to extract the third - level thermal - infrared deep - layer feature and the third - level visible - light deep - layer feature. After the third - level thermal - infrared deep - layer feature and the third - level visible - light deep - layer feature are interacted through the cross - modal interaction module CIM, they are respectively input into the convolutional layer P4 to extract the fourth - level thermal - infrared deep - layer feature and the fourth - level visible - light deep - layer feature. After the fourth - level thermal - infrared deep - layer feature and the fourth - level visible - light deep - layer feature are interacted through the cross - modal interaction module CIM, they are respectively input into the convolutional layer P5 to extract the fifth - level thermal - infrared deep - layer feature and the fifth - level visible - light deep - layer feature. In the present invention, the convolutional layer P1 is a C3 network, the convolutional layer P2 is a C3 + C2f network, the convolutional layer P3 is a C3 + C2f network, the convolutional layer P4 is a C3 + C2f network, and the convolutional layer P5 is a C3 + C2f + SPPF network.

[0066] The network input of the designed end-to-end multi-modal fusion detection model is the registered visible light image and thermal infrared image. To fully exploit the complementary information of multi-modal features and meet the requirements of real-time performance and accuracy for the landing point detection task, an adaptive fusion design is adopted throughout the entire process from the dual-branch feature extraction module to the gated fusion module and then to the loss function, realizing the end-to-end multi-modal and multi-scale optimization of dynamic features. This not only preserves the uniqueness of each modal feature but also achieves information complementarity and deep fusion through cross-modal interaction.

[0067] The dual-branch feature extraction module is the foundation of the multi-modal fusion detection model in the entire detection. It extracts visible light image features and thermal infrared image features respectively. The dual-branch feature extraction module is based on the YOLO-backbone variant structure. This module makes full use of the complementary information of the dual-modal data through offset learning and attention mechanism to achieve cross-modal feature alignment and interaction from shallow to deep layers. To address the possible spatial deviation and local geometric distortion problems between modalities, a recalibration module RCM is introduced in the shallow layer to achieve spatial alignment through two-level calibration. In the deep layer, a cross-modal interaction module CIM based on linear attention is adopted. Through linear transformation, cross-attention calculation, and feed-forward processing, it realizes feature complementarity and efficient interaction at the semantic level, supplementing the information of different modalities into each other's modalities. The network architecture of the dual-branch feature extraction module utilizes multi-level feature information. Through dynamic feature alignment and interactive fusion, it obtains fusion features that are superior to traditional methods in both edge details and high-level semantics, providing more accurate and robust feature support for subsequent object detection.

[0068] As Figure 6 shown, in the recalibration module RCM, for the thermal infrared feature map F inf and the visible light feature map F vis are first concatenated in the channel dimension to obtain a joint input feature map. Let the joint input feature map F joint ∈R b×2c×h×w . For simplicity of representation, without considering the batch dimension, it is denoted as having a shape of (2c, h, w), where 2c is the number of channels. The joint input feature map undergoes a shared convolution module to achieve feature dimensionality reduction and obtain a general representation, providing general information for subsequent offset prediction. After the shared convolutional layer, it is divided into a thermal infrared offset prediction sub-network and a visible light offset prediction sub-network for calculating the offsets of each modality. For the thermal infrared offset prediction sub-network, the general representation is first adjusted in the channel dimension through a 1×1 convolution, and then the number of parameters and computational amount are reduced through depthwise separable convolution and restored to the required number of channels, thereby outputting the thermal infrared offset vector δ inf , whose dimension is b×2k 2×h×w, where k is the size of the convolutional sampling window. For example, if the sampling window is 3×3, then k = 3, to obtain an offset estimation with low computational cost. Similarly, for the visible light offset prediction sub-network, the general representation is further adjusted in the channel dimension through 1×1 convolution, and then the visible light offset vector δ is output through depthwise separable convolution vis , to obtain an offset estimation with low computational cost; the thermal infrared offset vector δ inf is used as the sampling point offset and applied to the thermal infrared modality feature map F inf to obtain the corrected aligned thermal infrared modality feature map The visible light offset vector δ vis is used as the sampling point offset and applied to the visible light modality feature map F vis to obtain the corrected aligned visible light modality feature map

[0069] As Figure 7 shown, the Cross-modal Interaction Module (CIM) obtains a feature table with highly complementary semantic information between modalities. In the cross-modal interaction module, first, the thermal infrared feature map F inf and the visible light feature map F vis are linearly mapped respectively to generate (Q inf , K inf , V inf ) and (Q vis , K vis , V vis) The vector; then through the cross-attention operation, the Q vector of the thermal infrared feature is subjected to an attention operation with the visible light K vector and the visible light V vector of the visible light feature to obtain the thermal infrared fusion feature; the Q vector of the visible light feature is subjected to an attention operation with the thermal infrared K vector and the thermal infrared V vector of the thermal infrared feature to obtain the visible light fusion feature. To enhance the stability and feature expression ability of the model, a residual connection is introduced after the cross-attention output to ensure that the feature representation of the original modality is not lost during the information fusion process. The visible light fusion feature and the visible light feature are added element-wise and normalized to output the visible light fusion normalized feature. The visible light fusion normalized feature is processed by a feed-forward neural network and then added element-wise with the visible light fusion normalized feature again and normalized, and finally the optimized and enhanced interactive visible light feature is output. The thermal infrared fusion feature and the thermal infrared feature are added element-wise and normalized to output the thermal infrared fusion normalized feature. The thermal infrared fusion normalized feature is processed by a feed-forward neural network and then added element-wise with the thermal infrared fusion normalized feature again and normalized, and finally the optimized and enhanced interactive thermal infrared feature is output. Here, the feed-forward neural network is a feed-forward neural network composed of a linear transformation and a ReLU activation function.

[0070] In the gated fusion module, the thermal infrared deep features and the visible light deep features of the third, fourth, and fifth levels are respectively input into the same-scale fusion module CMFM for multi-scale fusion to obtain the third, fourth, and fifth level fusion features. The fifth level fusion feature is obtained through the convolution operation C2f to obtain the fifth level fusion intermediate feature. The fifth level fusion intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the fourth level fusion feature to obtain the fourth level recombination intermediate feature. The fourth level recombination intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the third level fusion feature to obtain the third level recombination feature. The third level recombination feature is downsampled, the fifth level fusion intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the fourth level recombination intermediate feature to obtain the fourth level recombination feature. The fourth level recombination feature is downsampled and input into the cross-scale recombination module OAGM together with the fifth level fusion intermediate feature to obtain the fifth level recombination feature; the third, fourth, and fifth level recombination features are respectively input into the detection head, and then the class information and localization information corresponding to the explosion target are output through post-processing and drawn in the original image to obtain an image with the explosion target bounding box.

[0071] In the gated fusion module, the multi-scale bimodal features are further fused and strengthened. In the gated fusion module, first, at the same scale, the same-scale fusion module CMFM is used to achieve the fusion output of visible light and thermal infrared features; second, at different scales, the cross-scale recombination module OAGM is used to fuse the features of each layer to achieve cross-scale information integration. The gated fusion module adopts the same-scale interaction and cross-scale weighting strategies, effectively coordinating the complementary relationship between multi-scale multi-modal features, and maintaining the flexibility of feature allocation through the gated fusion mechanism, adaptively capturing target features of different sizes. The gated fusion module realizes the efficient interaction of information between different modalities and scales while maintaining light weight, significantly improving the expression ability of the fused features. The efficient combination of multi-modal, multi-scale, attention mechanism and gated mechanism in the present invention enables the model to more accurately capture the complementary information of multi-modal data and achieve robust detection results in complex scenarios.

[0072] As Figure 8 shown, in the same-scale fusion module CMFM, at each scale, the module is used to realize the interaction and integration of two-modal features. For the same-scale input feature maps F vis and F inf , the size is [B, C, H, W], where B represents the batch size, C represents the number of channels, and H and W represent the height and width of the feature map respectively. In the same-scale fusion module CMFM, first, the input is concatenated and reduced in dimension by 1×1 convolution, and divided into G groups in the channel dimension to obtain the grouped feature X g ; then, adaptive average pooling operations are performed on X g in the width dimension and height dimension respectively to capture spatial context information from the horizontal and vertical directions; then the spatial context information is concatenated and feature fusion is performed through 1×1 convolution, and then the Split operation is used to separate the fused features back into the height and width components, and the Sigmoid activation function σ is applied to generate the spatial attention weight A spatial ; the spatial attention weight A spatial is applied to the grouped feature X g to obtain the enhanced feature X1; similarly, the grouped feature X g can extract the local spatial feature X2 through 3×3 convolution; the global average pooling operation, Reshape operation and Softmax are respectively performed on the enhanced feature X1 and the local spatial feature X2 to calculate their respective channel attention weights A ch1 and A ch2 ; at the same time, after the enhanced feature X1 and the local spatial feature X2 are reshaped, the cross-attention operation is performed, that is, the reshaped enhanced feature X1 is multiplied by the channel attention weight A ch2 , the reshaped local spatial feature X2 is multiplied by the channel attention weight A ch1 , and added to obtain the cross-attention Across ; Then, the cross-attention is reshaped to the original spatial dimension, and the Sigmoid activation function is applied to obtain the final weight A; the final weight is applied to the grouped feature X g , and they are recombined in the channel dimension to obtain the final fused feature F fuse .

[0073] As Figure 9 shown, the cross-scale recombination module is the Orthogonal Attention Gated Module (OAGM), which combines the orthogonal attention mechanism and the gated network to achieve cross-scale adaptive feature fusion; the features from different scales in the cross-scale recombination module are respectively denoted as F1 and F2, first concatenated in the channel dimension to obtain the joint input of the feature map; after dimensionality reduction by 1×1 convolution, the joint input with the same shape as the original feature map is obtained; then a random matrix with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonal filters, denoted as orthogonal filter G ortho ; The orthogonal filter performs a pointwise multiplication operation with the joint input feature map to obtain a one-dimensional orthogonal weight vector as the orthogonal feature of each channel; the Softmax operation is performed on the orthogonal weight vector, and the weight matrix A is adaptively learned according to the current feature relationship. The value range of the weight matrix A is [0,1]. The gated fusion mechanism distributes the weights of different scales. Denote the weight corresponding to the feature F1 as A and the weight corresponding to the feature F2 as 1 - A. After weighted combination, residual connection is performed and added to the original feature F1 and the original feature F2 to obtain the final fusion result F. The gated fusion mechanism of the cross-scale recombination module can adaptively adjust the contributions of features at different scales to detection, taking into account both shallow details and high-level semantics; the fused features are further processed through a series of convolutional layers and cross-layer fusion to enhance the robustness of feature expression.

[0074] During the optimization process of the multi-modal fusion detection model of the present invention, an improved loss function is adopted. The improved loss function not only considers the target box overlap degree, but also adds auxiliary bounding boxes and vector angle constraints to optimize for targets of different sizes; while ensuring the convergence speed of targets at each scale, it effectively improves the detection accuracy and the accuracy of boundary localization, enhances the perception ability of targets of different sizes. The improved loss function is as follows:

[0075] Loss = λ1L cls + λ2L Inner-SIoU + λ3L DFL

[0076] L cls is the classification loss, and L Inner-SIoU is the localization loss, and LDFL It is the distribution focal loss, and λ1, λ2, and λ3 are the weight coefficients corresponding to each loss.

[0077] S5. Determine the landing point image coordinates through the explosion bounding box. According to the coordinate mapping relationship between the landing point image coordinates and the predefined reference points on the target area image and the reference projection target surface, map the calculated landing point image coordinates to the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface;

[0078] Adopt the landing point calculation algorithm. Based on the bounding box on the image corresponding to the explosion phenomenon provided by the fusion detection algorithm, correct the position of the landing point, and obtain the image coordinates of the landing point from the explosion area; According to the pre-stored coordinate mapping relationship of the reference points on the target area image and the reference projection target surface, map the image coordinates of the landing point to the reference projection target surface, and obtain the position coordinate information of the landing point in the target area, then the specific coordinates of the landing point compared with the target area plane can be known;

[0079] The landing point detection algorithm includes the following steps: Obtain and store the bounding boxes and corresponding category information of the candidate areas containing explosion phenomena in consecutive frames for subsequent judgment and statistics of the landing points; Analyze the change rate and stability of the recorded bounding boxes and category information, calculate the change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames, and at the same time, by statistically analyzing the consecutive frame data, eliminate the abnormal frames caused by instantaneous noise or misdetection to ensure that only the detection results truly reflecting the initial moment of the explosion are retained; According to the bounding boxes and category information in the candidate frames, select the moment when a certain frame first meets the conditions of a stable bounding box and a clear category as the initial moment of the explosion, and extract the bounding box in this frame as the basis for calculating the landing point; Use the determined bounding box to calculate the center point as the preliminary landing point image coordinates; With the help of the additional information detected in multiple subsequent frames, supplement and correct the coordinates calculated initially, such as analyzing the change trend of the landing point coordinates in consecutive frames through linear regression, or performing mean operation on the coordinates of each frame, or learning and fitting through a machine learning model; Finally, obtain the corrected accurate image coordinates, and output the finally calculated landing point image coordinates as the position basis corresponding to the initial moment of the explosion.

[0080] Embodiment 2

[0081] As Figure 4 shown, the present invention discloses a projectile landing point detection system based on dual optical end-to-end fusion, including: a data acquisition module for acquiring visible light imaging data and thermal infrared imaging data of the entire target area.

[0082] A preprocessing module for preprocessing visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target area. Specifically: configure the acquisition time sequence for the visible light imaging data and the thermal infrared imaging data so that each frame of visible light image and thermal infrared image is synchronously aligned in time, and then register and align each frame of visible light image and thermal infrared image spatially to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target area. In the present invention, data is collected using a dual-camera. The visible light lens and the thermal infrared lens have different resolutions and field of views for their respective imaging. Spatio-temporal registration can achieve alignment in terms of resolution and field of view. However, due to the errors of imaging distortion and pseudo-coaxial errors, there are still slight offsets in local areas of the images.

[0083] A target area mapping module for projecting the ground target area on the spatio-temporally registered visible light image and thermal infrared image onto the corresponding reference projection target surface. A mapping relationship of the position information between the ground target area on the spatio-temporally registered visible light image and thermal infrared image and the reference projection target surface is pre-established and stored in the target area mapping module, and the position information is coordinate information.

[0084] Select a number of reference points in the target area in advance. The number of reference points is greater than 3, and the connected area formed by the reference points should cover the entire area of the target area. The reference points can be selected around the target area, and the reference points exhibit obvious features in both the visible light and infrared thermal imaging images. The obvious features include, but are not limited to, morphological features, texture features, edge features, or color features. Or a calibration reference object is pre-made, with recognizable marking features on its surface, and corresponding feature points are selected.

[0085] Record the relative position information of a number of reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information. Determine the image coordinate positions of a number of reference points in the target area image, denoted as image position information. Based on the target position information and the image position information, obtain the coordinate mapping relationship of a number of reference points between the target area image and the reference projection target surface.

[0086] An end-to-end fusion detection module for inputting the registered visible light image and thermal infrared image into a multi-modal fusion detection model, and judging frame by frame whether an explosion phenomenon is detected in the image to obtain the bounding box corresponding to the explosion target in the image.

[0087] The multi-modal fusion detection model includes a dual-branch feature extraction module, a gated fusion module, and a detection head. The dual-branch feature extraction module uses a dual-branch structure to process visible light images and thermal infrared images respectively. The thermal infrared image and the visible light image respectively pass through the convolutional layer P1 to extract the first-level thermal infrared shallow features and the first-level visible light shallow features. After being corrected by the recalibration module RCM, the first-level thermal infrared shallow features and the first-level visible light shallow features are respectively input into the convolutional layer P2 to extract the second-level thermal infrared shallow features and the second-level visible light shallow features. After being corrected by the recalibration module RCM, the second-level thermal infrared shallow features and the second-level visible light shallow features are respectively input into the convolutional layer P3 to extract the third-level thermal infrared deep features and the third-level visible light deep features. After the third-level thermal infrared deep features and the third-level visible light deep features are interacted through the cross-modal interaction module CIM, they are respectively input into the convolutional layer P4 to extract the fourth-level thermal infrared deep features and the fourth-level visible light deep features. After the fourth-level thermal infrared deep features and the fourth-level visible light deep features are interacted through the cross-modal interaction module CIM, they are respectively input into the convolutional layer P5 to extract the fifth-level thermal infrared deep features and the fifth-level visible light deep features. In the present invention, the convolutional layer P1 is a C3 network, the convolutional layer P2 is a C3 + C2f network, the convolutional layer P3 is a C3 + C2f network, the convolutional layer P4 is a C3 + C2f network, and the convolutional layer P5 is a C3 + C2f + SPPF network.

[0088] The network input of the designed end-to-end multi-modal fusion detection model is the registered visible light image and thermal infrared image. To fully exploit the complementary information of multi-modal features and meet the requirements of real-time performance and accuracy for the landing point detection task, an adaptive fusion design is adopted throughout the entire stage from the dual-branch feature extraction module to the gated fusion module and then to the loss function, realizing the end-to-end multi-modal and multi-scale optimization of dynamic features, which not only retains the uniqueness of each modal feature but also achieves information complementarity and deep fusion through cross-modal interaction.

[0089] The dual-branch feature extraction module is the foundation of the multi-modal fusion detection model in the whole detection, which extracts visible light image features and thermal infrared image features respectively. The dual-branch feature extraction module is based on the YOLO-backbone variant structure. This module makes full use of the complementary information of the dual-modal data through offset learning and attention mechanism to achieve cross-modal feature alignment and interaction from shallow to deep layers. Aiming at the possible spatial deviation and local geometric distortion problems between modalities, a recalibration module RCM is introduced in the shallow layer to achieve spatial alignment through two-level calibration. In the deep layer, a cross-modal interaction module CIM based on linear attention is adopted. Through linear transformation, cross-attention calculation and feed-forward processing, feature complementarity and efficient interaction at the semantic level are realized, and the information of different modalities is supplemented to the other modality. The network architecture of the dual-branch feature extraction module utilizes multi-level feature information, and through dynamic feature alignment and interaction fusion, fusion features that are superior to traditional methods in both edge details and high-level semantics are obtained, providing more accurate and robust feature support for subsequent object detection.

[0090] In the recalibration module RCM, for the thermal infrared feature map F inf and the visible light feature map F vis are first concatenated in the channel dimension to obtain a joint input feature map. Let the joint input feature map F joiny ∈R b×2c×h×w , for the convenience of representation, the batch dimension is not considered, and the shape is denoted as (2c, h, w), where 2c is the number of channels; the joint input feature map undergoes a shared convolution module to achieve feature dimensionality reduction and obtain a general representation, providing general information for subsequent offset prediction; after the shared convolution layer, it is divided into a thermal infrared offset prediction sub-network and a visible light offset prediction sub-network for calculating the offsets of each modality; for the thermal infrared offset prediction sub-network, the general representation is first adjusted in the channel dimension through a 1×1 convolution, and then the number of parameters and computational amount are reduced through a depthwise separable convolution and restored to the required number of channels, thereby outputting the thermal infrared offset vector δ inf , whose dimension is b×2k 2 ×h×w, where k is the convolution sampling window size. For example, if the sampling window is 3×3, then k = 3, obtaining a low-computation offset estimate; similarly, for the visible light offset prediction sub-network, the general representation is adjusted in the channel dimension through a 1×1 convolution, and then the visible light offset vector δ vis is output through a depthwise separable convolution, obtaining a low-computation offset estimate; taking the thermal infrared offset vector δ inf as the sampling point offset and applying it to the thermal infrared modality feature map F inf , the corrected aligned thermal infrared modality feature map is obtained Taking the visible light offset vector δvis As the sampling point offset, it is applied to the visible light modality feature map F vis , and the corrected aligned visible light modality feature map is obtained

[0091] The Cross-modal Interaction Module (CIM) obtains a feature table with highly complementary semantic information between modalities. In the cross-modal interaction module, first, the thermal infrared feature map F inf and the visible light feature map F vis are linearly mapped respectively to generate the (Q inf , K inf , V inf ) and (Q vis , K vis , V vis ) vectors corresponding to the features of each modality; then, through the cross-attention operation, the Q vector of the thermal infrared feature is subjected to an attention operation with the visible light K vector and the visible light V vector of the visible light feature to obtain the thermal infrared fusion feature; the Q vector of the visible light feature is subjected to an attention operation with the thermal infrared K vector and the thermal infrared V vector of the thermal infrared feature to obtain the visible light fusion feature. To enhance the stability and feature expression ability of the model, a residual connection is introduced after the cross-attention output to ensure that the feature representation of the original modality is not lost during the information fusion process. The visible light fusion feature and the visible light feature are added element-wise and normalized to output the visible light fusion normalized feature. After being processed by the feed-forward neural network, the visible light fusion normalized feature is added element-wise with the visible light fusion normalized feature again and normalized, and finally, the optimized and enhanced interactive visible light feature is output. The thermal infrared fusion feature and the thermal infrared feature are added element-wise and normalized to output the thermal infrared fusion normalized feature. After being processed by the feed-forward neural network, the thermal infrared fusion normalized feature is added element-wise with the thermal infrared fusion normalized feature again and normalized, and finally, the optimized and enhanced interactive thermal infrared feature is output. Here, the feed-forward neural network is a feed-forward neural network composed of a single linear transformation and a ReLU activation function.

[0092] In the gated fusion module, the thermal infrared deep features and visible light deep features of the third, fourth, and fifth levels are respectively input into the same-scale fusion module CMFM for multi-scale fusion to obtain the third, fourth, and fifth level fusion features. The fifth level fusion feature is convolved by the convolution operation C2f to obtain the fifth level fusion intermediate feature. The fifth level fusion intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the fourth level fusion feature to obtain the fourth level recombination intermediate feature. The fourth level recombination intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the third level fusion feature to obtain the third level recombination feature. The third level recombination feature is downsampled, the fifth level fusion intermediate feature is upsampled and input into the cross-scale recombination module OAGM together with the fourth level recombination intermediate feature to obtain the fourth level recombination feature. The fourth level recombination feature is downsampled and input into the cross-scale recombination module OAGM together with the fifth level fusion intermediate feature to obtain the fifth level recombination feature. The third, fourth, and fifth level recombination features are respectively input into the detection head, and then the class information and localization information corresponding to the explosion target are output after post-processing and drawn in the original image to obtain an image with the bounding box of the explosion target.

[0093] In the gated fusion module, the multi-modal features of each scale are further fused and enhanced. In the gated fusion module, first, the fusion output of visible light and thermal infrared features is realized through the same-scale fusion module CMFM at the same scale. Secondly, the features of each layer are fused using the cross-scale recombination module OAGM at different scales to achieve cross-scale information integration. The gated fusion module adopts the same-scale interaction and cross-scale weighting strategies, effectively coordinating the complementary relationship between the multi-modal features of each scale, and maintaining the flexibility of feature allocation through the gated fusion mechanism, adaptively capturing the target features of different sizes. The gated fusion module realizes the efficient interaction of information between different modalities and different scales while maintaining lightweight, significantly improving the expression ability of the fusion features. The efficient combination of multi-modal, multi-scale, attention mechanism, and gated mechanism in the present invention enables the model to more accurately capture the complementary information of multi-modal data and achieve robust detection results in complex scenarios.

[0094] In the same-scale fusion module CMFM, the interaction and integration of the two-modal features are realized using the module at each scale. For the same-scale input feature maps F vis and F inf , the size is [B, C, H, W], where B represents the batch size, C represents the number of channels, and H and W respectively represent the height and width of the feature map. In the same-scale fusion module CMFM, first, the input is concatenated and reduced in dimension by 1×1 convolution, and divided into G groups in the channel dimension to obtain the grouped feature X g ; then X gAdaptive average pooling operations are performed separately in the width dimension and the height dimension to capture spatial context information in the horizontal and vertical directions; then the spatial context information is concatenated and feature fusion is performed through 1×1 convolution, and then the Split operation separates the fused features back into the height and width components, and the Sigmoid activation function σ is applied to generate the spatial attention weight A spatial ; Apply the spatial attention weight A spatial to the grouped feature X g to obtain the enhanced feature X1; Similarly, the grouped feature X g can extract local spatial features X2 through 3×3 convolution; The global average pooling operation, Reshape operation, and Softmax are respectively performed on the enhanced feature X1 and the local spatial feature X2 to calculate their respective channel attention weights A ch1 and A ch2 ; At the same time, after performing the Reshape operation on the enhanced feature X1 and the local spatial feature X2, the cross-attention operation is performed, that is, the reshaped enhanced feature X1 is multiplied by the channel attention weight A ch2 , the reshaped local spatial feature X2 is multiplied by the channel attention weight A ch1 , and added to obtain the cross-attention A cross ; Then the cross-attention is Reshaped to the original spatial size, and the Sigmoid activation function is applied to obtain the final weight A; The final weight is applied to the grouped feature X g , and recombined in the channel dimension to obtain the final fused feature F fuse .

[0095] The cross-scale recombination module is the Orthogonal Attention Gated Module (OAGM), which combines the orthogonal attention mechanism and the gated network to achieve cross-scale adaptive feature fusion; The features from different scales in the cross-scale recombination module are respectively represented as F1 and F2, first concatenated in the channel dimension to obtain the joint input of the feature map; After 1×1 convolution for dimensionality reduction, a joint input with the same shape as the original feature map is obtained; Then a random matrix with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonal filters, denoted as the orthogonal filter G ortho; The orthogonalization filter performs a pointwise multiplication operation with the combined input feature map to obtain a one-dimensional orthogonalization weight vector, which serves as the orthogonalized feature for each channel. A Softmax operation is performed on the orthogonalization weight vector to adaptively learn the weight matrix A based on the current feature relationship. The value range of the weight matrix A is [0, 1]. The gated fusion mechanism assigns weights to different scales. Denote the weight corresponding to feature F1 as A and the weight corresponding to feature F2 as 1 - A. After weighted combination, a residual connection is performed and added to the original feature F1 and the original feature F1 to obtain the final fusion result F. The gated fusion mechanism of the cross-scale recombination module can adaptively adjust the contributions of different-scale features to detection, taking into account both shallow details and high-level semantics; the fused features are further processed through a series of convolutional layers and cross-layer fusion to enhance the robustness of feature expression.

[0096] During the optimization process of the multi-modal fusion detection model of the present invention, an improved loss function is adopted. The improved loss function not only considers the target box overlap degree but also adds auxiliary bounding boxes and vector angle constraints to optimize for targets of different sizes; while ensuring the convergence speed of targets at each scale, it effectively improves the detection accuracy and the accuracy of boundary localization, enhances the perception ability for targets of different sizes. The improved loss function is:

[0097] Loss = λ1L cls + λ2L Inner-SIoI + λ3L DFL

[0098] L cls is the classification loss, L Inner-SIoU is the localization loss, L DFL is the distribution focal loss, and λ1, λ2, and λ3 are the weight coefficients corresponding to each loss.

[0099] The landing point calculation module is used to calculate and correct the position of the landing point according to the explosion bounding box by using the landing point calculation algorithm, and determine the landing point image coordinates from the explosion area; according to the coordinate mapping relationship of the reference point on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface, obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface. In the landing point calculation module, the bounding boxes and corresponding category information of the candidate areas containing explosion phenomena in consecutive frames are acquired and stored, the change rate and stability of the recorded bounding boxes and category information are analyzed, the change trends of the positions, sizes and category confidence degrees of the bounding boxes between frames are calculated, and at the same time, by statistically analyzing the consecutive frame data, the abnormal frames caused by instantaneous noise or false detection are eliminated; according to the bounding boxes and category information in the candidate frames, select the moment when a certain frame first meets the conditions of stable bounding box and clear category as the explosion initial moment, and extract the bounding box in this frame as the basis for landing point calculation; use the determined bounding box to calculate the center point as the preliminary landing point image coordinates; with the help of the additional information detected in multiple subsequent frames, supplement and correct the preliminary calculated landing point image coordinates, finally obtain the corrected landing point image coordinates, and output the corrected landing point image coordinates as the position basis corresponding to the explosion initial moment.

[0100] Embodiment 3

[0101] Another embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a method for detecting the landing point of a projectile based on dual optical end-to-end fusion as described above.

[0102] The electronic device may include: a processor, a memory, a bus, and a communication interface. The processor, the communication interface, and the memory are connected through the bus; a computer program executable on the processor is stored in the memory, and when the processor runs the computer program, it executes a method for detecting the landing point of a projectile based on dual optical end-to-end fusion provided in any one of the foregoing embodiments of the present invention.

[0103] Among them, the memory may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface (which can be wired or wireless), the communication connection between this device network element and at least one other network element is realized, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used.

[0104] The bus can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. Among them, the memory is used to store programs. After receiving the execution instruction, the processor executes the program. Any of the embodiments of the method for detecting the projectile impact point based on dual optical end-to-end fusion disclosed in the foregoing embodiments of the present invention can be applied to or implemented by the processor.

[0105] The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above-mentioned processor can be a general-purpose processor, which may include a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed by the hardware decoding processor, or executed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0106] The electronic device provided by the embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the method adopted, run, or implemented by it.

[0107] Embodiment 4

[0108] Another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program is executed by a processor to implement the method of any of the above embodiments. The computer-readable storage medium is an optical disc, on which a computer program (i.e., a program product) is stored. When the computer program is run by the processor, it will execute the method provided by any of the foregoing embodiments.

[0109] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other optical and magnetic storage media, which will not be elaborated here one by one.

[0110] The computer-readable storage medium provided by the above embodiments of the present application and the method provided by the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run, or implemented by the application programs stored therein.

Claims

1. A bullet impact point detection method based on dual optical end-to-end fusion, characterized in that It includes the following steps: S1. Collect visible light imaging data and thermal infrared imaging data of all target areas; S2. Preprocess the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas; S3. Project the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces; S4. Input the registered visible light images and thermal infrared images into a multi-modal fusion detection model, and frame by frame determine whether an explosion phenomenon is detected in the images, and obtain the bounding boxes corresponding to the explosion targets in the images; S5. According to the explosion bounding boxes, use a landing point calculation algorithm to calculate and correct the position of the landing point, and determine the landing point image coordinates in the explosion area; according to the coordinate mapping relationship of the reference points on the target area image and the reference projection target surface, map the landing point image coordinates to the reference projection target surface to obtain the position coordinate information of the landing point within the reference projection target surface, and record and output the relative position information of the landing point on the reference projection target surface.

2. The method for detecting the projectile impact point based on dual optical end-to-end fusion according to claim 1, wherein The specific steps of step S2 include the following steps: Configure the acquisition time sequence for the visible light imaging data and thermal infrared imaging data, so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then register and align each frame of visible light image and thermal infrared image in space to obtain visible light images and thermal infrared images with spatio-temporal registration for the corresponding target areas.

3. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 1, characterized in that: The specific steps of step S3 include the following steps: Pre-establish and store the mapping relationship of the position information between the ground target areas on the spatio-temporally registered visible light images and thermal infrared images and the reference projection target surface, and the position information is coordinate information; Select several reference points in the target area in advance. The number of reference points is greater than 3, and the connected area formed by the reference points should cover the entire area of the target area; the reference points can be selected around the target area, and the reference points show obvious features in both visible light and thermal infrared images. The obvious features include but are not limited to morphological features, texture features, edge features or color features; or pre-fabricate a reference object for calibration, whose surface has recognizable marking features, and select the corresponding feature points; Record the relative position information of several reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as target position information; determine the image coordinate positions of several reference points in the target area image, denoted as image position information; according to the image position information and the target position information, obtain the coordinate mapping relationship of several reference points on the target area image and the reference projection target surface.

4. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 1, characterized in that, The multimodal fusion detection model includes a dual-branch feature extraction module, a gated fusion module, and a detection head. The dual-branch feature extraction module adopts a dual-branch structure to process visible light images and thermal infrared images respectively. The thermal infrared image and the visible light image respectively pass through a convolutional layer to extract the first-level shallow thermal infrared features and the first-level shallow visible light features. After being corrected by a recalibration module, the first-level shallow thermal infrared features and the first-level shallow visible light features are respectively input into a convolutional layer to extract the second-level shallow thermal infrared features and the second-level shallow visible light features. After being corrected by a recalibration module, the second-level shallow thermal infrared features and the second-level shallow visible light features are respectively input into a convolutional layer to extract the third-level deep thermal infrared features and the third-level deep visible light features. After the third-level deep thermal infrared features and the third-level deep visible light features are interacted through a cross-modal interaction module, they are respectively input into a convolutional layer to extract the fourth-level deep thermal infrared features and the fourth-level deep visible light features. After the fourth-level deep thermal infrared features and the fourth-level deep visible light features are interacted through a cross-modal interaction module, they are respectively input into a convolutional layer to extract the fifth-level deep thermal infrared features and the fifth-level deep visible light features; In the gated fusion module, the third-level, fourth-level, and fifth-level deep thermal infrared features and deep visible light features are respectively input into the same-scale fusion module for multi-scale fusion to obtain the third-level, fourth-level, and fifth-level fusion features. The fifth-level fusion feature is obtained through a convolutional operation to obtain the fifth-level fusion intermediate feature. The fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level fusion feature to obtain the fourth-level recombination intermediate feature. The fourth-level recombination intermediate feature is upsampled and input into the cross-scale recombination module together with the third-level fusion feature to obtain the third-level recombination feature. The third-level recombination feature is downsampled, the fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level recombination intermediate feature to obtain the fourth-level recombination feature. The fourth-level recombination feature is downsampled and input into the cross-scale recombination module together with the fifth-level fusion intermediate feature to obtain the fifth-level recombination feature; The third-level, fourth-level, and fifth-level recombination features are respectively input into the detection head, and then the category information and localization information corresponding to the explosion target are output through post-processing and drawn in the original image to obtain an image with the bounding box of the explosion target.

5. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 4, characterized in that In the recalibration module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input feature map; the joint input feature map undergoes shared convolution to achieve feature dimensionality reduction to obtain a general representation; then the general representation undergoes depthwise separable convolution operation to output a thermal infrared offset vector and a visible light offset vector. The thermal infrared offset vector is used as the sampling point offset and applied to the thermal infrared feature map to obtain the corrected aligned thermal infrared feature map; The visible light offset vector is used as the sampling point offset and applied to the visible light feature map to obtain the corrected aligned visible light feature map.

6. The method for detecting the projectile landing point based on dual optical end-to-end fusion according to claim 4, wherein In the cross-modal interaction module, first, linear mapping is performed on the thermal infrared feature map and the visible light feature map respectively to generate Q, K, and V vectors corresponding to the features of each modality. Then, through cross-attention operation, the Q vector of the thermal infrared feature map is subjected to attention calculation with the K vector and V vector of the visible light feature map to obtain the thermal infrared fusion feature map; the Q vector of the visible light feature map is subjected to attention calculation with the K vector and V vector of the thermal infrared feature map to obtain the visible light fusion feature map; the visible light fusion feature map and the visible light feature map are added element by element and normalized to output the visible light fusion normalized feature map. The visible light fusion normalized feature map is processed by a feed-forward neural network and then added element by element and normalized again with the visible light fusion normalized feature map to finally output the optimized and enhanced interactive visible light feature map; the thermal infrared fusion feature map and the thermal infrared feature map are added element by element and normalized to output the thermal infrared fusion normalized feature map. The thermal infrared fusion normalized feature map is processed by a feed-forward neural network and then added element by element and normalized again with the thermal infrared fusion normalized feature map to finally output the optimized and enhanced interactive thermal infrared feature map.

7. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 4, characterized in that, In the same-scale fusion module, first, two feature maps input at the same scale are concatenated and convolutionally downsampled to obtain grouped features; then, adaptive average pooling operations are performed on the grouped features in the width dimension and the height dimension respectively to capture spatial context information from both the horizontal and vertical directions; then, the spatial context information is concatenated and feature fusion is performed through convolution, and then the Split operation is performed to separate the fused features back into height and width components, and the Sigmoid activation function is applied to generate spatial attention weights; the spatial attention weights are applied to the grouped features to obtain enhanced features; at the same time, local spatial features are extracted from the grouped features through convolution; The channel attention weights of the enhanced features and the channel attention weights of the local spatial features are calculated respectively after global average pooling operation, Reshape operation, and Softmax on the enhanced features and the local spatial features; at the same time, after the Reshape reshaping operation is performed on the enhanced features and the local spatial features, a cross-attention operation is performed, that is, the reshaped enhanced features are multiplied by the channel attention weights of the local spatial features, and the reshaped local spatial features are multiplied by the channel attention weights of the enhanced features, and added to obtain the cross-attention; then, the cross-attention is Reshaped to the original spatial size, and the Sigmoid activation function is applied to obtain the final weights; the final weights are applied to the grouped features and recombined in the channel dimension to obtain the final fused features.

8. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 4, characterized in that, In the cross-scale recombination module, first, two feature maps of different scales input are concatenated in the channel dimension to obtain a joint input of the feature maps; after convolutional downsampling, a joint input with the same shape as the original feature map is obtained; then, a random matrix with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonal filters; the orthogonal filters are subjected to pointwise product operation with the joint input feature map to obtain a one-dimensional orthogonal weight vector; Perform a Softmax operation on the orthogonalized weight vectors, adaptively learn the weight matrix based on the current feature relationships, and the gating fusion mechanism distributes weights of different scales, performs weighted combination and then residual connection, and adds them to the original two features respectively to obtain the final fused features.

9. A method for detecting the impact point of a projectile based on dual optical end-to-end fusion according to claim 1, characterized in that, The specific steps of step S5 are as follows: Obtain and store the bounding boxes containing explosion events and the corresponding class information in consecutive frames, perform rate-of-change and stability analysis on the recorded bounding boxes and class information, calculate the change trends of the positions, sizes, and class confidences of the bounding boxes between frames, and at the same time, by statistically analyzing the consecutive frame data, eliminate abnormal results caused by instantaneous noise or misdetection; According to the bounding boxes and class information in the candidate frames, screen for the moment when a frame first satisfies the conditions of a stable bounding box and a clear class as the initial explosion moment, and extract the bounding box in this frame as the basis for impact point calculation; Use the determined bounding box to calculate the center point as the preliminary impact point image coordinates; With the help of additional information from multi-frame detections in subsequent frames, supplement and correct the preliminary calculated impact point image coordinates, finally obtain the corrected impact point image coordinates, and output the corrected impact point image coordinates as the position basis corresponding to the initial explosion moment.

10. A projectile impact point detection system based on dual optical end-to-end fusion, characterized in that Including: A data acquisition module for acquiring visible light imaging data and thermal infrared imaging data of all target areas; A preprocessing module for preprocessing the visible light imaging data and thermal infrared imaging data to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas; A target area mapping module for projecting the ground target areas on the spatio-temporally registered visible light images and thermal infrared images onto the corresponding reference projection target surfaces; An end-to-end fusion detection module for inputting the registered visible light images and thermal infrared images into a multi-modal fusion detection model, and judging frame by frame whether an explosion phenomenon is detected in the images to obtain the bounding boxes corresponding to the explosion targets in the images; An impact point calculation module for calculating and correcting the position of the impact point using an impact point calculation algorithm based on the explosion bounding box, and determining the impact point image coordinates in the explosion area; According to the coordinate mapping relationship between the reference point on the target area image and the reference projection target surface, map the impact point image coordinates to the reference projection target surface to obtain the position coordinate information of the impact point within the reference projection target surface, and record and output the relative position information of the impact point on the reference projection target surface.

11. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 10, characterized in that: In the preprocessing module, the acquisition time sequence of the visible light imaging data and the thermal infrared imaging data is configured so that each frame of visible light image and thermal infrared image is synchronized and aligned in time, and then each frame of visible light image and thermal infrared image is registered and aligned in space to obtain visible light images and thermal infrared images with spatio-temporal registration of the corresponding target areas.

12. A projectile landing point detection system based on dual optical end-to-end fusion according to claim 10, characterized in that: In the target area mapping module, a mapping relationship of the position information between the ground target areas on the spatio-temporally registered visible light images and thermal infrared images and the reference projection target surface is established and stored in advance, and the position information is coordinate information; Select a number of reference points in the target area in advance. The number of reference points is greater than 3, and the connected area formed by the reference points should cover the entire area of the target area. The reference points can be selected around the target area, and the reference points show obvious features in both visible light and thermal infrared images. The obvious features include, but are not limited to, morphological features, texture features, edge features, or color features. Or prepare a calibration reference object in advance, whose surface has identifiable marking features, and select corresponding feature points. Record the relative position information of a number of reference points, and establish a reference projection target surface based on this information. Denote the position coordinates of the reference points on the reference projection target surface as the target position information. Determine the image coordinate positions of a number of reference points in the target area image, denoted as the image position information. According to the image position information and the target position information, obtain the coordinate mapping relationship of a number of reference points between the target area image and the reference projection target surface.

13. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 10, characterized in that: The multi-modal fusion detection model includes a dual-branch feature extraction module, a gated fusion module, and a detection head. Among them, the dual-branch feature extraction module uses a dual-branch structure to process visible light images and thermal infrared images respectively. The thermal infrared image and the visible light image are respectively passed through convolutional layers to extract the first-level thermal infrared shallow features and the first-level visible light shallow features. After being corrected by the recalibration module, the first-level thermal infrared shallow features and the first-level visible light shallow features are respectively input into convolutional layers to extract the second-level thermal infrared shallow features and the second-level visible light shallow features. After being corrected by the recalibration module, the second-level thermal infrared shallow features and the second-level visible light shallow features are respectively input into convolutional layers to extract the third-level thermal infrared deep features and the third-level visible light deep features. After the third-level thermal infrared deep features and the third-level visible light deep features are interacted through the cross-modal interaction module, they are respectively input into convolutional layers to extract the fourth-level thermal infrared deep features and the fourth-level visible light deep features. After the fourth-level thermal infrared deep features and the fourth-level visible light deep features are interacted through the cross-modal interaction module, they are respectively input into convolutional layers to extract the fifth-level thermal infrared deep features and the fifth-level visible light deep features. In the gated fusion module, the third-level, fourth-level, and fifth-level thermal infrared deep features and visible light deep features are respectively input into the same-scale fusion module for multi-scale fusion to obtain the third-level, fourth-level, and fifth-level fusion features. The fifth-level fusion feature is obtained through convolutional operations to get the fifth-level fusion intermediate feature. The fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level fusion feature to obtain the fourth-level recombination intermediate feature. The fourth-level recombination intermediate feature is upsampled and input into the cross-scale recombination module together with the third-level fusion feature to obtain the third-level recombination feature. The third-level recombination feature is downsampled, the fifth-level fusion intermediate feature is upsampled and input into the cross-scale recombination module together with the fourth-level recombination intermediate feature to obtain the fourth-level recombination feature. The fourth-level recombination feature is downsampled and input into the cross-scale recombination module together with the fifth-level fusion intermediate feature to obtain the fifth-level recombination feature. The third-level, fourth-level, and fifth-level recombination features are respectively input into the detection head, and then post-processed to output the category information and positioning information corresponding to the explosion target, which are drawn in the original image to obtain an image with the explosion target bounding box.

14. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 13, characterized in that: In the re - calibration module, the thermal infrared feature map and the visible light feature map are first concatenated in the channel dimension to obtain a joint input feature map. The joint input feature map undergoes shared convolution to achieve feature dimensionality reduction and obtain a general representation. Then, after depth - separable convolution operations on the general representation, a thermal infrared offset vector and a visible light offset vector are output. The thermal infrared offset vector is used as the sampling point offset and applied to the thermal infrared feature map to obtain a corrected aligned thermal infrared feature map. The visible light offset vector is used as the sampling point offset and applied to the visible light feature map to obtain a corrected aligned visible light feature map.

15. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 13, characterized in that: In the cross - modal interaction module, first, linear mappings are respectively performed on the thermal infrared feature map and the visible light feature map to generate Q, K, and V vectors corresponding to the features of each modality. Then, through cross - attention operations, the Q vector of the thermal infrared feature map is subjected to attention operations with the K and V vectors of the visible light feature map to obtain a thermal infrared fused feature map. The Q vector of the visible light feature map is subjected to attention operations with the K and V vectors of the thermal infrared feature map to obtain a visible light fused feature map. The visible light fused feature map is added element - by - element to the visible light feature map and normalized to output a visible light fused and normalized feature map. The visible light fused and normalized feature map is processed by a feed - forward neural network and then added element - by - element to the visible light fused and normalized feature map again and normalized, finally outputting an optimized and enhanced interactive visible light feature map. The thermal infrared fused feature map is added element - by - element to the thermal infrared feature map and normalized to output a thermal infrared fused and normalized feature map. The thermal infrared fused and normalized feature map is processed by a feed - forward neural network and then added element - by - element to the thermal infrared fused and normalized feature map again and normalized, finally outputting an optimized and enhanced interactive thermal infrared feature map.

16. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 13, characterized in that: In the same - scale fusion module, first, two feature maps of the same scale input are concatenated and convolutionally reduced in dimensionality to obtain grouped features. Then, adaptive average pooling operations are respectively performed on the grouped features in the width dimension and the height dimension to capture spatial context information from both the horizontal and vertical directions. Then, the spatial context information is concatenated and feature - fused through convolution, and then a Split operation is performed to separate the fused features back into height and width components, and a Sigmoid activation function is applied to generate spatial attention weights. The spatial attention weights are applied to the grouped features to obtain enhanced features. At the same time, local spatial features are extracted from the grouped features through convolution. The channel attention weights of the enhanced features and the local spatial features are calculated respectively after the global average pooling operation, the Reshape operation, and the Softmax on the enhanced features and the local spatial features; at the same time, after the enhanced features and the local spatial features are reshaped by the Reshape operation, the cross-attention operation is performed, that is, the reshaped enhanced features are multiplied by the channel attention weights of the local spatial features, and the reshaped local spatial features are multiplied by the channel attention weights of the enhanced features, and then added to obtain the cross-attention; then the cross-attention is reshaped to the original spatial size, and the Sigmoid activation function is applied to obtain the final weights; the final weights are applied to the grouped features and recombined in the channel dimension to obtain the final fused features.

17. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 13, characterized in that: In the cross-scale recombination module, first, two feature maps of different scales in the input are concatenated in the channel dimension to obtain the joint input of the feature maps; after convolution for dimensionality reduction, the joint input with the same shape as the original feature map is obtained; then a random matrix with the same shape as the joint input feature map is introduced, and Schmidt orthogonalization is performed to obtain a set of orthogonalized filters; the orthogonalized filters are subjected to a pointwise multiplication operation with the joint input feature map to obtain a one-dimensional orthogonalized weight vector; The Softmax operation is performed on the orthogonalized weight vector, and the weight matrix is adaptively learned according to the current feature relationship. The gated fusion mechanism distributes the weights of different scales, performs weighted combination and then residual connection, and adds them to the original two features respectively to obtain the final fused features.

18. A bullet impact point detection system based on dual optical end-to-end fusion according to claim 10, characterized in that: In the landing point calculation module, the bounding boxes and corresponding category information containing explosion events in consecutive frames are acquired and stored, and the change rate and stability of the recorded bounding boxes and category information are analyzed, and the change trends of the positions, sizes, and category confidence levels of the bounding boxes between frames are calculated. At the same time, by statistically analyzing the consecutive frame data, the abnormal results caused by instantaneous noise or misdetection are eliminated; according to the bounding boxes and category information in the candidate frames, the moment when a frame first satisfies the conditions of stable bounding box and clear category is selected as the initial explosion moment, and the bounding box in this frame is extracted as the basis for landing point calculation; using the determined bounding box, the center point is calculated as the preliminary landing point image coordinates; with the help of the additional information detected in multiple subsequent frames, the preliminary calculated landing point image coordinates are supplemented and corrected, and finally the corrected landing point image coordinates are obtained, and the corrected landing point image coordinates are used as the position basis corresponding to the initial explosion moment for output.

19. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement a projectile landing point detection method based on dual optical end-to-end fusion as described in any one of claims 1-9.

20. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to implement a projectile landing point detection method based on dual optical end-to-end fusion as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Detection method and device for shell explosion information and storage medium

    CN114549498A

  • Geographic positioning method for shell drop point of double-point fixed camera

    CN117392233A

Cited By

  • Lightweight human body detection and distance measurement method and system suitable for stage lamp

    CN121074953A

  • Image alignment fusion method and system based on template matching and GIFNet, and medium

    CN121147268A

  • Image alignment fusion method and system based on template matching and GIFNet, and medium

    CN121147268B

  • A power distribution network defect detection method, device, medium and equipment

    CN122510266A