Target detection methods, electronic devices and storage media

By acquiring aligned infrared, visible light, and radar signals through a common aperture device, and performing feature extraction and fusion, the detection accuracy problem caused by sensor architecture is solved, and efficient, all-weather target detection is achieved.

CN121811206BActive Publication Date: 2026-07-31QIANYUAN NATIONAL LABORATORY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QIANYUAN NATIONAL LABORATORY
Filing Date
2026-03-10
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, discrete sensor architectures lead to line-of-sight deviations and temporal jitter, decision-level fusion loses semantic cues, and data-level fusion ignores sensor differences, resulting in poor target detection accuracy.

Method used

A common aperture device is used to acquire time-aligned infrared images, visible light images, and radar signals. Feature extraction and stitching fusion are then performed. Parallel detection by multiple detectors and result integration are used to achieve a deep integration of optical and radar features.

Benefits of technology

Eliminate spatiotemporal biases at the source, improve the comprehensiveness and robustness of detection, adapt to harsh environments, and achieve all-weather target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121811206B_ABST
    Figure CN121811206B_ABST
Patent Text Reader

Abstract

This application provides a target detection method, electronic device, and storage medium. The method includes: acquiring current data output by a common aperture device; stitching and fusing infrared and visible light images and extracting features to obtain an optical feature map; extracting features from radar signals to obtain a radar feature map; pairing the features of each channel in the optical feature map and the features of each channel in the radar feature map to obtain multiple feature pairs; performing feature alignment and fusion on each feature pair to obtain fused features; distributing each fused feature to multiple detection heads; each detection head generating an original detection result based on each fused feature; and integrating and stitching the original detection results to obtain a target detection result. This application acquires and outputs current data through a common aperture device, enabling temporal and spatial alignment of the current data. By performing feature alignment and fusion on the formed feature pairs, it achieves a deep integration of optical and radar features, compensating for the deficiencies of single-modal detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of target detection technology, and more specifically, to a target detection method, electronic device, and storage medium. Background Technology

[0002] All-weather, all-time, and high-precision target detection capabilities have become a core prerequisite for the reliable operation of security systems across various fields. Complex environments continuously challenge the perception capabilities of single sensors: visible light imaging is easily limited by lighting and visibility; infrared thermal imaging loses contrast and lacks texture detail when the temperature difference between the target and background is close; while millimeter-wave or lidar possesses penetration and ranging accuracy, its spatial resolution is limited, making it difficult to support fine-grained target classification and small target identification. Therefore, fusing multi-source heterogeneous sensor information to complement each other's shortcomings has become an inevitable technical path to improve environmental adaptability and detection reliability.

[0003] Current mainstream multimodal target detection technologies generally adopt a discrete sensor architecture, meaning that visible light, infrared, and radar devices are independently installed in different physical locations. This approach aligns the coordinate systems of each sensor by calibrating parameters and uses interpolation algorithms to unify timestamps and spatial grids. Based on this, fusion strategies are mainly divided into two categories: one is decision-level fusion, where each modality runs its detection model independently, and the final result is output through weighted voting or confidence fusion; the other is data-level fusion, which involves simply concatenating the original image and one-dimensional radar signal at the input and feeding them into a unified neural network for end-to-end learning.

[0004] However, the aforementioned existing technologies have the following drawbacks: First, the discrete hardware architecture results in inherent line-of-sight bias and temporal jitter in the original data, inevitably introducing interpolation artifacts and geometric distortions in the subsequent registration process, thus destroying the accurate correspondence between multimodal data. Second, decision-level fusion loses a massive amount of intermediate semantic clues. Third, data-level fusion ignores the essential differences between visible light, infrared images, and radar signals, and the forced stitching leads to blurred learning targets for the network, easily causing modal interference and feature collapse, significantly restricting the ability to model cross-modal deep associations, and resulting in poor detection accuracy. Summary of the Invention

[0005] The purpose of this application is to address the shortcomings of the prior art by providing a target detection method, electronic device, and storage medium to solve the problem of poor target detection accuracy in the prior art.

[0006] To achieve the above objectives, the technical solution adopted in this application is as follows:

[0007] Firstly, this application provides a target detection method, the method comprising:

[0008] The common aperture device is used to acquire current data for the current scene, including infrared images, visible light images and radar signals. The common aperture device is used to output the current data that is time-aligned based on the received raw signals.

[0009] The infrared image and the visible light image are stitched together and their features are extracted to obtain an optical feature map. The radar signal is then subjected to feature extraction to obtain a radar feature map. The number of channel features in the optical feature map is the same as the number of channel features in the radar feature map.

[0010] The features of each channel in the optical feature map and the features of each channel in the radar feature map are paired to obtain multiple feature pairs. The feature pairs are then aligned and fused to obtain the fused features corresponding to each feature pair.

[0011] The fusion features are distributed to multiple detection heads, and each detection head generates an original detection result based on the fusion features. The original detection results output by each detection head are integrated and spliced ​​to obtain the target detection result of the current scene.

[0012] Optionally, the step of stitching and fusing the infrared image and the visible light image and extracting features to obtain an optical feature map includes:

[0013] The infrared image and the visible light image are respectively subjected to alignment and registration processing, size standardization processing and data normalization processing to obtain the processed infrared image and the processed visible light image.

[0014] The processed infrared image and the processed visible light image are subjected to channel stitching and fusion as well as data enhancement to obtain dual-modal optical image data.

[0015] The dual-modal optical image data is input into a deep convolutional backbone network for layer-by-layer downsampling to obtain multiple optical layer features;

[0016] Upsampling and lateral connection are performed on the optical level features to obtain a multi-scale feature pyramid. Based on the attention mechanism, feature enhancement is performed on the multi-scale feature pyramid to obtain an optical feature map.

[0017] Optionally, feature extraction is performed on the radar signal to obtain a radar feature map, including:

[0018] The radar signal is preprocessed to obtain the processed radar signal;

[0019] Multi-dimensional feature extraction is performed on the processed radar signal to obtain a radar feature matrix;

[0020] The radar feature matrix is ​​projected at multiple scales to obtain features at multiple scales.

[0021] Based on the channel attention mechanism, attention enhancement is performed on each scale feature to obtain the radar feature map.

[0022] Optionally, the step of performing multi-dimensional feature extraction on the processed radar signal to obtain a radar feature matrix includes:

[0023] The processed radar signal is subjected to time-domain feature extraction, frequency-domain feature extraction, time-frequency-domain feature extraction, and statistical feature extraction to obtain time-domain features, frequency-domain features, time-frequency-domain features, and statistical features.

[0024] The time-domain features, frequency-domain features, time-frequency-domain features, and statistical features are concatenated to obtain the concatenated features;

[0025] The spliced ​​features are then encoded to generate a radar feature matrix.

[0026] Optionally, the channel features in the optical feature map include a first optical feature, a second optical feature, a third optical feature, and a fourth optical feature, with the resolutions of the first, second, and third optical features decreasing sequentially, and the fourth optical feature being a deep semantic feature; the channel features in the radar feature map include a first radar feature, a second radar feature, a third radar feature, and a fourth radar feature, with the feature size of the first optical feature being consistent with the feature size of the first radar feature, the feature size of the second optical feature being consistent with the feature size of the second radar feature, the feature size of the third optical feature being consistent with the feature size of the third radar feature, and the feature size of the fourth optical feature being consistent with the feature size of the fourth radar feature;

[0027] The process of pairing the features of each channel in the optical feature map and each channel in the radar feature map yields multiple feature pairs, including:

[0028] The first optical feature and the first radar feature are used as a first feature pair, the second optical feature and the second radar feature are used as a second feature pair, the third optical feature and the third radar feature are used as a third feature pair, and the fourth optical feature and the fourth radar feature are used as a fourth feature pair.

[0029] Optionally, the step of performing feature alignment and fusion on each feature pair to obtain the fused feature corresponding to each feature pair includes:

[0030] The first optical feature and the first radar feature in the first feature pair are respectively subjected to feature alignment processing and cross-modal attention enhancement processing to obtain the first enhanced optical feature and the first enhanced radar feature; the first enhanced optical feature and the first enhanced radar feature are subjected to feature splicing and fusion to obtain the first spliced ​​feature; the first spliced ​​feature is weighted and calibrated based on the attention mechanism to obtain the first fused feature;

[0031] The second optical feature and the second radar feature in the second feature pair are respectively subjected to feature alignment processing, and a single-channel weight map is generated based on the obtained second processed optical feature and second processed radar feature; based on the gated fusion mechanism, the second stitching feature is determined based on the single-channel weight map, the second processed optical feature and the second processed radar feature; the second fusion feature is determined based on the second stitching feature and the second processed optical feature.

[0032] The third optical feature and the third radar feature in the third feature pair are respectively subjected to feature alignment processing, and the resulting third processed optical feature and the third processed radar feature are subjected to feature addition processing to obtain the third stitched feature; based on the gated fusion mechanism, the gated fusion feature is determined according to the third stitched feature and the third processed optical feature; the gated fusion feature and the third stitched feature are fused to obtain the adjusted feature; based on the nonlinear activation function, the third fused feature is determined according to the adjusted feature;

[0033] The fourth optical feature and the fourth radar feature in the fourth feature pair are respectively subjected to feature alignment processing, and the obtained fourth processed optical feature and fourth processed radar feature are enhanced based on the interactive attention mechanism to obtain the fourth enhanced optical feature and the fourth enhanced radar feature; the fourth enhanced optical feature and the fourth enhanced radar feature are spliced ​​and compressed to obtain the fourth spliced ​​feature; the long-range dependency relationship of the fourth spliced ​​feature is captured based on the global context fusion mechanism to obtain the fourth fused feature.

[0034] Optionally, the process of integrating and splicing the original detection results output by each detection head to obtain the target detection result for the current scene includes:

[0035] The original detection results are concatenated to obtain the output tensor;

[0036] Based on the output tensor, bounding box decoding calculation, confidence calculation, class probability calculation, confidence threshold filtering, class score threshold filtering, and non-maximum elimination processing are performed respectively to obtain the target detection result.

[0037] Optionally, the step of performing bounding box decoding calculation, confidence calculation, class probability calculation, confidence threshold filtering, class score threshold filtering, and non-maximum elimination processing on the output tensor to obtain the target detection result includes:

[0038] For the output tensor, perform bounding box decoding calculation, confidence calculation, and class probability calculation respectively to obtain multiple initial bounding boxes and the confidence and class probability corresponding to each initial bounding box;

[0039] Based on a preset confidence threshold, each initial bounding box is filtered according to its confidence level, and based on a preset class score threshold, each initial bounding box is filtered according to its class probability, resulting in multiple filtered bounding boxes.

[0040] Based on the nonmaximum suppression strategy, redundancy is eliminated from the multiple filtered bounding boxes to obtain the target detection result.

[0041] Secondly, this application provides a target detection device, the device comprising:

[0042] The acquisition module is used to acquire current data for the current scene output by the common aperture device. The current data includes infrared images, visible light images, and radar signals. The common aperture device is used to output the current data that is aligned in time according to the received raw signals.

[0043] The feature extraction module is used to stitch and fuse the infrared image and the visible light image and extract features to obtain an optical feature map, and to extract features from the radar signal to obtain a radar feature map. The number of channel features in the optical feature map is the same as the number of channel features in the radar feature map.

[0044] The fusion module is used to pair the channel features in the optical feature map and the channel features in the radar feature map to obtain multiple feature pairs, and to perform feature alignment and fusion on each feature pair to obtain the fused feature corresponding to each feature pair.

[0045] The detection module is used to distribute the fused features to multiple detection heads, each detection head generates an original detection result based on the fused features, and the original detection results output by each detection head are integrated and spliced ​​to obtain the target detection result of the current scene.

[0046] Thirdly, this application provides an electronic device, including: a processor, a storage medium, and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the target detection method described above.

[0047] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the target detection method described above.

[0048] The beneficial effects of this application are as follows: It acquires current data for the current scene output by a common aperture device, including infrared images, visible light images, and radar signals, eliminating spatiotemporal biases at the source and avoiding post-processing registration errors and computational costs. The infrared and visible light images are stitched together and feature extracted to obtain an optical feature map, and the radar signal is used to extract features to obtain a radar feature map. Features from each channel in the optical and radar feature maps are paired to obtain multiple feature pairs, and these feature pairs are then aligned and fused to obtain the fused features corresponding to each pair. Through channel feature pairing and alignment fusion, a deep integration of optical and radar features is achieved, fully exploring cross-modal correlation information and compensating for the detection deficiencies of single-modal detection. The fused features are distributed to multiple detection heads, each generating original detection results based on its respective fused features. The original detection results output by each detection head are integrated and stitched together to obtain the target detection results for the current scene. Combining parallel detection by multiple detection heads and result integration accurately extracts target information at different scales, improving the comprehensiveness of detection. Overall, it achieves efficient fusion and accurate detection of multimodal data, enhances the robustness of target detection, and can adapt to harsh environments such as nighttime and fog, enabling all-weather target detection. Attached Figure Description

[0049] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a schematic diagram illustrating an application scenario of a target detection method provided in an embodiment of this application;

[0051] Figure 2 This is a schematic flowchart of a target detection method provided in an embodiment of this application;

[0052] Figure 3 This is a schematic diagram of a process for obtaining an optical feature map provided in an embodiment of this application;

[0053] Figure 4 This is a schematic diagram of another process for obtaining an optical feature map provided in an embodiment of this application;

[0054] Figure 5 This is a schematic diagram of a process for obtaining a radar feature map provided in an embodiment of this application;

[0055] Figure 6 This is a schematic diagram of another process for obtaining a radar feature map provided in an embodiment of this application;

[0056] Figure 7 This is a schematic diagram of a process for obtaining a radar feature matrix provided in an embodiment of this application;

[0057] Figure 8 This is a schematic diagram of a process for obtaining the fused features corresponding to each feature pair lock, provided in an embodiment of this application;

[0058] Figure 9 This is a schematic diagram of another process for obtaining the fused features corresponding to each feature pair lock, provided in an embodiment of this application;

[0059] Figure 10 This is a schematic diagram of a process for obtaining target detection results in the current scene, provided by an embodiment of this application.

[0060] Figure 11 This is a schematic diagram of the structure of a target detection device provided in an embodiment of this application;

[0061] Figure 12 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the accompanying drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.

[0063] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0064] It should be noted that the term "comprising" will be used in the embodiments of this application to indicate the presence of the features declared thereafter, but does not exclude the addition of other features.

[0065] Existing target detection methods suffer from problems such as line-of-sight deviation and temporal jitter, loss of massive intermediate semantic clues, modal interference, and feature collapse. To address these issues, this application proposes a target detection method. This method acquires temporally aligned current data from a common-aperture device, including infrared images, visible light images, and radar signals, thereby eliminating spatiotemporal deviations at the source and avoiding line-of-sight deviation and temporal jitter. Then, the infrared image and visible light image are stitched together and their features extracted. Next, features are extracted from the radar feature map. The extracted feature maps are then paired, and the paired feature pairs are aligned, fused, and detected by the detection head to obtain the target detection result. This method accurately extracts target information at different scales, improving the comprehensiveness of detection.

[0066] Next, we will introduce the application scenarios of target detection methods. Figure 1 This is a schematic diagram illustrating an application scenario of a target detection method provided in an embodiment of this application. For example... Figure 1 As shown, the target detection method is applied to electronic devices with computing capabilities, such as controllers. The electronic device is connected to a common aperture device.

[0067] Among them, such as Figure 1 As shown, the common aperture device consists of an optical sensor, a primary mirror, a dichroic mirror, and a feed source. The primary mirror integrates the visible light and infrared optical detection channels with the radar signal detection channel within the same optical aperture. The original light signal and radar reflection signal from the external scene are received by the primary mirror and then processed by the dichroic mirror to separate and collect the visible light and infrared signals. The feed source receives and converts the radar signal, ensuring complete synchronization of the acquisition time of each channel's signal.

[0068] Figure 2 This is a schematic flowchart of a target detection method provided in an embodiment of this application. Figure 2 As shown, the target detection method will be introduced next.

[0069] S201. Obtain current data for the current scene output by the common aperture device. The current data includes infrared images, visible light images, and radar signals. The common aperture device is used to output time-aligned current data based on the received raw signals.

[0070] Optionally, the common aperture device can directly output infrared images, visible light images, and radar signals at the same time and in the same field of view, i.e., time-aligned current data.

[0071] S202. The infrared image and the visible light image are stitched together and their features are extracted to obtain an optical feature map. The radar signal is then subjected to feature extraction to obtain a radar feature map. The number of channel features in the optical feature map is the same as the number of channel features in the radar feature map.

[0072] Specifically, for infrared and visible light images: First, the infrared and visible light images are preprocessed with alignment and registration, size standardization, data normalization, and data augmentation to eliminate local biases in image acquisition, unify data formats, and improve the model's generalization ability. Then, the preprocessed dual-modal images are fused by channel stitching to form dual-modal optical image data containing multi-channel information. Finally, this data is input into a deep convolutional backbone network, which is enhanced by downsampling, upsampling, lateral connections, and attention mechanisms to extract optical feature maps containing multi-scale and multi-semantic information.

[0073] For radar signals, preprocessing is first performed, including length standardization, filtering and denoising, and data normalization, to obtain standardized radar data with high signal-to-noise ratio and stable distribution. Then, features are extracted in parallel from multiple dimensions, such as time domain, frequency domain, time-frequency domain, and statistical dimensions. The radar feature matrix is ​​generated through concatenation and encoding using a Multi-Layer Perceptron (MLP) network. Finally, the feature matrix undergoes multi-scale projection transformation and is calibrated using a channel attention mechanism to generate a radar feature map.

[0074] Optionally, the optical feature map and the radar feature map have the same number of channel dimensions to ensure that cross-modal features can be paired, aligned and fused, eliminating the difference in feature dimensions between modes.

[0075] S203. Pair the features of each channel in the optical feature map and the features of each channel in the radar feature map to obtain multiple feature pairs, and perform feature alignment and fusion on each feature pair to obtain the fused features corresponding to each feature pair.

[0076] Optionally, the channel features that match the dimension, scale, and semantic information in the optical feature map and the radar feature map are combined one by one to form feature pairs. Each feature pair contains optical modal features and radar modal features at the same scale.

[0077] Optionally, based on the premise of consistent channel feature quantity, and following the principles of scale matching and semantic matching, channel features of the same scale and semantics in the optical feature map and radar feature map are paired one-to-one to obtain multiple feature pairs. For example, high-resolution optical channel features are paired with high-resolution radar channel features, and deep semantic optical channel features are paired with deep semantic radar channel features.

[0078] Specifically, for each feature pair, feature alignment is first performed through operations such as convolution to eliminate subtle deviations in spatial resolution and feature dimension across modal features. Then, based on the scale and semantic characteristics of the feature pair, corresponding fusion strategies are adopted for feature fusion. These fusion strategies include cross-modal attention, gated fusion, element-wise addition, and interactive attention. By exploring the deep correlation between optical and radar features, it is possible to enhance and correct features from one modality for features from another.

[0079] S204. Distribute each fusion feature to multiple detection heads, and each detection head generates an original detection result based on each fusion feature. Integrate and splice the original detection results output by each detection head to obtain the target detection result of the current scene.

[0080] The detection head is a network module in the model used to detect objects from the fused features. It can extract detection information such as the bounding box, category, and confidence score of the target from the fused features. Different detection heads are adapted to fused features of different scales.

[0081] Specifically, fused features of different scales and semantics are distributed to multiple corresponding parallel detection heads. Each detection head independently extracts target information from the received fused features through convolutional layers, generating raw detection results containing bounding box positions, target confidence, and class scores. Then, the raw detection results output by all detection heads are integrated and concatenated along the tensor dimension to form a raw output tensor containing candidate target information at all scales and locations. Finally, this tensor undergoes post-processing such as bounding box decoding, confidence threshold filtering, class score threshold filtering, and non-maximum suppression to remove redundant and erroneous detection information, select the target detection information with the highest confidence, and ultimately obtain the target detection results for the current scene.

[0082] In this embodiment, current data for the current scene output by the common aperture device is acquired. This current data includes infrared images, visible light images, and radar signals, eliminating spatiotemporal biases at the source and avoiding post-processing registration errors and computational costs. The infrared and visible light images are stitched together and feature extracted to obtain an optical feature map. Features are extracted from the radar signal to obtain a radar feature map. Features from each channel in both the optical and radar feature maps are paired to obtain multiple feature pairs. These feature pairs are then aligned and fused to obtain the fused features corresponding to each pair. Through channel feature pairing and alignment fusion, a deep integration of optical and radar features is achieved, fully exploring cross-modal correlation information and compensating for the detection deficiencies of single-modal detection. The fused features are distributed to multiple detection heads, each generating original detection results based on its respective fused features. The original detection results output by each detection head are integrated and stitched together to obtain the target detection results for the current scene. Combining parallel detection by multiple detection heads and result integration accurately extracts target information at different scales, improving the comprehensiveness of detection. Overall, it achieves efficient fusion and accurate detection of multimodal data, enhances the robustness of target detection, and can adapt to harsh environments such as nighttime and fog, enabling all-weather target detection.

[0083] The specific methods for obtaining optical feature maps and radar feature maps will be introduced next.

[0084] Figure 3 This is a schematic diagram of a process for obtaining an optical feature map provided in an embodiment of this application. Figure 4 This is another schematic diagram of the process for obtaining an optical feature map provided in an embodiment of this application, such as... Figure 3 and Figure 4 As shown:

[0085] S301. Alignment and registration processing, size standardization processing, and data normalization processing are performed on the infrared image and the visible light image respectively to obtain the processed infrared image and the processed visible light image.

[0086] Specifically, key feature points, such as edges and corners, are first extracted from the infrared and visible light images. The spatial transformation matrix of the feature points is then calculated. Finally, the infrared and visible light images are mapped to the coordinate system of the other image through affine transformation to complete spatial alignment.

[0087] Based on the input requirements of the deep convolutional backbone network, the two aligned images are scaled to the same fixed size, such as 1024×1024, to ensure the consistency of the convolutional kernel's sliding stride and receptive field.

[0088] The pixel values ​​of the two images are then normalized to stabilize the data distribution and prevent gradient explosion or slow convergence during model training caused by the high dynamic range of infrared images. The final outputs are the processed infrared image and the processed visible light image.

[0089] S302. Perform channel stitching and fusion and data enhancement on the processed infrared image and the processed visible light image to obtain dual-modal optical image data.

[0090] The processed infrared and visible light images are stitched together along the channel dimension to generate a tensor of shape [H, W, 4], where H is the height and W is the width. Random data augmentation is then applied to the stitched tensor, such as random horizontal flipping, ±10° rotation, and scaling by 0.8-1.2 times, as well as random brightness adjustments to the infrared channels to simulate the thermal radiation intensity of different environments. This avoids model overfitting and adapts to changes in lighting and viewing angles in real-world scenes.

[0091] S303. Input the dual-modal optical image data into the deep convolutional backbone network for layer-by-layer downsampling to obtain multiple optical layer features.

[0092] First, the bimodal optical image data is adjusted to the network input format and input into a preset deep convolutional backbone network for layer-by-layer downsampling. This embodiment introduces four layers of optical layer features, but does not limit the number of layers and features in actual implementation: the output of layer 1 is at a resolution of H / 2 × W / 2 with 64 channels, representing shallow detail features; the output of layer 2 is at a resolution of H / 4 × W / 4 with 128 channels, representing mid-level texture features; the output of layer 3 is at a resolution of H / 8 × W / 8 with 256 channels, representing mid-to-deep semantic features; and the output of layer 4 is at a resolution of H / 16 × W / 16 with 512 channels, representing deep abstract features.

[0093] S304. Upsample and laterally connect the features at each optical level to obtain a multi-scale feature pyramid. Based on the attention mechanism, enhance the features of the multi-scale feature pyramid to obtain an optical feature map.

[0094] Specifically, the deepest optical layer features are upsampled by a factor of 2 and then horizontally connected to the features of the next higher level. For example, an H / 16 × W / 16 optical layer feature is upsampled by a factor of 2 and then horizontally connected to an H / 8 × W / 8 optical layer feature. This process is repeated to sequentially upsample deep features and fuse them with shallow features, ultimately generating four feature maps of different resolutions, forming a multi-scale feature pyramid. It is worth noting that this embodiment uses four optical layer features as an example; the number of optical layer features can be adjusted as needed in actual implementation.

[0095] For each scale feature map in the feature pyramid, channel attention and spatial attention are applied separately. For example, channel weights are calculated using the SE module, and then a spatial weight map is generated through convolution to weight and enhance the feature map. Then, the number of channels in the enhanced feature map is adjusted through convolution to make it completely consistent with the number of channels in the subsequent radar feature map, and finally, an optical feature map is output.

[0096] exist Figure 4 The example demonstrates optical feature extraction at four scales: P3 level feature extraction, P4 level feature extraction, P5 level feature extraction, and deep feature extraction. The output dimensions for P3 level feature extraction are 256×128×128, P4 level feature extraction is 512×64×64, P5 level feature extraction is 512×32×32, and deep feature extraction is 1024×32×32.

[0097] In this embodiment, a multi-scale feature pyramid is used to adapt to the detection of targets of different sizes, an attention mechanism is used to improve the representation ability of target features, and dimensional uniformity ensures the feasibility of cross-modal fusion.

[0098] Figure 5 This is a schematic diagram of a process for obtaining a radar feature map provided in an embodiment of this application. Figure 6 This is another schematic diagram of the process for obtaining a radar feature map provided in an embodiment of this application, such as... Figure 5 and Figure 6 As shown:

[0099] S501. Preprocess the radar signal to obtain the processed radar signal.

[0100] Specifically, preprocessing can include filtering, baseline calibration, and data normalization. Filtering employs adaptive filtering or wavelet denoising algorithms to remove background clutter and high-frequency electronic noise while retaining the effective components of the target echo. Baseline calibration involves eliminating baseline drift through polynomial fitting to ensure a uniform signal amplitude reference. Length normalization involves truncating or padding radar signals of different durations to a preset fixed length, such as 1024 sampling points, to ensure consistent input dimensions for subsequent feature extraction. Data normalization maps signal amplitudes to the 0-1 range, reducing the impact of signal strength fluctuations on feature extraction.

[0101] S502. Perform multi-dimensional feature extraction on the processed radar signal to obtain the radar feature matrix.

[0102] The radar signal features are extracted in parallel from four dimensions: time domain, frequency domain, time-frequency domain, and statistical dimension. The feature vectors extracted from these four dimensions are then concatenated to form a one-dimensional long feature vector. This long feature vector is then input into a multilayer perceptron for nonlinear encoding, mapping the feature vector into a fixed-dimensional two-dimensional matrix, ultimately yielding the radar feature matrix.

[0103] S503. Perform multi-scale projection on the radar feature matrix to obtain features at multiple scales.

[0104] Optionally, when selecting the scale, the scale on which the radar feature needs to be projected can be determined based on the scale of the optical feature pyramid. For example, if the scale of the optical feature pyramid is 4, then the scale feature of the radar is also 4.

[0105] Specifically, the radar feature matrix is ​​first adjusted for the number of channels using a 1×1 convolution, and then downsampled using convolutional or pooling layers of different lengths to generate multiple feature maps with different resolutions. Next, the deep, low-resolution features are upsampled to calibrate the spatial dimensions of the feature maps, ensuring that the resolution of each radar feature at each scale matches that of the corresponding optical feature. This results in radar feature maps at multiple scales, i.e., multiple scale features, where the spatial dimensions of each scale feature match the corresponding scale of the optical feature map.

[0106] S504. Based on the channel attention mechanism, attention enhancement is performed on features at each scale to obtain radar feature maps.

[0107] Optionally, for each scale feature map, the SE module is connected. The SE module compresses the features of each channel into a scalar through global average pooling, learns the weight coefficients of each channel through two fully connected layers, assigns high weights to important channels or low weights to redundant channels, and then multiplies the learned weight coefficients with the original scale feature map channel by channel to complete feature enhancement.

[0108] Optionally, the number of channels in the enhanced feature maps at each scale can be adjusted by 1×1 convolution to make them exactly the same as the number of channel features in the optical feature map;

[0109] Optionally, enhanced feature maps at multiple scales can be integrated into a single feature set to ultimately output a radar feature map.

[0110] For example, the formula for obtaining the radar feature map is as follows: Formula (1):

[0111] (1)

[0112] in, and These are the weights of the convolutional layer. It is the scale feature of the i-th scale. This indicates channel multiplication.

[0113] For example, in Figure 6 The example shows four scale projections: scale 1 projection, scale 2 projection, scale 3 projection, and scale 4 projection. The scale 1 projection corresponds to an output dimension of 256×128×128, the scale 2 projection corresponds to an output dimension of 512×64×64, the scale 3 projection corresponds to an output dimension of 512×32×32, and the scale 4 projection corresponds to an output dimension of 1024×32×32.

[0114] In this embodiment, multi-dimensional feature extraction is used to mine all-round information of radar signals, and finally outputs a radar feature map that meets the requirements of cross-modal pairing and fusion.

[0115] Next, refer to Figure 7 The specific method for extracting multi-dimensional features from the processed radar signal to obtain the radar feature matrix in step S502 above is described. Specifically, Figure 7 This is a schematic diagram of a process for obtaining a radar feature matrix provided in an embodiment of this application.

[0116] S701. Perform time-domain feature extraction, frequency-domain feature extraction, time-frequency-domain feature extraction, and statistical feature extraction on the processed radar signal to obtain time-domain features, frequency-domain features, time-frequency-domain features, and statistical features.

[0117] Among these, time-domain feature extraction involves mining features from the time-domain sequence of the processed radar signal to extract relevant features that reflect the change of signal amplitude over time, specifically including extracting the signal's mean, variance, and zero-crossing rate; frequency-domain feature extraction involves first converting the processed radar signal from the time domain to the frequency domain, and then extracting relevant frequency-domain features that reflect the target's motion and structural characteristics, specifically including performing a Fourier transform on the signal and extracting its spectral centroid; time-frequency domain feature extraction generates a time-frequency map of the processed radar signal through time-frequency transformation, and then extracts relevant features that reflect the dynamic change of signal frequency over time, specifically including using wavelet transform to extract energy distribution features; and statistical feature extraction targets the overall data distribution characteristics of the processed radar signal, extracting relevant statistical features that reflect the global distribution pattern of the signal, specifically including extracting kurtosis features.

[0118] Specifically, the time-domain features, frequency-domain features, time-frequency-domain features, and statistical features obtained after the above-mentioned feature extraction operations are all presented in the form of one-dimensional feature vectors. Each type of feature vector independently represents the core characteristics of the radar signal in the corresponding dimension.

[0119] S702. Perform feature concatenation on time-domain features, frequency-domain features, time-frequency-domain features, and statistical features to obtain concatenated features.

[0120] Specifically, the four independent one-dimensional feature vectors are sequentially connected according to the preset dimensional direction to complete the feature splicing and fusion process, and finally a one-dimensional long feature vector containing full-dimensional information of the radar signal is obtained, that is, the spliced ​​feature.

[0121] This feature concatenation operation can fully preserve the target information covered by time domain, frequency domain, time-frequency domain, and statistical features, avoiding the information loss problem caused by single-dimensional features.

[0122] S703. Perform feature encoding on the spliced ​​features to generate a radar feature matrix.

[0123] Specifically, the concatenated features are input into a multilayer perceptron network. The fully connected layers of the multilayer perceptron perform nonlinear mapping and dimensional transformation on the one-dimensional concatenated features, transforming the one-dimensional feature vector into feature data in the form of a two-dimensional matrix.

[0124] In the feature encoding process, the multilayer perceptron can autonomously learn the deep relationships between features, enhance the feature representation capability for effective target detection, filter redundant feature information, and adjust the dimensional specifications of the two-dimensional matrix to the standard specifications required for subsequent multi-scale projection operations.

[0125] Optionally, during the feature encoding process, the number of network layers and the number of neurons in the fully connected layers of the multilayer perceptron can be adjusted according to the needs of the actual application scenario to achieve better feature encoding results.

[0126] In this embodiment, key features of radar signals are extracted from four perspectives, and the four types of information are then integrated together, which not only retains the unique information of each type of feature, but also forms a unified feature that can comprehensively characterize radar signals.

[0127] Optionally, the channel features in the optical feature map include a first optical feature, a second optical feature, a third optical feature, and a fourth optical feature, with the resolutions of the first, second, and third optical features decreasing sequentially, and the fourth optical feature being a deep semantic feature. The channel features in the radar feature map include a first radar feature, a second radar feature, a third radar feature, and a fourth radar feature, with the feature size of the first optical feature being the same as that of the first radar feature, the feature size of the second optical feature being the same as that of the second radar feature, the feature size of the third optical feature being the same as that of the third radar feature, and the feature size of the fourth optical feature being the same as that of the fourth radar feature.

[0128] Optionally, the first optical feature and the first radar feature are used as the first feature pair, the second optical feature and the second radar feature are used as the second feature pair, the third optical feature and the third radar feature are used as the third feature pair, and the fourth optical feature and the fourth radar feature are used as the fourth feature pair.

[0129] Based on the size matching relationship between the optical features and radar features mentioned above, optionally, this embodiment performs feature pairing operations according to the pairing principle of the same size and the same level. Specifically, the first optical feature is combined with the first radar feature to form a first feature pair, the second optical feature is combined with the second radar feature to form a second feature pair, the third optical feature is combined with the third radar feature to form a third feature pair, and the fourth optical feature is combined with the fourth radar feature to form a fourth feature pair. Through this pairing method, each feature pair contains optical modal features and radar modal features at the same scale or the same semantic level, laying the foundation for subsequent adoption of differentiated alignment and fusion strategies for different feature pairs, and ensuring that cross-modal fusion operations can be accurately applied to the corresponding feature level.

[0130] Figure 8 This is a schematic diagram of a process for obtaining the fused features corresponding to each feature pair lock, provided in an embodiment of this application. Figure 9 This is a schematic diagram illustrating another process for obtaining the fused features corresponding to each feature pair lock, provided in an embodiment of this application. For example... Figure 8 and Figure 9 As shown, the specific method for obtaining the fusion features will be introduced next.

[0131] S801. Perform feature alignment and cross-modal attention enhancement processing on the first optical feature and the first radar feature in the first feature pair respectively to obtain the first enhanced optical feature and the first enhanced radar feature. Perform feature stitching and fusion on the first enhanced optical feature and the first enhanced radar feature to obtain the first stitched feature. Perform weight calibration on the first stitched feature based on the attention mechanism to obtain the first fused feature.

[0132] First, feature alignment processing is performed on the first optical feature and the first radar feature to eliminate minor deviations in spatial resolution and channel dimension of cross-modal features, ensuring that the feature dimensions of the two are completely matched.

[0133] Based on this, cross-modal attention enhancement processing is performed. The cross-modal attention mechanism is used to mine the correlation information between the first optical feature and the first radar feature, enhance the feature and suppress background redundancy features, and obtain the first enhanced optical feature and the first enhanced radar feature.

[0134] Specifically, the first enhanced optical feature and the first enhanced radar feature are concatenated and fused in the channel dimension through a 3×3 convolution, integrating all the information of the dual-modal enhanced features to generate the first concatenated feature. Then, based on the attention mechanism, the channel and spatial dimensions of the first concatenated feature are weighted and calibrated to give the effective feature a higher weight coefficient, and finally the first fused feature with dual-modal advantages and enhanced representation ability is obtained.

[0135] S802: Perform feature alignment processing on the second optical feature and the second radar feature in the second feature pair respectively, and generate a single-channel weight map based on the obtained second processed optical feature and second processed radar feature. Based on the gated fusion mechanism, determine the second stitching feature according to the single-channel weight map, the second processed optical feature, and the second processed radar feature. Determine the second fused feature according to the second stitching feature and the second processed optical feature.

[0136] First, the second optical feature and the second radar feature are aligned by 1×1 convolution to obtain the second processed optical feature and the second processed radar feature with dimension matching. Based on the above two types of processed features, a single-channel weight map is generated by convolution operation and normalization operation. This weight map can represent the importance ratio of optical features and radar features under different spatial locations.

[0137] Based on the gating fusion mechanism, the single-channel weight map is applied to the second processed optical feature and the second processed radar feature respectively. The two types of features are dynamically fused according to the weight ratio to obtain the second stitched feature. Then, the second stitched feature is residually connected with the second processed optical feature, retaining the basic information of the optical feature while incorporating the complementary information of the radar feature, and finally determining the second fused feature.

[0138] S803. The third optical feature and the third radar feature in the third feature pair are respectively subjected to feature alignment processing, and the resulting processed third optical feature and processed third radar feature are then summed to obtain the third stitched feature. Based on a gated fusion mechanism, a gated fusion feature is determined according to the third stitched feature and the processed third optical feature. The gated fusion feature and the third stitched feature are fused to obtain the adjusted feature. Based on a nonlinear activation function, the third fusion feature is determined according to the adjusted feature.

[0139] First, the third optical feature and the third radar feature are subjected to 1×1 convolution alignment processing to obtain the third processed optical feature and the third processed radar feature.

[0140] The two types of processed features are added element by element to quickly integrate the basic information of the dual-modal features and generate a third concatenated feature.

[0141] Based on the gated fusion mechanism, the third stitching feature is used as input and the third processed optical feature is used as reference. The gated weight is dynamically generated and applied to the third stitching feature to obtain the gated fusion feature. The gated fusion feature is then fused with the third stitching feature channel by channel to further calibrate the feature information and obtain the adjusted feature. Finally, the adjusted feature is input into a nonlinear activation function to introduce a nonlinear transformation to enhance the complex representation ability of the feature and ultimately determine the third fusion feature.

[0142] S804. The fourth optical feature and the fourth radar feature in the fourth feature pair are respectively subjected to feature alignment processing. Based on the interactive attention mechanism, the obtained fourth processed optical feature and fourth processed radar feature are enhanced to obtain the fourth enhanced optical feature and the fourth enhanced radar feature. The fourth enhanced optical feature and the fourth enhanced radar feature are then concatenated and compressed to obtain the fourth concatenated feature. Based on the global context fusion mechanism, the long-range dependencies of the fourth concatenated feature are captured to obtain the fourth fused feature.

[0143] Specifically, the fourth optical feature and the fourth radar feature are first processed by 1×1 convolution to perform feature alignment and channel matching, resulting in the fourth processed optical feature and the fourth processed radar feature.

[0144] Based on the interactive attention mechanism, the two types of features guide each other in the generation of attention weights. That is, the optical features guide the attention focus of the radar features, and the radar features guide the attention focus of the optical features. After the enhancement process is completed, the fourth enhanced optical feature and the fourth enhanced radar feature are obtained.

[0145] After concatenating the two types of enhanced features, dimensionality is compressed using a 1×1 convolution to remove redundant information and integrate core features, generating a fourth concatenated feature. Then, based on a global context fusion mechanism, long-distance dependencies between different spatial locations in the fourth concatenated feature are captured, and global associations between deep semantic features are mined, ultimately obtaining a fourth fused feature that can represent the high-level semantic attributes of the target.

[0146] In this embodiment, four sets of fused features are finally output, which are adapted to the detection requirements of small targets, medium targets, large targets and deep semantics, respectively.

[0147] Next, refer to Figure 10 This section describes the specific process in step S204 above of integrating and splicing the original detection results output by each detection head to obtain the target detection results for the current scene.

[0148] Optionally, the original detection results can be spliced ​​together to obtain the output tensor.

[0149] Each original detection result is unprocessed detection data generated by multiple parallel detection heads from the fusion features of the corresponding scale. It includes information such as the initial bounding box coordinates, initial confidence, and initial class score of the target. The dimensions and format of the original detection results output by different detection heads are kept consistent.

[0150] Specifically, the raw detection results output by all detection heads are concatenated and fused according to the tensor's dimensional direction, integrating them into a multidimensional tensor containing detection information for all scales and all candidate targets—the output tensor. This output tensor completely preserves all the raw detection information output by the multiple detection heads, ensuring that subsequent post-processing can cover the screening and optimization of all candidate targets. The tensor's dimensions can include, for example, batch dimension or spatial dimension.

[0151] Optionally, based on the output tensor, bounding box decoding calculation, confidence calculation, class probability calculation, confidence threshold filtering, class score threshold filtering, and non-maximum elimination processing are performed respectively to obtain the target detection result.

[0152] First, based on the initial parameters in the output tensor, the following steps are performed: bounding box decoding calculation converts the bounding box offsets output by the model into bounding box coordinates in the actual pixel coordinate system, restoring the true position of the target in the image; confidence calculation is based on the confidence branch data in the output tensor to calculate the probability value of each candidate bounding box containing the target; class probability calculation is based on the class branch data in the output tensor to calculate the probability distribution of each candidate bounding box for each class.

[0153] Specifically, after completing the above calculations, filtering and elimination operations are performed sequentially: confidence threshold filtering sets a preset confidence threshold, removes candidate bounding boxes with confidence levels below the threshold, and retains target candidates with high confidence; category score threshold filtering sets a preset category score threshold, removes candidate bounding boxes with category probabilities below the threshold, and further filters effective targets; non-maximum elimination processing traverses the remaining candidate bounding boxes by category, removes redundant bounding boxes with overlap levels above a preset threshold, and retains only the bounding box with the highest confidence level corresponding to the same target.

[0154] Optionally, non-maximum elimination processing can employ optimization algorithms such as Soft-NMS to avoid removing valid targets due to excessive bounding box overlap. Finally, through the above series of processing, accurate target detection results containing target bounding box coordinates, target category, and target confidence are obtained.

[0155] In this embodiment, the accuracy of the target detection results is ensured by filtering out accurate and non-redundant target detection results from the original detection information through the processing of the original detection results.

[0156] Next, we will introduce the process of processing the output tensor in the above steps.

[0157] Optionally, based on a preset confidence threshold, each initial bounding box is filtered according to its confidence level, and based on a preset class score threshold, each initial bounding box is filtered according to its class probability, resulting in multiple filtered bounding boxes.

[0158] The preset confidence threshold and the preset category score threshold are both numerical thresholds pre-set according to the actual target detection scenario and detection requirements, and can be flexibly adjusted according to the requirements of detection accuracy and recall.

[0159] The confidence level of the initial bounding box represents the probability that the bounding box contains a target, and the class probability of the initial bounding box represents the probability that the target in the bounding box belongs to a certain preset class.

[0160] Specifically, a confidence filtering operation is first performed, comparing the confidence of all initial bounding boxes with a preset confidence threshold one by one, removing initial bounding boxes with a confidence lower than the threshold, and retaining initial bounding boxes with a confidence greater than or equal to the threshold. This operation can effectively filter out invalid bounding boxes that only contain background areas and have no actual detection targets.

[0161] Building upon the confidence-based filtering, a further category probability filtering operation is performed. Specifically, the category probabilities of each category in the initial bounding boxes retained after confidence-based filtering are compared one by one with a preset category score threshold. Initial bounding boxes whose probabilities for all categories are below the threshold are removed, retaining only those initial bounding boxes whose probability for any category is greater than or equal to the threshold. This operation can further filter out candidate bounding boxes with ambiguous target classifications or low classification confidence.

[0162] After the above two-layer threshold filtering operation, the final number of retained candidate bounding boxes are the multiple filtered bounding boxes.

[0163] Optionally, based on a nonmaximum suppression strategy, redundancy is removed from multiple filtered bounding boxes to obtain the target detection result.

[0164] Among them, the nonmaximum suppression strategy is a classic algorithm strategy in the field of object detection used to eliminate duplicate detection boxes. Specifically, it sorts candidate bounding boxes in the same category according to their confidence level and removes redundant bounding boxes with nonmaximum confidence levels and overlap with high-confidence bounding boxes exceeding a preset overlap threshold.

[0165] Specifically, all filtered bounding boxes are first classified according to the target category, so that bounding boxes of the same category are grouped together. Then, the filtered bounding boxes in each group are sorted in descending order of confidence. Starting from the filtered bounding box with the highest confidence in each group, the intersection-union ratio (IUR) of the bounding box with the other bounding boxes in the group is calculated in turn. The other bounding boxes with IUR exceeding the preset overlap threshold are identified as redundant bounding boxes and are removed. Only the bounding box with the highest confidence is retained.

[0166] Following the above logic, the filtered bounding boxes in each group are subjected to non-maximum suppression redundancy removal. Finally, all the remaining bounding boxes, along with their corresponding target categories and confidence information, constitute the target detection result. In this result, each actual detected target corresponds to a unique bounding box, and there are no issues of duplicate detection or redundant annotation.

[0167] In this embodiment, the number of candidate boxes in subsequent processing is greatly reduced by filtering and redundancy removal steps, thereby improving detection efficiency.

[0168] As an optional implementation, the model training phase of the object detection method in this embodiment can be based on a graphics processing unit (GPU) as an accelerated computing platform, and training can be performed on a pre-defined category object detection dataset. This platform includes third-party libraries such as Python, PyTorch, and OpenCV. Furthermore, the model inference phase can flexibly utilize either a central processing unit (CPU) or a GPU for inference.

[0169] Based on the same inventive concept, this application also provides a target detection device corresponding to the target detection method. Since the principle of the device in this application is similar to that of the target detection method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0170] Reference Figure 11 The diagram shown is a schematic representation of a target detection device according to an embodiment of this application. The device includes: an acquisition module 1101, a feature extraction module 1102, a fusion module 1103, and a detection module 1104; wherein:

[0171] The acquisition module 1101 is used to acquire current data for the current scene output by the common aperture device. The current data includes infrared images, visible light images and radar signals. The common aperture device is used to output the current data that is aligned in time according to the received raw signals.

[0172] The feature extraction module 1102 is used to stitch and fuse the infrared image and the visible light image and extract features to obtain an optical feature map, and to extract features from the radar signal to obtain a radar feature map. The number of channel features in the optical feature map is the same as the number of channel features in the radar feature map.

[0173] The fusion module 1103 is used to pair the channel features in the optical feature map and the channel features in the radar feature map to obtain multiple feature pairs, and to perform feature alignment and fusion on each feature pair to obtain the fused feature corresponding to each feature pair.

[0174] The detection module 1104 is used to distribute the fused features to multiple detection heads, each detection head generates an original detection result based on the fused features, and integrates and splices the original detection results output by each detection head to obtain the target detection result of the current scene.

[0175] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0176] This application also provides an electronic device, such as... Figure 12 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this application, including a processor 1201, a memory 1202, and a bus. The memory 1202 stores machine-readable instructions executable by the processor 1201. When the computer device is running, the processor 1201 and the memory 1202 communicate via the bus, and the processor 1201 executes the machine-readable instructions to perform the processing of the target detection method described above.

[0177] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the target detection method described above.

[0178] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some communication interfaces; the indirect coupling or communication connection of devices or modules can be electrical, mechanical, or other forms.

[0179] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0180] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A target detection method, characterized in that, The method includes: The current data for the current scene is obtained from the output of the common aperture device. The current data includes infrared images, visible light images, and radar signals. The common aperture device is used to output the current data that is aligned in time according to the received raw signals. The infrared image and the visible light image are stitched together and their features are extracted to obtain an optical feature map. The radar signal is then subjected to feature extraction to obtain a radar feature map. The number of channel features in the optical feature map is the same as the number of channel features in the radar feature map. The channel features in the optical feature map include a first optical feature, a second optical feature, a third optical feature, and a fourth optical feature. The channel features in the radar feature map include a first radar feature, a second radar feature, a third radar feature, and a fourth radar feature. The first optical feature and the first radar feature are used as a first feature pair, the second optical feature and the second radar feature are used as a second feature pair, the third optical feature and the third radar feature are used as a third feature pair, and the fourth optical feature and the fourth radar feature are used as a fourth feature pair. Feature alignment is performed on the optical and radar features in each feature pair. Cross-modal attention enhancement processing is applied to the optical and radar features in the first feature pair. The enhanced optical and radar features are then spliced ​​and fused to obtain a first spliced ​​feature. Weight calibration is performed on the first spliced ​​feature based on an attention mechanism to obtain a first fused feature. A single-channel weight map is generated based on the optical and radar features in the second feature pair. A second spliced ​​feature is determined based on a gated fusion mechanism, using the single-channel weight map, the optical features in the second feature pair, and the radar features. Based on the second stitching feature and the optical features in the second feature pair, a second fusion feature is determined; the optical features and radar features in the third feature pair are subjected to feature addition processing to obtain a third stitching feature; based on a gated fusion mechanism, a gated fusion feature is determined based on the third stitching feature and the optical features in the third feature pair; the gated fusion feature and the third stitching feature are fused to obtain an adjusted feature; based on a nonlinear activation function, a third fusion feature is determined based on the adjusted feature; based on an interactive attention mechanism, the optical features and radar features in the fourth feature pair are enhanced to obtain a fourth enhanced optical feature and a fourth enhanced radar feature; the fourth enhanced optical feature and the fourth enhanced radar feature are stitched and compressed to obtain a fourth stitching feature; based on a global context fusion mechanism, the long-range dependencies of the fourth stitching feature are captured to obtain a fourth fusion feature; Each fusion feature is distributed to multiple detection heads, and each detection head generates an original detection result based on the fusion feature. The original detection results output by each detection head are integrated and spliced ​​together to obtain the target detection result of the current scene.

2. The target detection method according to claim 1, characterized in that, The process of stitching and fusing the infrared image and the visible light image, and extracting features to obtain an optical feature map, includes: The infrared image and the visible light image are respectively subjected to alignment and registration processing, size standardization processing and data normalization processing to obtain the processed infrared image and the processed visible light image. The processed infrared image and the processed visible light image are subjected to channel stitching and fusion as well as data enhancement to obtain dual-modal optical image data. The dual-modal optical image data is input into a deep convolutional backbone network for layer-by-layer downsampling to obtain multiple optical layer features; Upsampling and lateral connection are performed on the optical level features to obtain a multi-scale feature pyramid. Based on the attention mechanism, feature enhancement is performed on the multi-scale feature pyramid to obtain an optical feature map.

3. The target detection method according to claim 1, characterized in that, Feature extraction is performed on the radar signal to obtain a radar feature map, including: The radar signal is preprocessed to obtain the processed radar signal; Multi-dimensional feature extraction is performed on the processed radar signal to obtain a radar feature matrix; The radar feature matrix is ​​projected at multiple scales to obtain features at multiple scales. Based on the channel attention mechanism, attention enhancement is performed on each scale feature to obtain the radar feature map.

4. The target detection method according to claim 3, characterized in that, The step of extracting multi-dimensional features from the processed radar signal to obtain a radar feature matrix includes: The processed radar signal is subjected to time-domain feature extraction, frequency-domain feature extraction, time-frequency-domain feature extraction, and statistical feature extraction to obtain time-domain features, frequency-domain features, time-frequency-domain features, and statistical features. The time-domain features, frequency-domain features, time-frequency-domain features, and statistical features are concatenated to obtain the concatenated features; The spliced ​​features are then encoded to generate a radar feature matrix.

5. The target detection method according to claim 1, characterized in that, The resolutions of the first optical feature, the second optical feature, and the third optical feature decrease sequentially, and the fourth optical feature is a deep semantic feature; the feature size of the first optical feature is the same as the feature size of the first radar feature, the feature size of the second optical feature is the same as the feature size of the second radar feature, the feature size of the third optical feature is the same as the feature size of the third radar feature, and the feature size of the fourth optical feature is the same as the feature size of the fourth radar feature.

6. The target detection method according to claim 1, characterized in that, The process of integrating and splicing the raw detection results output by each detection head to obtain the target detection result of the current scene includes: The original detection results are concatenated to obtain the output tensor; Based on the output tensor, bounding box decoding calculation, confidence calculation, class probability calculation, confidence threshold filtering, class score threshold filtering, and non-maximum elimination processing are performed respectively to obtain the target detection result.

7. The target detection method according to claim 6, characterized in that, The process involves performing bounding box decoding calculation, confidence score calculation, class probability calculation, confidence score threshold filtering, class score threshold filtering, and non-maximum elimination processing on the output tensor to obtain the target detection result, including: For the output tensor, perform bounding box decoding calculation, confidence calculation, and class probability calculation respectively to obtain multiple initial bounding boxes and the confidence and class probability corresponding to each initial bounding box; Based on a preset confidence threshold, each initial bounding box is filtered according to its confidence level, and based on a preset class score threshold, each initial bounding box is filtered according to its class probability, resulting in multiple filtered bounding boxes. Based on the nonmaximum suppression strategy, redundancy is eliminated from the multiple filtered bounding boxes to obtain the target detection result.

8. An electronic device, characterized in that, include: The processor and memory, the memory storing machine-readable instructions executable by the processor, wherein when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the target detection method as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the target detection method as described in any one of claims 1 to 7.