Unmanned aerial vehicle multi-source image target detection method and system based on multi-level fusion

CN118052975BActive Publication Date: 2026-09-22BEIJING JIAOTONG UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410069717.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2026-09-22
Estimated Expiration
2044-01-18

AI Technical Summary

Technical Problem

该方法通过图像层融合,即融合红外与可见光图像深度特征,提高了目标检测精度;通过结果层融合,即基于DS证据理论的可见光与红外决策级融合,实现了鲁棒性目标检测识别结果;解决了复杂环境下单传感器目标探测效率低、虚警率高等问题,提高多源日标检测的准确性和鲁棒性

Benefits of technology

本发明实施例提供的一种基于多层次融合的无人机异源图像目标检测方法及系统,采用以下步骤构成的技术方案:S101:获得可见光图像序列和红外光图像序列;可见光图像序列中包括多帧拍摄时间连续的可见光图像;红外光图像序列含有多帧拍摄时间连续的红外光图像; S102:提取可见光图像序列的多种可见光特征,对多种可见光特征进行融合,得到目标的可见光融合特征;提取红外光图像序列的多种红外光特征,对多种红外光特征进行融合,得到目标的红外光融合特征;S103:基于DS证据理论对所述多种可见光特征进行可信度融合,得到第一置信度;基于DS证据理论对所述多种红外光特征进行可信度融合,得到第二置信度;S104:基于所述可见光融合特征进行目标检测,得到第一检测目标;基于所述红外光融合特征进行目标检测,得到第二检测目标;S105:若第一检测目标和第二检测目标为同一个目标,对可见光融合特征和红外光融合特征进行融合,得到目标融合特征;对目标融合特征进行可信度分析,获得目标融合特征的第三置信度;若第三置信度大于预设值,确定检测到的第一检测目标和第二检测目标为真目标;S106:若第一检测目标和第二检测目标不是同一个目标,若第一检测目标是特定目标,按照设定步长调低第一置信度的值,若第二检测目标是特定目标,按照设定步长调低第二置信度的值;S107:基于调低以后的第一置信度或者第二置信度,确定基于可见光融合特征和红外光融合特征是否可以同时检测到所述特定目标;S108:若同时检测到所述特定目标,基于第一置信度的当前取值和第二置信度的当前取值,获得当前融合置信度;若当前融合置信度大于预设值,确定检测到的特定目标为真目标;S109:若未同时检测到所述特定目标,且检测到特定目标的传感器的置信度小于设定阈值,确定所检测到的目标为伪目标;S110:若未同时检测到所述特定目标,且检测到特定目标的传感器的置信度大于或者等于设定阈值,执行步骤S106~S109的操作。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118052975B_ABST
    Figure CN118052975B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-level fusion-based unmanned aerial vehicle heterogenous image target detection method and system.The present application carries out fusion operation in image input level and result output level, carries out fusion to the feature of visible light image sequence and infrared light image sequence, carries out confidence analysis and fusion based on DS evidence theory to the feature extracted, realizes high-precision target detection recognition result by visible light and infrared image target detection of depth feature fusion, by the design of the multi-source feature fusion module based on deep convolution network, fusion infrared and visible light image depth feature, improve target detection precision, solve the problem of low efficiency, false alarm rate and other problems of single sensor target detection in complex environment, improve the accuracy and robustness of multi-source day standard detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and more specifically, to a method and system for detecting heterogeneous UAV images based on multi-level fusion. Background Technology

[0002] With the rapid development of computer vision technology, various camera-based target detection devices are widely used in various fields. Currently, target detection devices are installed on drones to capture and track targets. Target detection is the foundation of tracking. Currently, most scenarios use image data acquired by a single sensor for target detection. However, in complex environments such as high-speed movement, blurred images, rainy days, and low visibility, single-sensor target detection has low accuracy and a high false alarm rate. Summary of the Invention

[0003] To address the shortcomings of existing technologies, the present invention aims to provide a method and system for detecting heterogeneous targets in UAV images based on multi-level fusion. This method improves target detection accuracy through image-level fusion, namely, fusing depth features from infrared and visible light images; and achieves robust target detection and recognition results through result-level fusion, namely, decision-level fusion of visible light and infrared based on DS evidence theory. It solves the problems of low efficiency and high false alarm rate in single-sensor target detection under complex environments, thereby improving the accuracy and robustness of multi-source target detection.

[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, embodiments of the present invention provide a method for detecting heterogeneous targets in UAV images based on multi-level fusion, comprising: S101: Obtain a visible light image sequence and an infrared light image sequence; the visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared light image sequence contains multiple frames of infrared light images captured in consecutive time. S102: Extract multiple visible light features from the visible light image sequence, fuse the multiple visible light features to obtain the visible light fusion feature of the target; extract multiple infrared light features from the infrared light image sequence, fuse the multiple infrared light features to obtain the infrared light fusion feature of the target; S103: Based on the DS evidence theory, the credibility of the multiple visible light features is fused to obtain a first confidence level; based on the DS evidence theory, the credibility of the multiple infrared light features is fused to obtain a second confidence level; S104: Target detection is performed based on the visible light fusion features to obtain a first detected target; target detection is performed based on the infrared light fusion features to obtain a second detected target; S105: If the first detection target and the second detection target are the same target, fuse the visible light fusion feature and the infrared light fusion feature to obtain the target fusion feature; perform a confidence analysis on the target fusion feature to obtain the third confidence level of the target fusion feature; if the third confidence level is greater than a preset value, determine that the detected first detection target and the second detection target are true targets; S106: If the first detection target and the second detection target are not the same target, if the first detection target is a specific target, lower the value of the second confidence level according to the adaptation step size, and if the second detection target is a specific target, lower the value of the first confidence level according to the adaptation step size. S107: Based on the lowered first or second confidence level, determine whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features; S108: If the specific target is detected simultaneously, obtain the current fused confidence level based on the current value of the first confidence level and the current value of the second confidence level; if the current fused confidence level is greater than a preset value, determine that the detected specific target is a real target; S109: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target; S110: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to the set threshold, execute the operations of steps S106 to S109.

[0005] Optionally, based on the current value of the first confidence level and the current value of the second confidence level, the current fusion confidence level is obtained, including: The current fusion confidence is obtained by weighting and summing the current values ​​of the first confidence level and the current values ​​of the second confidence level.

[0006] Optionally, the visible light image sequence includes multiple visible light features such as texture features, edge features, and color features; the fusion of multiple visible light features to obtain the visible light fusion features of the target includes: Texture features, color features, and edge features are weighted and fused according to the correspondence of pixel positions to obtain visible light fused features.

[0007] Optionally, the visible light image sequence may also include multi-frame motion features, histogram features, multi-scale features, aspect ratio features, and spatial scene analysis features.

[0008] Optionally, target detection is performed based on the visible light fusion features to obtain a first detected target, including: When the visible light fusion features are input into the target detection model, the target detection model identifies the first target corresponding to the visible light fusion features.

[0009] Secondly, embodiments of the present invention provide a multi-level fusion-based UAV heterogeneous image target detection system, the system comprising: The acquisition module is used to acquire visible light image sequences and infrared light image sequences; the visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared light image sequence contains multiple frames of infrared light images captured in consecutive time. The feature extraction module is used to extract multiple visible light features from visible light image sequences, fuse these features to obtain the visible light fusion features of the target; and to extract multiple infrared light features from infrared light image sequences, fuse these features to obtain the infrared light fusion features of the target. The confidence fusion module is used to fuse the confidence of the multiple visible light features based on the DS evidence theory to obtain a first confidence level; and to fuse the confidence of the multiple infrared light features based on the DS evidence theory to obtain a second confidence level. The target detection module is used to perform target detection based on the visible light fusion features to obtain a first detected target; and to perform target detection based on the infrared light fusion features to obtain a second detected target. If the first and second detected targets are the same target, the visible light fusion features and the infrared light fusion features are fused to obtain a target fusion feature. A confidence analysis is performed on the target fusion feature to obtain a third confidence level. If the third confidence level is greater than a preset value, the detected first and second detected targets are determined to be true targets. If the first and second detected targets are not the same target, and if the first detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size; if the second detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size. The first confidence level should be lowered by adjusting the step size. Based on the lowered first or second confidence level, it is determined whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. If the specific target is detected simultaneously, the current fusion confidence level is obtained based on the current values ​​of the first and second confidence levels. If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to a set threshold, the above steps are performed.

[0010] The multi-level fusion-based UAV heterogeneous image target detection method and system described in this invention have the following advantages: The present invention provides a method and system for detecting heterogeneous targets in UAV images based on multi-level fusion, which adopts the following steps: S101: Obtain a visible light image sequence and an infrared light image sequence; the visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared light image sequence contains multiple frames of infrared light images captured in consecutive time. S102: Extract multiple visible light features from the visible light image sequence, fuse these features to obtain the visible light fusion feature of the target; extract multiple infrared light features from the infrared light image sequence, fuse these features to obtain the infrared light fusion feature of the target; S103: Perform confidence fusion on the multiple visible light features based on DS evidence theory to obtain a first confidence level; perform confidence fusion on the multiple infrared light features based on DS evidence theory to obtain a second confidence level; S104: Perform target detection based on the visible light fusion feature to obtain a first detected target; perform target detection based on the infrared light fusion feature to obtain a second detected target; S105: If the first detected target and the second detected target are the same target, fuse the visible light fusion feature and the infrared light fusion feature to obtain the target fusion feature; perform confidence analysis on the target fusion feature to obtain a third confidence level of the target fusion feature; if the third confidence level is greater than a preset value, determine that the detected first and second detected targets are true targets. S106: If the first detection target and the second detection target are not the same target, if the first detection target is a specific target, lower the value of the first confidence level by a set step size; if the second detection target is a specific target, lower the value of the second confidence level by a set step size. S107: Based on the lowered first or second confidence level, determine whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. S108: If the specific target is detected simultaneously, obtain the current fusion confidence level based on the current value of the first confidence level and the current value of the second confidence level. If the current fusion confidence level is greater than a preset value, determine that the detected specific target is a true target. S109: If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is less than a set threshold, determine that the detected target is a false target. S110: If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is greater than or equal to a set threshold, execute steps S106 to S109.

[0011] By employing the above technical solution, multiple visible light features are extracted from the visible light image sequence and then fused to obtain the target's visible light fusion feature. This results in the visible light fusion feature containing more visible light features, enabling it to more accurately represent the target it refers to at the visible light level. Similarly, the obtained infrared fusion feature contains more infrared features, allowing it to more accurately represent the target it refers to at the infrared level. For evaluating the confidence level of the visible light fusion feature, the confidence level of the multiple visible light features is fused based on the DS evidence theory to obtain a first confidence level. Using this first confidence level as the confidence level of the visible light fusion feature demonstrates high accuracy in confidence assessment. Similarly, using a second confidence level as the confidence level of the infrared fusion feature also demonstrates high accuracy in confidence assessment. Then, target detection is performed based on the visible light fusion feature and the infrared fusion feature, respectively, yielding the first and second detected targets. For both visible and infrared light levels, based on the high accuracy of target characterization using visible light fusion features and infrared light fusion features, the accuracy of detecting the first and second targets using these two features is also high. This means that the confidence level of the first and second targets detected using visible light fusion features is high. Compared to target detection based on single features, the accuracy and reliability of target detection using the above method are improved.

[0012] Even though the accuracy of the first and second detection targets has been improved, this solution does not stop there. After obtaining the first and second detection targets, steps S105 to S109 are performed to further combine the two for target identification: First, it is determined whether the first and second detection targets are the same target. If they are the same target, the solution goes further by fusing the visible light fusion features and the infrared light fusion features to obtain the target fusion features. Then, a confidence analysis is performed on the target fusion features to obtain the confidence level of the target fusion features, called the third confidence level. Then, based on the third confidence level, it is determined whether the target detected by the two types of sensors is a real target (the visible light image sequence and the infrared light image sequence are acquired by the visible light sensor and the infrared light sensor, respectively). Specifically, if the third confidence level is greater than a preset value, the detected first and second detection targets are determined to be real targets, thus improving the accuracy of target detection. Furthermore, if the first detection target and the second detection target are not the same target, and if the first detection target is a specific target, the value of the first confidence level is lowered by a set step size; if the second detection target is a specific target, the value of the second confidence level is lowered by a set step size. Based on the lowered first or second confidence level, it is determined whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. If the specific target is detected simultaneously, the current fusion confidence level is obtained based on the current value of the first confidence level and the current value of the second confidence level. If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is less than a set threshold, the detected target is determined to be a false target. If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is greater than or equal to the set threshold, the above steps are executed until the detected target is determined to be a true target or the confidence level of the sensor that detected the specific target is less than the set threshold. By adopting the above scheme, the accuracy has been greatly improved, especially for the detection of targets with ghosting, thus reducing the false alarm rate of target detection.

[0013] In summary, this invention achieves image fusion at the input layer and result (confidence score) fusion at the output layer, and performs target recognition and judgment based on the fusion result, thus improving the accuracy of target detection. Specifically, image layer fusion: High-precision target detection and recognition results are achieved through visible light and infrared image target detection and recognition technology using deep feature fusion. Addressing the high-precision perception requirements of various scenes, an end-to-end visible light and infrared image target detection and recognition network is constructed. By designing a multi-source feature fusion module based on a deep convolutional network, depth features from infrared and visible light images are fused to improve target detection accuracy. Result layer fusion: Robust target detection and recognition results are achieved through decision-level fusion of visible light and infrared images based on DS evidence theory. Addressing the robust perception requirements of various scenes, techniques such as multi-feature extraction methods for targets, multi-feature confidence analysis methods for targets, and decision-level fusion based on DS evidence theory are studied to solve problems such as low efficiency and high false alarm rate of single-sensor target detection in complex environments, thereby improving the accuracy and robustness of multi-source target detection. Attached Figure Description

[0014] The present invention includes the following figures: Figure 1 This is a flowchart of a method for detecting heterogeneous targets in UAV images based on multi-level fusion, provided by an embodiment of the present invention. Figure 2 This is a flowchart of another UAV heterogeneous image target detection method based on multi-level fusion provided in an embodiment of the present invention. Figure 3 This is a block structure diagram of an electronic device provided in an embodiment of the present invention.

[0015] In the diagram: 500, bus; 501, receiver; 502, processor; 503, transmitter; 504, memory; 505, bus interface. Detailed Implementation

[0016] The present invention will be further described in detail below with reference to the accompanying drawings.

[0017] Example 1 like Figure 1 As shown, this embodiment of the invention provides a method for detecting targets in heterogeneous images of unmanned aerial vehicles based on multi-level fusion, including the steps described in S201~S210 below.

[0018] S201: Obtain visible light image sequences and infrared light image sequences.

[0019] The visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared image sequence contains multiple frames of infrared light images captured in consecutive time. The visible light image sequence and the infrared image sequence are acquired by a visible light sensor and an infrared light sensor, respectively. As an optional implementation, the visible light sensor and the infrared light sensor are mounted on the drone.

[0020] S202: Extract multiple visible light features from the visible light image sequence, fuse the multiple visible light features to obtain the visible light fusion feature of the target; extract multiple infrared light features from the infrared light image sequence, fuse the multiple infrared light features to obtain the infrared light fusion feature of the target.

[0021] In this embodiment of the invention, different extraction methods can be used for different features of the visible light image sequence. In this embodiment, various visible light features may include multi-frame motion features, histogram features, multi-scale features, aspect ratio features, texture features, edge features, color features, and spatial scene features. Among these, histogram features, multi-scale features, aspect ratio features, texture features, edge features, color features, and spatial scene features currently have mature and effective feature extraction methods. For example, the Gray Level Co-occurrence Matrix (GLCM) can be used to extract the texture features of an image. Since the visible light image sequence contains multiple frames of visible light images, its overall texture features can be obtained by extracting the texture features of each visible light image using the GLCM, and then weighting and filtering the texture features of multiple visible light images to obtain the overall texture features of the visible light image sequence. For multi-frame motion features, optical flow can be used to extract them. The motion trajectories obtained by optical flow tracking can be used to construct a motion matrix, which is then used as the multi-frame motion features.

[0022] Similarly, the extraction methods for different features of infrared light image sequences can be the same as those for different features of visible light image sequences.

[0023] To fuse multiple visible light features to obtain the visible light fusion feature of a target, the multiple visible light features can be weighted and summed to obtain the overall visible light fusion feature, and the multiple infrared light features can be weighted and summed to obtain the overall infrared light fusion feature.

[0024] For example, a visible light image sequence may contain multiple visible light features, including texture features, edge features, and color features. By fusing these multiple visible light features, we can obtain the visible light fusion features of the target. This involves weighting and fusing the texture features, color features, and edge features according to the correspondence between pixel positions to obtain the visible light fusion features.

[0025] Optionally, the visible light image sequence may also include multi-frame motion features, histogram features, multi-scale features, aspect ratio features, and spatial scene analysis features.

[0026] Optionally, various infrared light features include multi-frame motion features, histogram features, multi-scale features, aspect ratio features, texture features, edge features, color features, and spatial scene features. The methods for extracting various infrared light features from infrared light image sequences are the same as those for extracting various visible light features from visible light image sequences, and will not be repeated here.

[0027] By employing the above scheme, multiple visible light features are extracted from the visible light image sequence and fused to obtain the target's visible light fusion feature. This results in the visible light fusion feature containing more visible light features, enabling it to more accurately represent the target it represents at the visible light level. Similarly, the infrared fusion feature contains more infrared features, allowing it to more accurately represent the target it represents at the infrared level. Based on this, target detection is then performed using both visible light and infrared fusion features, yielding the first and second detected targets, respectively. For both the visible and infrared levels, the accuracy of target characterization based on these features is high, and the accuracy of the first and second detected targets obtained using both is also high. Compared to target detection based on single features, the accuracy and reliability of target detection using this method are improved. In other words, since both visible light and infrared fusion features accurately characterize the targets they represent, the detection results obtained using both are highly accurate.

[0028] However, up to this point, target detection has been based on infrared and visible light sensors respectively, which are all one-sided detection methods. Even though the accuracy of target detection is already high, there is still a high false alarm rate for single-sided sensor target detection, especially for situations involving ghost images. Therefore, to address these issues, the technical solution proposed in this application includes the following steps to further improve the accuracy of target detection.

[0029] S203: Based on the DS evidence theory, the credibility of the multiple visible light features is fused to obtain a first confidence level; based on the DS evidence theory, the credibility of the multiple infrared light features is fused to obtain a second confidence level.

[0030] Specifically, a confidence analysis can be performed on each of the multiple visible light features to obtain the confidence level of each feature. Then, the confidence levels of the multiple visible light features are fused using the DS evidence theory to obtain a first confidence level. This first confidence level is used as the confidence level of the fused visible light feature obtained above. The specific method for performing the confidence analysis on each visible light feature can be as follows: for each specific feature in the visible light image sequence, the reciprocal of the variance of the specific features of multiple visible light images is used as the confidence level of the specific feature of the visible light image sequence. Alternatively, traditional confidence analysis methods can be used to analyze the confidence level of the specific feature. The specific feature refers to the visible light feature among the multiple visible light features.

[0031] Similarly, a confidence analysis is performed on each of the multiple infrared light features to obtain the confidence level of each feature. Then, the confidence levels of the multiple infrared light features are fused using the DS evidence theory to obtain a second confidence level. This second confidence level is used as the confidence level of the fused infrared light feature obtained above. The specific method for performing the confidence analysis on each infrared light feature can be as follows: for each specific feature in the infrared light image sequence, the inverse of the variance of the specific features of multiple infrared light images is used as the confidence level of the specific feature of the infrared light image sequence. Alternatively, traditional confidence analysis methods can be used to analyze the confidence level of the specific feature. The specific feature refers to the infrared light feature among the multiple infrared light features. After obtaining the confidence level of each infrared light feature, the confidence levels of the multiple infrared light features are fused using the DS evidence theory to obtain the second confidence level.

[0032] By adopting the above scheme, the first confidence level as the confidence level of visible light fusion features has high accuracy in confidence assessment. Similarly, the second confidence level as the confidence level of infrared light fusion features has high accuracy in confidence assessment.

[0033] S204: Target detection is performed based on the visible light fusion features to obtain a first detection target; target detection is performed based on the infrared light fusion features to obtain a second detection target.

[0034] In this embodiment of the invention, the first confidence level and the second confidence level can be obtained by using the confidence levels of the target detection model output for the first detection target and the second detection target, that is, using the confidence level of the first detection target as the first confidence level and the confidence level of the second detection target as the second confidence level.

[0035] As an optional implementation, target detection is performed based on the visible light fusion features to obtain a first detection target, including: When the visible light fusion features are input into the target detection model, the target detection model identifies the first target corresponding to the visible light fusion features.

[0036] Target detection is performed based on the infrared light fusion features to obtain a second detection target, including: The infrared light fusion features are input into the target detection model, and the target detection model identifies the second target corresponding to the infrared light fusion features.

[0037] In this embodiment of the invention, the target detection model may be a convolutional neural network (CNN) or a recurrent neural network (RNN).

[0038] By adopting the above scheme, based on the high accuracy of target characterization using visible light fusion features and infrared light fusion features, the accuracy of the first and second detection targets obtained by detection using both is also high. That is, the confidence level of the first detection target obtained by visible light fusion feature detection is high, and the confidence level of the second detection target is also high. Compared with target detection based on single features, the accuracy and reliability of target detection performed by the above method are improved.

[0039] Even though the accuracy of the first and second detection targets has been improved, this solution does not stop there. After obtaining the first and second detection targets, steps S205 to S209 are performed to further combine the two for target identification, thereby further improving the accuracy of target detection and reducing the false alarm rate.

[0040] After step S204, it is determined whether the first detection target and the second detection target are the same target. If the first detection target and the second detection target are the same target, step S205 is executed.

[0041] S205: If the first detection target and the second detection target are the same target, fuse the visible light fusion feature and the infrared light fusion feature to obtain the target fusion feature; perform a confidence analysis on the target fusion feature to obtain the third confidence level of the target fusion feature.

[0042] Determine whether the third confidence level is greater than the preset value.

[0043] S2051: If the third confidence level is greater than the preset value, determine that the first and second detected targets are true targets.

[0044] The specific method for performing credibility analysis on the target fusion features to obtain the third confidence level of the target fusion features can refer to the specific method for performing credibility analysis on each of the multiple visible light features described above, and will not be repeated here. The target fusion features are obtained by fusing the visible light fusion features and the infrared light fusion features, which can be done by weighting the pixel values ​​of the feature points according to the correspondence between pixel points. In this embodiment of the invention, the preset value can be 0.55, 0.6, 0.65, 0.7, etc., and can be set according to the actual scenario.

[0045] S2052: If the third confidence level is less than or equal to the preset value, determine that the detected first and second detection targets are false targets.

[0046] By employing the above scheme, targets are detected in images acquired by both visible light and infrared sensors. Furthermore, features from the two sensors are fused, and confidence analysis is performed on the fused features. The calculated confidence level (third confidence level) is used to evaluate whether the detected target is a genuine target. Compared to target detection based solely on single-sensor images, this method achieves higher accuracy and a lower false alarm rate. Moreover, using the confidence level of the fused features as a criterion for overall detection accuracy further controls the false alarm rate. If the third confidence level is greater than a preset value, the first and second detected targets are determined to be genuine targets, improving the accuracy of genuine target detection and reducing the false alarm rate.

[0047] S206: If the first detection target and the second detection target are not the same target, if the first detection target is a specific target, lower the value of the second confidence level according to the adaptation step size, and if the second detection target is a specific target, lower the value of the first confidence level according to the adaptation step size.

[0048] Optionally, step S206 may further include determining whether the second detection target is a specific target.

[0049] In this embodiment, if the first detection target and the second detection target are not the same target, it means that the infrared camera (sensor) and the visible light camera (sensor) are detecting different targets. Without further judgment, there will be a high false alarm rate. That is to say, only when both cameras (with overlapping fields of view) detect the same target can it be ensured that the detected target is real. To improve the accuracy of target detection, if the first detection target and the second detection target are not the same target, and the first detection target is a specific target, the value of the second confidence level is lowered by a set step size; if the second detection target is a specific target, the value of the first confidence level is lowered by an adaptive step size, that is, the confidence level of the sensor that did not detect the specific target is lowered.

[0050] In this embodiment of the invention, the specific target can be set to a real target, not a phantom, or the detection target corresponding to the higher confidence value between the first confidence level and the second confidence level can be set as the specific target. For example, if the first confidence level is greater than or equal to the second confidence level, then the first detection target is set as the specific target; if the second confidence level is greater than the first confidence level, then the second detection target is set as the specific target.

[0051] If the first detected target is a specific target, it means that a real target was detected based on visible light fusion features. However, if the target was not detected based on infrared light fusion features, then the target output by the infrared light fusion feature detection is not the specific target; it could be another target (object). Since the target detection model outputs targets enclosed by the highest confidence bounding box, and targets enclosed by lower confidence bounding boxes are not included in the model's output, if the second detected target output by the infrared light fusion feature detection model is determined to be not the specific target, then the confidence of the second detected target can be reduced (lowering the second confidence level). This means the target output by the infrared light fusion feature detection model is no longer the second detected target. Then, this output target is compared with the first target to see if they are the same target. The adaptation step size is adaptively adjusted. It can be set to be less than the difference between the second confidence level and the first sequential confidence level, or greater than the difference between the second confidence level and the second sequential confidence level. The first sequential confidence level is the confidence level of detection boxes in other detection frames where the confidence level is less than the second confidence level but greater than other confidence levels, based on infrared light fusion features. The second sequential confidence level is the confidence level of detection boxes in other detection frames where the confidence level is less than the first sequential confidence level but greater than other confidence levels, based on infrared light fusion features.

[0052] If the second detected target is a specific target, it means that a real target was detected based on infrared light fusion features. However, the target was not detected based on visible light fusion features. This means the target output by the visible light fusion feature detection is not the specific target; it could be another target (object). Since the target detection model outputs results for targets enclosed by the highest confidence bounding box, and targets enclosed by lower confidence bounding boxes are not included in the model's output, it's determined that the first detected target output by the target detection model based on visible light fusion features is not the specific target. Therefore, the confidence of the first detected target can be lowered (reducing the first confidence level), so that the target output by the target detection model based on visible light fusion features is no longer the first detected target. Then, this output target is compared with the second target to see if they are the same target. The adaptation step size is adaptively adjusted. It can be set to be less than the difference between the first confidence level and the third sequential confidence level, or greater than the difference between the first confidence level and the fourth sequential confidence level. The third sequential confidence level is the confidence level of detection boxes in other detection frames where the confidence level is less than the first confidence level but greater than other confidence levels, based on visible light fusion features. The fourth sequential confidence level is the confidence level of detection boxes in other detection frames where the confidence level is less than the third sequential confidence level but greater than other confidence levels, based on visible light fusion features.

[0053] After adjusting the confidence level, further judgment is needed to see if the two sensors detect the same target simultaneously.

[0054] S207: Based on the lowered first or second confidence level, determine whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features.

[0055] Based on the lowered first or second confidence level, the specific method for determining whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features is as follows: If the first detection target is a specific target, the value of the second confidence level is lowered according to the adaptation step size to obtain the first-order detection result of the adjusted infrared light sensor. This detection result includes the first-order detected target and the first-order confidence level. Then, the first-order detected target is compared with the first detected target to see if they are the same target.

[0056] If the second detection target is a specific target, the value of the first confidence level is lowered according to the adaptation step size to obtain the adjusted third-order detection result of the visible light sensor. This detection result includes the third-order detected target and the third-order confidence level. Then, the third-order detected target is compared with the second detected target to see if they are the same target.

[0057] If the specific target is detected simultaneously in step S207, the following step S208 is executed.

[0058] S208: If the specific target is detected simultaneously, obtain the current fused confidence level based on the current value of the first confidence level and the current value of the second confidence level.

[0059] The current fusion confidence level is determined to be greater than the preset value.

[0060] S2081: If the current fusion confidence is greater than the preset value, determine that the detected specific target is a true target.

[0061] At this point, based on step S206, if the first detection target is a specific target, the current value of the first confidence remains unchanged, or is the value of the first confidence in step S206, while the current value of the second confidence is adjusted to the first sequential confidence, or the current value of the second confidence is equal to the value of the second confidence in step S206 minus the adaptation step size.

[0062] If the second detection target is a specific target, the current value of the second confidence level is still the value of the second confidence level in step S206, while the current value of the first confidence level is adjusted to the third sequential confidence level, or the current value of the first confidence level is equal to the value of the first confidence level in step S206 minus the adaptation step size.

[0063] Optionally, the current fusion confidence is obtained based on the current value of the first confidence and the current value of the second confidence by performing a weighted sum of the current values ​​of the first confidence and the current values ​​of the second confidence.

[0064] If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. The preset value can be 0.4, 0.5, 0.55, 0.56, 0.6, etc., and can be set according to the actual needs of the scenario.

[0065] S2082: If the current fusion confidence is less than or equal to the preset value, the detected specific target is determined to be a false target.

[0066] In this embodiment of the invention, if two sensors detect the same target simultaneously, a fusion confidence score should be used for further judgment, which can improve the accuracy of target detection and reduce the false alarm rate of target detection.

[0067] If the specific target is not detected simultaneously in step S207, proceed to step 209.

[0068] S209: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target.

[0069] Step S209 also includes determining whether the confidence level of the sensor that did not detect the specific target is less than a set threshold.

[0070] In other words, if the confidence level of the sensor that did not detect the specific target (if the visible light sensor detects the real target, the sensor that did not detect the specific target is the infrared light sensor; if the infrared light sensor detects the real target, the sensor that did not detect the specific target is the visible light sensor, i.e., the first confidence level or the second confidence level) decreases to a certain level, and neither sensor can detect the same target at the same time, it means that the sensor that did not detect the specific target did not detect the target, and the detected specific target is identified as a false target.

[0071] The threshold value can be any value between 0.2 and 0.8, such as 0.3, 0.4, 0.5, 0.6, etc.

[0072] S210: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to the set threshold, execute the operations of steps S206 to S209.

[0073] In other words, if a specific target is not detected simultaneously, the sensor that did not detect the target should continuously adjust its output and compare the outputs until the outputs of the two sensors indicate the same target, or the confidence level becomes too low to be reliable. This improves the accuracy of target detection.

[0074] In summary, the process first determines whether the first and second detection targets are the same target. If they are, it goes further by fusing the visible light fusion features and infrared light fusion features to obtain target fusion features. Then, it performs a confidence analysis on the target fusion features to obtain the confidence level of the target fusion features, called the third confidence level. Based on the third confidence level, it determines whether the target detected by the two types of sensors is a real target (the visible light image sequence and the infrared light image sequence are acquired by the visible light sensor and the infrared light sensor, respectively). Specifically, if the third confidence level is greater than a preset value, the detected first and second detection targets are determined to be real targets, thus improving the accuracy of target detection. Furthermore, if the first detection target and the second detection target are not the same target, and if the first detection target is a specific target, the value of the first confidence level is lowered by a set step size; if the second detection target is a specific target, the value of the second confidence level is lowered by a set step size. Based on the lowered first or second confidence level, it is determined whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. If the specific target is detected simultaneously, the current fusion confidence level is obtained based on the current value of the first confidence level and the current value of the second confidence level. If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is less than a set threshold, the detected target is determined to be a false target. If the specific target is not detected simultaneously, and the confidence level of the sensor that detected the specific target is greater than or equal to the set threshold, the above steps are executed until the detected target is determined to be a true target or the confidence level of the sensor that detected the specific target is less than the set threshold. By adopting the above scheme, the accuracy has been greatly improved, especially for the detection of targets with ghosting, thus reducing the false alarm rate of target detection.

[0075] Example 1 This invention also provides another method for detecting heterogeneous targets in UAV images based on multi-level fusion, such as... Figure 2 As shown, the UAV heterogeneous image target detection method based on multi-level fusion includes the technical solutions described in steps S101 to S110: S101: Obtain visible light image sequences and infrared light image sequences.

[0076] S102: Extract multiple visible light features from the visible light image sequence, fuse the multiple visible light features to obtain the visible light fusion feature of the target; extract multiple infrared light features from the infrared light image sequence, fuse the multiple infrared light features to obtain the infrared light fusion feature of the target.

[0077] S103: Based on the DS evidence theory, the credibility of the multiple visible light features is fused to obtain a first confidence level; based on the DS evidence theory, the credibility of the multiple infrared light features is fused to obtain a second confidence level.

[0078] S104: Target detection is performed based on the visible light fusion features to obtain a first detection target; target detection is performed based on the infrared light fusion features to obtain a second detection target.

[0079] S105: If the first detection target and the second detection target are the same target, fuse the visible light fusion feature and the infrared light fusion feature to obtain the target fusion feature; perform a confidence analysis on the target fusion feature to obtain the third confidence level of the target fusion feature; if the third confidence level is greater than the preset value, determine that the detected first detection target and the second detection target are true targets.

[0080] S106: If the first detection target and the second detection target are not the same target, if the first detection target is a specific target, lower the value of the second confidence level according to the adaptation step size, and if the second detection target is a specific target, lower the value of the first confidence level according to the adaptation step size.

[0081] S107: Based on the lowered first or second confidence level, determine whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features.

[0082] S108: If the specific target is detected simultaneously, obtain the current fused confidence level based on the current value of the first confidence level and the current value of the second confidence level; if the current fused confidence level is greater than a preset value, determine that the detected specific target is a true target.

[0083] S109: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target.

[0084] S110: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to the set threshold, execute the operations of steps S106 to S109.

[0085] In this embodiment of the invention, steps S101 to S110 correspond to steps S201 to S210 in the above embodiment, and the implementation methods of the steps are the same. For details, please refer to the above description of steps S201 to S210, which will not be repeated here.

[0086] In summary, by fusing the images from the input layer and the results (confidence scores) from the output layer, and then using the fused results for target recognition and judgment, the accuracy of target detection is improved. Image layer fusion: High-precision target detection and recognition results are achieved through deep feature fusion of visible light and infrared images. In response to the high-precision perception requirements of the scene, the construction of an end-to-end visible light and infrared image target detection and recognition network is studied. By designing a multi-source feature fusion module based on deep convolutional network, the depth features of infrared and visible light images are fused to improve the target detection accuracy.

[0087] Result-level fusion: Robust target detection and recognition results are achieved through decision-level fusion of visible light and infrared based on DS evidence theory. In response to the robust perception requirements of the field, this study investigates target multi-feature extraction methods, target multi-feature credibility analysis methods, and DS evidence theory decision-level fusion technologies to solve problems such as low efficiency and high false alarm rate of single-sensor target detection in complex environments, thereby improving the accuracy and robustness of multi-source target detection.

[0088] Based on multi-layered fusion and judgment, the accuracy of target detection is improved and the false alarm rate is reduced.

[0089] Based on the UAV heterogeneous image target detection method based on multi-level fusion provided in the above embodiments, this invention provides a UAV heterogeneous image target detection system based on multi-level fusion for performing the above method. The system includes: The acquisition module is used to acquire visible light image sequences and infrared light image sequences; the visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared light image sequence contains multiple frames of infrared light images captured in consecutive time.

[0090] The feature extraction module is used to extract multiple visible light features from the visible light image sequence, fuse the multiple visible light features to obtain the visible light fusion feature of the target; and to extract multiple infrared light features from the infrared light image sequence, fuse the multiple infrared light features to obtain the infrared light fusion feature of the target.

[0091] The confidence fusion module is used to perform confidence fusion on the multiple visible light features based on DS evidence theory to obtain a first confidence level; and to perform confidence fusion on the multiple infrared light features based on DS evidence theory to obtain a second confidence level.

[0092] The target detection module is used to perform target detection based on the visible light fusion features to obtain a first detected target; and to perform target detection based on the infrared light fusion features to obtain a second detected target. If the first and second detected targets are the same target, the visible light fusion features and the infrared light fusion features are fused to obtain a target fusion feature. A confidence analysis is performed on the target fusion feature to obtain a third confidence level. If the third confidence level is greater than a preset value, the detected first and second detected targets are determined to be true targets. If the first and second detected targets are not the same target, and if the first detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size; if the second detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size. The first confidence level should be lowered by adjusting the step size. Based on the lowered first or second confidence level, it is determined whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. If the specific target is detected simultaneously, the current fusion confidence level is obtained based on the current values ​​of the first and second confidence levels. If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to a set threshold, the above steps are performed.

[0093] The specific manner in which each module performs its operations in the system described in the above embodiments has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0094] This invention also provides an electronic device, such as... Figure 3 As shown, it includes a memory 504, a processor 502, and a computer program stored in the memory 504 and executable on the processor 502. When the processor 502 executes the program, it implements the steps of any of the methods of the UAV heterogeneous image target detection method based on multi-level fusion described above.

[0095] Among them, Figure 3In this document, a bus architecture (represented by bus 500) is used. Bus 500 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 502 and memory represented by memory 504. Bus 500 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 505 provides an interface between bus 500 and receiver 501 and transmitter 503. Receiver 501 and transmitter 503 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 502 is responsible for managing bus 500 and general processing, while memory 504 can be used to store data used by processor 502 during operation.

[0096] This invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims listing several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.

[0097] The contents not described in detail in this specification are existing technologies known to those skilled in the art.

Claims

1. A method for detecting heterogeneous targets in UAV images based on multi-level fusion, characterized in that, include: S101: Obtain a visible light image sequence and an infrared light image sequence; the visible light image sequence includes multiple frames of visible light images captured in consecutive time. An infrared image sequence contains multiple frames of infrared images captured in consecutive time. S102: Extract multiple visible light features from the visible light image sequence, fuse the multiple visible light features to obtain the visible light fusion feature of the target; extract multiple infrared light features from the infrared light image sequence, fuse the multiple infrared light features to obtain the infrared light fusion feature of the target; S103: Based on the DS evidence theory, the credibility of the multiple visible light features is fused to obtain a first confidence level; based on the DS evidence theory, the credibility of the multiple infrared light features is fused to obtain a second confidence level; S104: Target detection is performed based on the visible light fusion features to obtain a first detected target; target detection is performed based on the infrared light fusion features to obtain a second detected target; S105: If the first detection target and the second detection target are the same target, fuse the visible light fusion feature and the infrared light fusion feature to obtain the target fusion feature; A credibility analysis is performed on the target fusion features to obtain the third confidence level of the target fusion features; If the third confidence level is greater than the preset value, the detected first and second targets are determined to be true targets. S106: If the first detection target and the second detection target are not the same target, if the first detection target is a specific target, lower the value of the second confidence level according to the adaptation step size, and if the second detection target is a specific target, lower the value of the first confidence level according to the adaptation step size. S107: Based on the lowered first or second confidence level, determine whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features; S108: If the specific target is detected simultaneously, obtain the current fusion confidence based on the current value of the first confidence and the current value of the second confidence; If the current fusion confidence is greater than the preset value, the detected specific target is determined to be a real target; S109: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target; S110: If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to the set threshold, execute the operations of steps S106 to S109.

2. The UAV heterogeneous image target detection method based on multi-level fusion according to claim 1, characterized in that, Based on the current value of the first confidence level and the current value of the second confidence level, the current fusion confidence level is obtained, including: The current fusion confidence is obtained by weighting and summing the current values ​​of the first confidence level and the current values ​​of the second confidence level.

3. The UAV heterogeneous image target detection method based on multi-level fusion according to claim 1, characterized in that, Visible light image sequences possess various visible light features, including texture features, edge features, and color features; The process of fusing multiple visible light features to obtain the visible light fused features of the target includes: Texture features, color features, and edge features are weighted and fused according to the correspondence of pixel positions to obtain visible light fused features.

4. The UAV heterogeneous image target detection method based on multi-level fusion according to claim 1, characterized in that, The visible light image sequence also includes multiple visible light features such as multi-frame motion features, histogram features, multi-scale features, aspect ratio features, and spatial scene analysis features.

5. The UAV heterogeneous image target detection method based on multi-level fusion according to claim 1, characterized in that, Target detection is performed based on the visible light fusion features to obtain a first detection target, including: When the visible light fusion features are input into the target detection model, the target detection model identifies the first target corresponding to the visible light fusion features.

6. A multi-level fusion-based UAV heterogeneous image target detection system, characterized in that, The system includes: The acquisition module is used to acquire visible light image sequences and infrared light image sequences; the visible light image sequence includes multiple frames of visible light images captured in consecutive time; the infrared light image sequence contains multiple frames of infrared light images captured in consecutive time. The feature extraction module is used to extract multiple visible light features from visible light image sequences, fuse these features to obtain the visible light fusion features of the target; and to extract multiple infrared light features from infrared light image sequences, fuse these features to obtain the infrared light fusion features of the target. The confidence fusion module is used to fuse the confidence of the multiple visible light features based on the DS evidence theory to obtain a first confidence level; and to fuse the confidence of the multiple infrared light features based on the DS evidence theory to obtain a second confidence level. The target detection module is used to perform target detection based on the visible light fusion features to obtain a first detected target; and to perform target detection based on the infrared light fusion features to obtain a second detected target. If the first and second detected targets are the same target, the visible light fusion features and the infrared light fusion features are fused to obtain a target fusion feature. A confidence analysis is performed on the target fusion feature to obtain a third confidence level. If the third confidence level is greater than a preset value, the detected first and second detected targets are determined to be true targets. If the first and second detected targets are not the same target, and if the first detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size; if the second detected target is a specific target, the value of the second confidence level is lowered according to an adaptive step size. The first confidence level should be lowered by adjusting the step size. Based on the lowered first or second confidence level, it is determined whether the specific target can be detected simultaneously based on visible light fusion features and infrared light fusion features. If the specific target is detected simultaneously, the current fusion confidence level is obtained based on the current values ​​of the first and second confidence levels. If the current fusion confidence level is greater than a preset value, the detected specific target is determined to be a true target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is less than a set threshold, the detected target is determined to be a false target. If the specific target is not detected simultaneously, and the confidence level of the sensor that did not detect the specific target is greater than or equal to a set threshold, the above steps are performed.

7. The UAV heterogeneous image target detection system based on multi-level fusion according to claim 6, characterized in that, Based on the current value of the first confidence level and the current value of the second confidence level, the current fusion confidence level is obtained, including: The current fusion confidence is obtained by weighting and summing the current values ​​of the first confidence level and the current values ​​of the second confidence level.

8. The UAV heterogeneous image target detection system based on multi-level fusion according to claim 6, characterized in that, Visible light image sequences possess various visible light features, including texture features, edge features, and color features; The process of fusing multiple visible light features to obtain the visible light fused features of the target includes: Texture features, color features, and edge features are weighted and fused according to the correspondence of pixel positions to obtain visible light fused features.

9. The UAV heterogeneous image target detection system based on multi-level fusion according to claim 6, characterized in that, The visible light image sequence also includes multiple visible light features such as multi-frame motion features, histogram features, multi-scale features, aspect ratio features, and spatial scene analysis features.

10. The UAV heterogeneous image target detection system based on multi-level fusion according to claim 6, characterized in that, Target detection is performed based on the visible light fusion features to obtain a first detection target, including: When the visible light fusion features are input into the target detection model, the target detection model identifies the first target corresponding to the visible light fusion features.

Citation Information

Patent Citations

  • Target detection method based on visible light and infrared image fusion

    CN113763356A

  • Multi-source remote sensing image fusion target comprehensive detection method

    CN113963240A