Contrastive learning based multi-modal object detection method
Patent Information
- Application Number
- CN202410300689.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-03-15
AI Technical Summary
[0008]本发明的目的在于解决上述现有的技术存在的不足,提出一种基于对比学习的多模态目标检测方法,用于解决现有技术存在的检测精度和效率较低的技术问题
[0022]1.本发明在对基于对比学习的多模态目标检测网络进行训练的过程中,通过多模态目标检测网络和对比学习网络的学习结果对多模态目标检测网络的参数进行更新,能够拉近RGB图像表征与对应IR图像表征之间的距离,并推远RGB图像表征与其他IR图像表征之间的距离,使特征提取器特别关注RGB图像和IR图像的深层特征和信息融合,不但增大了两种模态之间的互信息,还提高多模态目标检测网络的鲁棒性和准确性,避免了现有技术对深层特征训练不足和语义信息提取不足的缺陷,有效提高了目标检测精度。
Smart Images

Figure CN118154844B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology and relates to a multimodal target detection method, specifically a multimodal target detection method based on contrastive learning, which can be used in fields such as autonomous driving and traffic safety. Background Technology
[0002] The primary goal of object detection is to identify the category and location of objects in an image. This is a core problem in image processing technology and is widely used in everyday applications such as pedestrian and vehicle detection. When detecting pedestrians and vehicles on roads, traditional visible light RGB sensors cannot consistently provide effective input information to the object detection network due to variations in weather and lighting conditions. For example, in adverse environments such as nighttime, heavy fog, and heavy rain, visible light RGB sensors may fail, making it difficult to accurately capture the boundary between the background and the target. Therefore, some research has introduced near-infrared (IR) images, which are unaffected by obstructions and low light, to enhance the input information. This can compensate for the information lost in visible light images and better achieve the object detection task. RGB and IR images have significantly different wavelength ranges and information content, representing different modalities of data. This type of object detection task, using data from multiple different modalities as input, is called multimodal object detection.
[0003] Multimodal target detection can effectively fuse information from different modal images by combining the rich color and texture information provided by RGB images under good lighting conditions with the thermal radiation contour information captured by IR images, thereby significantly improving detection performance.
[0004] Based on the stage in which the fusion module is located, multimodal object detection can be divided into detection methods based on early-stage fusion, mid-stage fusion, and late-stage fusion. Early-stage fusion-based methods involve pixel-level fusion of RGB and IR images before feature extraction; however, they suffer from severe information redundancy and cannot fully utilize the complementary information between different modalities. Late-stage fusion-based methods involve integrating the final results after feature extraction for different modalities using voting mechanisms or weighting; however, late-stage fusion essentially ignores the information interaction between modalities, resulting in limited improvement in object detection accuracy.
[0005] Mid-stage fusion-based detection methods extract and fuse features using convolutional neural networks. The fused features are then fed into other network layers for further learning. This fusion strategy offers better performance than both early and late-stage methods and has greater potential for accuracy improvement, showing significant promise in multimodal object detection. This approach typically involves building two parallel detection networks to extract features for each modality separately. The model structure of each detection network is similar to common object detection models, such as the YOLO series and R-CNN series.
[0006] To improve the accuracy of object detection, for example, the paper "Cross-Modality Fusion Transformer for Multispectral Object Detection" published by Fang Qingyun et al. in 2022 proposed a Transformer-based multimodal object detection method. This method first constructs two parallel object detection networks to extract features for each modality, and then builds a Transformer-based feature fusion module "CFT" between the two parallel networks. RGB and IR images from the training sample set are input into the corresponding object detection networks to obtain concatenated RGB and IR intermediate layer features. The CFT fusion module fuses the RGB and IR image features, and the detection head performs object detection on the fused features. The network is then iterated using an object detection loss function. Finally, the trained network model is used to obtain the object detection results for the test sample set.
[0007] This method leverages the self-attention capability of the Transformer module. During the feature extraction stage, its network integrates global contextual information, enabling robust fusion training between RGB and IR modalities. This robustly captures and utilizes the fusion information between RGB and IR images, helping the detection head to more accurately detect target categories and locations, thus improving accuracy. However, its shortcomings lie in its use of the Transformer's multi-head attention mechanism to mine dependencies between pixel samples within a single image, focusing particularly on contextual aggregation models while neglecting the training of deep features between images and the extraction of abstract semantic information. This results in relatively low target detection accuracy. Furthermore, when detecting test samples, the intermediate layer features of each sample still require complex CFT fusion, leading to excessively long detection times and low detection efficiency. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of the existing technologies mentioned above, and to propose a multimodal target detection method based on contrastive learning, which solves the technical problems of low detection accuracy and efficiency in the existing technologies.
[0009] To achieve the above objectives, the technical solution adopted by the present invention includes the following steps:
[0010] (1) Obtain the training sample set and the test sample set:
[0011] Acquire N visible light RGB images and N corresponding near-infrared IR images representing multiple target types. The nth visible light RGB image and its corresponding near-infrared IR image represent the same scene. Data augmentation is performed on M of these visible light RGB images and their corresponding near-infrared IR images, and the targets are then labeled. The M augmented visible light RGB images and their corresponding near-infrared IR images, along with their labels, form the training sample set. Simultaneously, the remaining R = NM augmented visible light RGB images and their corresponding near-infrared IR images form the test sample set, where N ≥ 10000, M > N / 2, and the mth augmented RGB image and its corresponding IR image are respectively... and
[0012] (2) Construct a multimodal target detection network model W based on contrastive learning:
[0013] Construct a cascaded multimodal target detection network W OD And contrastive learning network W CL The target detection network model W, where W OD It includes a first feature extractor and a second feature extractor with identical structures arranged in parallel, and a multi-scale detection head connected between the two feature extractors; W CL This includes a first multilayer perceptron MLP and a second multilayer perceptron MLP that are cascaded with the first feature extractor and the second feature extractor, respectively, and have the same structure.
[0014] (3) Define the loss function Loss of the multimodal object detection network model W based on contrastive learning:
[0015] Loss = Loss OD +λLoss CL
[0016] Among them, Loss OD Loss CL These represent the multimodal object detection network W. OD The loss value, contrastive learning network W CL The loss value, λ represents the loss. CL The weights;
[0017] (4) Iteratively train the object detection network model W:
[0018] The object detection network model W is iteratively trained using the training sample set to obtain the trained network model. * ;
[0019] (5) Obtain the target detection results:
[0020] The test sample set is used as the trained object detection network model W. * Input, multimodal target detection network The target in each test sample is detected, resulting in R detection results.
[0021] Compared with existing technologies, the present invention has the following advantages:
[0022] 1. In the process of training a multimodal object detection network based on contrastive learning, this invention updates the parameters of the multimodal object detection network by using the learning results of the multimodal object detection network and the contrastive learning network. This can narrow the distance between RGB image representations and their corresponding IR image representations, and widen the distance between RGB image representations and other IR image representations. This allows the feature extractor to pay special attention to the deep features and information fusion of RGB and IR images, which not only increases the mutual information between the two modalities, but also improves the robustness and accuracy of the multimodal object detection network. This avoids the shortcomings of existing technologies in terms of insufficient training of deep features and insufficient extraction of semantic information, and effectively improves the accuracy of object detection.
[0023] 2. The multimodal target detection network of the present invention uses contrastive learning as an auxiliary task during training, without the need to introduce a complex feature fusion module. While effectively enhancing the performance of the target detection network, the contrastive learning network does not participate in the detection process of test samples, thus avoiding the time-consuming feature fusion in existing technologies when detecting targets. This not only improves the target detection accuracy but also the target detection efficiency. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating the implementation of the present invention. Detailed Implementation
[0025] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0026] Reference Figure 1 The present invention includes the following steps:
[0027] Step 1) Obtain the training sample set and the test sample set:
[0028] Acquire N visible light RGB images and N corresponding near-infrared IR images representing multiple target types. The nth visible light RGB image and its corresponding near-infrared IR image represent the same scene. Data augmentation is performed on M of these visible light RGB images and their corresponding near-infrared IR images, and the targets are then labeled. The M augmented visible light RGB images and their corresponding near-infrared IR images, along with their labels, form the training sample set. Simultaneously, the remaining R = NM augmented visible light RGB images and their corresponding near-infrared IR images form the test sample set, where N ≥ 10000, M > N / 2, and the mth augmented RGB image and its corresponding IR image are respectively... and
[0029] In this embodiment, the RGB and IR images are real road images captured by a visible light camera and a FLIR BlackFly thermal imaging camera, respectively, and both images are 640×512×3 pixels in size. Each visible light RGB image and its corresponding near-infrared IR image are randomly scaled and cropped. Then, the scaled and cropped RGB and IR images are randomly rotated and flipped. Next, the rotated and flipped RGB and IR images undergo random brightness and contrast adjustments. Finally, random noise is added to the adjusted RGB and IR images to achieve data enhancement for each visible light RGB image and its corresponding near-infrared IR image, resulting in the data-enhanced m-th visible light RGB image. and its corresponding near-infrared (IR) image Data augmentation increases the diversity of data, allowing models to learn target features under different perspectives, scales, and lighting conditions. This reduces the risk of overfitting, improves the generalization and robustness of the target detection model, and thus enhances detection performance.
[0030] Step 2) Construct a multimodal object detection network model W based on contrastive learning:
[0031] Construct a cascaded multimodal target detection network W OD And contrastive learning network W CL The target detection network model W, where W OD It includes a first feature extractor and a second feature extractor with identical structures arranged in parallel, and a multi-scale detection head connected between the two feature extractors; W CL This includes a first multilayer perceptron MLP and a second multilayer perceptron MLP that are cascaded with the first feature extractor and the second feature extractor, respectively, and have the same structure.
[0032] Both the first and second feature extractors include multiple cascaded convolutional modules; the multi-scale detection head includes three parallel fusion modules and detection heads cascaded with the three fusion modules respectively; both the first and second multilayer perceptrons (MLPs) include two stacked fully connected layers and a normalization layer, as well as a ReLU activation function loaded between the two fully connected layers.
[0033] Step 3) Define the loss function Loss for the multimodal object detection network model W based on contrastive learning:
[0034] Loss = Loss OD +λLoss CL
[0035] Among them, Loss OD Loss CL These represent the multimodal object detection network W. OD The loss value, contrastive learning network W CL The loss value, λ represents the loss. CL The weights are λ = 0.1 in this embodiment. The parameters of the multimodal object detection network are updated jointly by the learning results of the multimodal object detection network and the contrastive learning network. This can narrow the distance between the RGB image representation and the corresponding IR image representation, and widen the distance between the RGB image representation and other IR image representations.
[0036] Step 4) Iteratively train the object detection network model W:
[0037] The object detection network model W is iteratively trained using the training sample set to obtain the trained network model. * ;
[0038] (4a) Initialize the number of iterations to t, the maximum number of iterations to T, T≥100, and the current network model W t The weight parameter in is ω t Initialize a vector queue containing K random vectors as a Queue, where K > M, and let t = 1; in this embodiment, K = 65536.
[0039] (4b)W OD The first and second feature extractors in the process respectively target... Multi-scale feature extraction is performed to obtain Corresponding multi-size feature sets The three fusion modules of the multi-scale detection head respectively process each and its corresponding Each and its corresponding Each and its corresponding Perform feature fusion to obtain a fused feature set. The three detectors in the multi-scale detector head respectively detect... Perform object detection to obtain detection results for M training samples, where, They represent The shallow, medium and deep features, They represent shallow, medium and deep features
[0040] (4c)W CL The first and second multilayer perceptron MLPs in the model respectively handle deep features. After performing nonlinear mapping and normalization, we obtain Corresponding RGB image representation vector IR image representation vector The queue is updated by replacing the same number of IR image representation vectors in the queue with the obtained M IR image representation vectors, and then the calculation is performed. and In the updated queue, except Similarity of the other K-1 vectors
[0041] and In the updated queue, except Similarity of the other K-1 vectors The calculation formulas are as follows:
[0042]
[0043]
[0044]
[0045]
[0046] Where τ represents the temperature hyperparameter, and These represent the mappings of two fully connected layers in the first multilayer perceptron. and Let τ and τ represent the mappings of the two fully connected layers of the second multilayer perceptron, respectively. ReLu(·) represents the nonlinear activation function, and Norm(·) represents the normalization operation. In this example, τ = 0.07.
[0047] (4d) Using a multimodal target detection network WOD Loss value OD And contrastive learning network W CL Loss value CL Calculate the loss value Loss of the network model W, and adjust the weight parameters ω based on the loss. t The model is updated to obtain the target detection network model W after this iteration;
[0048] Multimodal target detection network W OD Loss value OD And contrastive learning network W CL Loss value CL The loss value (Loss) of network model W is calculated using the following formulas:
[0049] Loss OD =λ1L conf +λ2L cls +λ3L box
[0050]
[0051] Among them, L conf W OD The squared error loss between the detected target confidence result and the true confidence, L cls W OD The binary cross-entropy loss between the detected target class and the true class, L box W OD The perfect intersection-union (CIU) loss between the detected bounding box locations and the ground truth bounding box locations, where λ1, λ2, and λ3 represent L... conf L cls L box The weights are λ = 0.1, λ1 = 1, λ2 = 0.5, and λ3 = 0.05.
[0052] The loss function of the contrastive learning network can narrow the distance between RGB image representations and their corresponding IR image representations, and widen the distance between RGB image representations and other IR image representations. This allows the feature extractor of the multimodal object detection network to pay special attention to the deep features and information fusion of RGB and IR images. This not only increases the mutual information between the two modalities, but also improves the robustness and accuracy of the multimodal object detection network. It avoids the shortcomings of existing technologies in terms of insufficient training of deep features and insufficient extraction of semantic information, and effectively improves the accuracy of object detection.
[0053] For network model W t The weight parameter ω in t The update is performed using the following formula:
[0054]
[0055] Where η represents the pre-set gradient descent parameter, ω t+1 Represents ω t The update results This indicates the partial derivative operation.
[0056] (4e) Determine whether t = T holds true. If so, obtain the trained object detection network model W. * Among them, the trained multimodal object detection network model W OD * Otherwise, let t = t + 1, W t =W, and execute step (4b).
[0057] Step 5) Obtain the target detection results:
[0058] The test sample set is used as the trained object detection network model W. * Input, multimodal target detection network For each test sample, targets are detected, resulting in R detection results. The contrastive learning network serves as an auxiliary task to the multimodal target detection network during training. While effectively enhancing the performance of the target detection network, it does not participate in the detection of test samples, avoiding the significant time spent on feature fusion during target detection. This not only improves target detection accuracy but also increases its efficiency.
[0059] The effects of the present invention will be further explained below with reference to simulation experiments.
[0060] 1. Simulation conditions:
[0061] The hardware platform for the simulation experiment is: a Xeon(R) Platinum 8255C CPU with a clock speed of 2.50GHz*96, 256GB of memory, and an NVIDIA A100 SXM4 80GB graphics processor. The software platform is: Ubuntu 18.04.5 operating system, Python 3.8.16, and PyTorch 1.12.1.
[0062] 2. Simulation content and result analysis:
[0063] The detection accuracy and efficiency of the present invention and the prior art were compared by simulation, and the results are shown in Table 1.
[0064] Table 1
[0065] Existing technology 77.6% 91.2ms This invention 78.6% 20.7ms
[0066] The detection time is the total time for detecting the test sample set divided by the number of test samples, i.e., the average time for detecting one sample; the detection accuracy is calculated using the following formula for the average precision mAP50:
[0067]
[0068]
[0069] Y represents all target categories; TP represents true positives, where the detector's predicted bounding box and the true label satisfy the intersection-union (IU) threshold; FP represents false positives, where the detector's predicted bounding box is not a true target; FN represents false negatives, where a true target was not detected. mAP50 is the average accuracy (AP) for all categories when the IU is 0.50.
[0070] As can be seen from Table 1, compared with the prior art, the detection accuracy of the present invention is significantly improved, while the detection time is significantly reduced and the detection efficiency is significantly improved.
Claims
1. A multimodal target detection method based on contrastive learning, characterized in that, The steps include the following: (1) Obtain the training sample set and the test sample set: Get multiple target types Visible light RGB images and their corresponding The first near-infrared IR image, Two visible light RGB images and their corresponding near-infrared IR images depict the same scene, and the images are compared to each other. After data augmentation of two visible light RGB images and their corresponding near-infrared IR images, the target is labeled, and then the data-augmented images are... The training sample set consists of 10 visible light RGB images and their corresponding near-infrared IR images and their labels, while the remaining data is augmented. The test sample set consists of 10 visible light RGB images and their corresponding near-infrared IR images, among which... , , No. The augmented RGB image and its corresponding IR image are respectively and ; (2) Construct a multimodal target detection network model based on contrastive learning : Constructing a cascaded multimodal target detection network and contrastive learning networks Object detection network model ,in The first feature extractor and the second feature extractor are composed of three cascaded convolutional modules arranged in parallel. Each of the convolutional modules at the same position in the first feature extractor and the second feature extractor is connected to a multi-scale detection head composed of a cascaded fusion module and a detection head. This includes a first multilayer perceptron MLP and a second multilayer perceptron MLP that are cascaded with the first feature extractor and the second feature extractor, respectively, and have the same structure. (3) Define a multimodal target detection network model based on contrastive learning. loss function : ; in, , These represent multimodal object detection networks. Loss value, contrastive learning network The loss value, express The weights; (4) Target detection network model Perform iterative training: The target detection network model was trained using the sample set. Perform iterative training to obtain a trained object detection network model. ; (5) Obtain target detection results: The test sample set is used as the trained object detection network model. Input, multimodal target detection network Detect the target in each test sample and obtain One test result.
2. The method according to claim 1, characterized in that, The step (1) described above is about aligning with... Data augmentation is performed on two visible light RGB images and their corresponding near-infrared IR images. The steps are as follows: For each visible light RGB image and its corresponding near-infrared IR image, random scaling and cropping are performed. Then, the randomly scaled and cropped RGB and IR images are randomly rotated and flipped. Next, random brightness and contrast adjustments are performed on the randomly rotated and flipped RGB and IR images. Finally, random noise is added to the RGB and IR images after random brightness and contrast adjustments. This process achieves data augmentation for each visible light RGB image and its corresponding near-infrared IR image, resulting in the data-augmented [image / description]. Visible light RGB image and its corresponding near-infrared (IR) image .
3. The method according to claim 1, characterized in that, The target detection network model described in step (2) ,in: Both the first and second multilayer perceptrons (MLPs) consist of two stacked fully connected layers and a normalization layer, as well as a ReLU activation function loaded between the two fully connected layers.
4. The method according to claim 1, characterized in that, The step (4) involves training the target detection network model using a training sample set. The iterative training process involves the following steps: (4a) Initialize the number of iterations to be The maximum number of iterations is , Current network model The weight parameters in are Initial storage has A vector queue of random vectors is , , and order ; (4b) The first and second feature extractors in the process respectively target... , Multi-scale feature extraction is performed to obtain , Corresponding multi-size feature sets , The three fusion modules of the multi-scale detection head respectively process each... and its corresponding Each and its corresponding Each and its corresponding Perform feature fusion to obtain a fused feature set. The three detectors in the multi-scale detector head respectively detect... , , Perform target detection and obtain The detection results of each training sample, among which, , , They represent The shallow, medium and deep features, , , They represent The shallow, medium and deep features, , , ; (4c) The first and second multilayer perceptron MLPs in the model respectively handle deep features. , After performing nonlinear mapping and normalization, we obtain , Corresponding RGB image representation vector IR image representation vector and through the obtained IR image representation vector pairs Replace the same number of IR image representation vectors to achieve [the desired result]. Update, then calculate and In the updated queue, except Other Similarity of vectors , ; (4d) Through a multimodal target detection network loss value and contrastive learning networks loss value Computational network model loss value and through For weight parameters The model is updated to obtain the object detection network model after this iteration. ; (4e) Judgment If true, then a well-trained object detection network model is obtained. Among them, the trained multimodal object detection network model Otherwise, let , Then proceed to step (4b).
5. The method according to claim 4, characterized in that, The steps described in step (4c) and In the updated queue, except Other Similarity of vectors , The calculation formulas are as follows: ; ; ; ; in, This indicates temperature hyperparameters. and These represent the mappings of two fully connected layers in the first multilayer perceptron. and These represent the mappings of the two fully connected layers of the second multilayer perceptron. Represents a non-linear activation function. This indicates a normalization operation.
6. The method according to claim 4, characterized in that, The multimodal target detection network described in step (4d) loss value and contrastive learning networks loss value Network Model loss value The calculation formulas are as follows: ; ; in, express The squared error loss between the detected target confidence score and the true confidence score. express The binary cross-entropy loss between the detected target class and the true class. express The cross-union loss between the detected bounding box locations and the ground truth bounding box locations. , , They represent , , The weight.
7. The method according to claim 4, characterized in that, The network model described in step (4d) weight parameters in The update is performed using the following formula: ; in, This represents the pre-set gradient descent parameters. express The update results This indicates the partial derivative operation.
Citation Information
Patent Citations
Infrared image and depth image bimodal target segmentation method and device
CN112115864A
Salient target detection method based on RGB-T multi-source image data
CN114898106A