Cross-modal contrastive learning method and system for rgb-d image dense prediction task

By constructing an RGB-D cross-modal self-supervised framework and utilizing local-global coupling modules and a cross-modal training paradigm, the data annotation challenge in dense RGB-D prediction tasks is solved, achieving efficient multi-scale feature learning and improving the accuracy and local detail capture capabilities of downstream tasks.

CN116434033BActive Publication Date: 2025-11-11SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310268449.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-20
Publication Date
2025-11-11
Estimated Expiration
2043-03-20

AI Technical Summary

Technical Problem

Existing technologies face challenges in dense RGB-D prediction tasks, including difficulties in data annotation, scarcity of labeled data, and the RGB-D cross-modal gap. Existing multimodal contrastive learning methods struggle to effectively extract local cues and detailed features.

Method used

An RGB-D cross-modal self-supervised framework is constructed, using the ResNet50 network structure as the encoder. It is pre-trained through a local-global coupling module and a cross-modal training paradigm, and multi-scale features are learned by utilizing spatially aware cross-modal contrast loss. By combining self-supervised learning and supervised training, the parameters are effectively transferred.

Benefits of technology

It effectively overcomes the problem of insufficient data and improves the accuracy of downstream tasks, especially in RGB-D salient object detection and semantic segmentation, where it enhances the ability to capture local details and outperforms traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116434033B_ABST
    Figure CN116434033B_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal contrast learning method and system for an RGB-D image dense prediction task, constructs an RGB-D cross-modal self-supervision framework to pre-train an encoder, inputs pre-trained encoder parameters into a network model of a downstream RGB-D dense prediction task, performs supervised training on the network model, obtains the trained network model of the downstream RGB-D dense prediction task, and completes inference and output of a prediction result; the method overcomes the problem of insufficient data, fills the gap of RGB-D cross-modal data, can extract multi-scale modal specific clues and heterogeneous cross-modal correlations through the pre-training method, and promotes multi-modal fusion of a downstream task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of computer vision and self-supervised contrastive learning technology, and particularly relates to RGB-D cross-modal semantic segmentation and salient object detection technology. It mainly involves a cross-modal contrastive learning method and system for dense prediction tasks of RGB-D images. Background Technology

[0002] Dense prediction is a task that classifies and predicts each pixel of an image. It is a fundamental field of computer vision that includes many visual tasks, such as salient object detection, which captures salient visual regions, and semantic segmentation, which classifies each pixel of an image scene.

[0003] In recent years, the development of depth sensors has brought additional stable geometry and contextual cues to traditional RGB-based computer vision systems. The resulting multimodal vision systems possess the complementarity of the two modalities, and joint inference significantly improves their accuracy and robustness. Given the powerful feature learning capabilities and the great success of deep learning tools, various methods for dense RGB-D prediction tasks based on convolutional neural networks (CNNs) have been proposed. To fully integrate the multi-scale cross-modal aspects of RGB-D pairs, many existing models are typically equipped with multiple cross-modal cross-level fusion paths and modules. This design introduces significant complexity to the model; a large number of parameters often require large-scale training data to ensure effectiveness, which presents a significant challenge because collecting multimodal data and labeling dense pixel-level tags is both expensive and labor-intensive. Previous work avoided the scarcity of labeled multimodal data by borrowing pre-trained weights from ImageNet as appropriate initializations for all modalities. However, the domain gap between ImageNet and dense prediction datasets, as well as the modal gap between RGB and depth, often leads to biased initialization and subsequent sub-optimization.

[0004] The rapid development of self-supervised learning (SSL) offers new possibilities for directly overcoming the data shortage problem in multimodal dense prediction. As one of the most promising directions in SSL, contrastive learning (aiming to learn invariant high-level features in image transformations) has been widely applied in multiple fields and has made significant progress in classification tasks. Most existing contrastive learning methods follow the instance recognition paradigm, classifying transformed versions of the input as images from the same source. This concept has been widely adopted and applied to cross-modal recognition of multimodal data, such as speech, video, text, and RGB-D images. Existing multimodal contrastive learning solutions focus on learning high-level global representations but have limited ability to extract local cues to infer details. Summary of the Invention

[0005] This invention addresses the problems of scarce labeled data in the RGB-D domain and insufficient targeted design for RGB-D cross-modal gaps in existing technologies. It provides a cross-modal contrastive learning method and system for dense RGB-D image prediction tasks. The method constructs an RGB-D cross-modal self-supervised framework to pre-train the encoder, and then inputs the pre-trained encoder parameters into the network model for the downstream RGB-D dense prediction task. The network model is then trained in a supervised manner to obtain the trained network model for the downstream RGB-D dense prediction task, which completes inference and outputs the prediction results. This invention overcomes the problem of insufficient data and bridges the data gap in RGB-D cross-modal data. The pre-training method of this invention can extract multi-scale modal-specific cues and heterogeneous cross-modal correlations, thereby promoting multi-modal fusion in downstream tasks.

[0006] To achieve the above objectives, the technical solution adopted by this invention is: a cross-modal contrastive learning method for dense prediction tasks of RGB-D images, which constructs an RGB-D cross-modal self-supervised framework to pre-train the encoder, inputs the pre-trained encoder parameters into the network model of the downstream RGB-D dense prediction task, performs supervised training on the network model, obtains the trained network model of the downstream RGB-D dense prediction task, and completes inference to output the prediction result;

[0007] The RGB-D cross-modal self-supervised framework includes at least a local-global coupling module and a cross-modal training paradigm. The local-global coupling module contains a spatially aware cross-modal contrast loss for a multi-scale projector, which can simultaneously learn features with global understanding, local context, and multi-scale cross-modal correlation during the pre-training phase. The cross-modal training paradigm uses only cross-modal consistency as the multi-modal contrast loss.

[0008] As an improvement to the present invention, this method specifically includes the following steps:

[0009] S1, Dataset Acquisition: Acquire RGB-D cross-modal image datasets, which are used for pre-training the encoder in the RGB-D cross-modal self-supervised framework and training the network model for the downstream RGB-D dense prediction task, respectively.

[0010] S2, Encoder selection: ResNet50 network structure is adopted as the encoder;

[0011] S3, Construct a self-supervised framework: Construct an RGB-D cross-modal self-supervised framework to pre-train the encoder selected in step S2. The pre-training provides RGB-D domain-specific model initialization parameters for downstream RGB-D dense prediction tasks.

[0012] S4, Network model training for the prediction task: Input the self-supervised pre-trained encoder parameters obtained in step S3 into the network model of the downstream RGB-D dense prediction task, and perform supervised training on the network model.

[0013] S5, Prediction Result Output: Based on different RGB-D tasks, the downstream RGB-D dense prediction task network model trained in step S4 is used to complete the inference and output the prediction result.

[0014] As an improvement of the present invention, in the data set of step S1, different RGB-D unlabeled image datasets are collected as training sets for the pre-training of the encoder by the RGB-D cross-modal self-supervised framework; and the publicly labeled datasets of the corresponding tasks are used as training sets for supervised learning of the downstream RGB-D dense prediction task.

[0015] As another improvement of the present invention, step S3 specifically includes:

[0016] S31: Each training sample image x will be transformed by two different data augmentation operations t′, t″∈Γ to obtain x′, x″, thereby generating different perspectives of the same sample, where Γ represents the data augmentation operations of random cropping and splicing, rotation and scaling and flipping;

[0017] S32: Each pair of RGB and Depth paired images is then input into the corresponding encoder after undergoing the same data augmentation process as in step S31. θ With momentum encoder e ξ Encode: f′=e θ (t′(x′)), f″=e ξ (t″(x″)), where f′ and f″ are the feature maps generated after encoding;

[0018] S33: Feature extraction is performed on the feature map generated in step S32 using a spatially aware local-global coupling module: f is used to obtain the global feature map y through a global pooling layer, and local feature maps F1 and F2 are obtained through two local pooling layers. The three feature maps y, F1, and F2 are unfolded into vectors, and the vectors are merged to make them a one-dimensional representation y containing multi-scale features. * =concat(y, F1, F2);

[0019] S34: Use one-dimensional vectors generated from different modalities for cross-modal contrastive learning pre-training, with the total loss function being the cross-modal loss function:

[0020]

[0021] In the formula, q r As anchor samples for updating the RGB modality encoder parameters, qd The anchor samples are used to update the parameters of the Depth mode encoder. and These are negative sample pools for two different modalities.

[0022] As another improvement of the present invention, in step S32, e ξ The moving average method is used for parameter updates, and the formula is defined as follows:

[0023]

[0024] Where ξ is e ξ The parameter set, For e θ The parameter set.

[0025] As another improvement of the present invention, the loss function in step S34 is a contrastive loss function, specifically defined as follows:

[0026]

[0027] In the formula, q and k + They are positive sample pairs, q and k - Negative sample pairs τ represents the negative sample vector of the negative sample pool, and τ represents the temperature coefficient.

[0028] To achieve the above objectives, the present invention also adopts the following technical solution: a cross-modal contrastive learning system for dense prediction tasks of RGB-D images, comprising a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the methods described above.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] (1) This invention adopts the pre-training paradigm of self-supervised learning. Compared with supervised learning, self-supervised learning uses unlabeled RGB-D datasets for pre-training, which overcomes the problem that the labeling of dense RGB-D prediction tasks is difficult and labeled data is scarce.

[0031] (2) This invention utilizes a customized local-global coupled contrast loss to fully leverage the spatial alignment and semantic consistency between RGB and Depth data to learn multi-scale-specific representations and cross-modal correlations, which efficiently meets the characteristics of dense prediction tasks and effectively improves the accuracy of downstream tasks.

[0032] (3) This invention is a contrastive learning framework tailored for multimodal dense prediction tasks. In RGB-D salient object detection and semantic segmentation, it is far superior to traditional contrastive learning benchmarks and has a better ability to capture local details. Attached Figure Description

[0033] Figure 1 This is a flowchart of the steps of the method of the present invention;

[0034] Figure 2 This is a schematic diagram of the RGB-D cross-modal self-supervised framework of the present invention;

[0035] Figure 3 This is a downstream task effect diagram in Embodiment 1 of the present invention. Detailed Implementation

[0036] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.

[0037] Example 1

[0038] Cross-modal contrastive learning methods for dense prediction tasks of RGB-D images, such as Figure 1 As shown, it includes the following steps:

[0039] Step S1: Obtain RGB-D cross-modal image datasets, which are used for pre-training the encoder and training the network model for the downstream RGB-D dense prediction task by the RGB-D cross-modal self-supervised framework. For different RGB-D dense prediction visual tasks, different unlabeled RGB-D image datasets are collected as the training set for the pre-training framework. For different RGB-D dense prediction visual tasks, the publicly labeled datasets of the corresponding tasks are used as the training set for supervised learning of the downstream tasks.

[0040] The salient object detection experiment in this embodiment includes 10 publicly available datasets, including NJU2K, NLPR, STERE, SIP, DES, LFSD, DUT, SSD, REDWEB-S, and COME-15K, totaling 25,233 RGB-D paired samples as a self-supervised pre-training dataset. In the downstream task stage, we use 1,485 and 700 images extracted from NJU2K and NLPR, respectively, as the training set for this example, while the test set is the test set from all the remaining datasets.

[0041] Step S2: Use the ResNet50 network structure as the encoder. This invention uses the MoCo contrastive learning self-supervised framework as the baseline framework. Considering the asymmetric structure in the MoCo framework, it is necessary to set up... Figure 2 The two encoders shown have corresponding momentum encoders, with RGB and Depth as anchor samples respectively, and all encoders are set to ResNet-50.

[0042] Contrastive learning aims to learn invariant high-level features in transformations of similar samples. Existing contrastive learning methods follow the instance recognition paradigm and are widely adopted for cross-modal recognition of multimodal data. However, cross-modal recognition solutions focus on learning high-level global representations, making it difficult to achieve the detailed learning requirements of pixel-level dense features. This paper proposes a cross-modal contrastive learning framework for RGB-D dense prediction visual tasks, with the overall structure as follows: Figure 2 As shown, it mainly includes two major designs: a spatially aware local-global coupling module and a cross-modal training paradigm.

[0043] 1) The spatially aware local-global coupling module includes a multi-scale projector spatially aware cross-modal contrast loss, enabling the model to fully utilize the natural pixel pairing information in RGB-D pairs during contrast learning. The self-supervised pre-training phase can simultaneously learn features with global understanding, local context, and multi-scale cross-modal relevance.

[0044] 2) The cross-modal training paradigm is the foundation of this invention's framework. This paradigm uses only cross-modal consistency as the multimodal contrastive loss. Unlike the hybrid contrastive paradigm, which may encounter conflicts between intramodal and cross-modal optimization objectives, this contrastive learning process effectively balances both intramodal features and cross-modal correlations.

[0045] Step S3: Construct a self-supervised framework: Construct an RGB-D cross-modal self-supervised framework to pre-train the encoder selected in step S2. The pre-training process specifically includes:

[0046] Step S31: Each training sample image x is transformed by two different data augmentation operations t′, t″∈Γ to obtain x′, x″, thereby generating different perspectives of the same sample. Γ represents the data augmentation operations of random cropping and splicing, rotation and scaling and flipping.

[0047] Step S32: Each pair of RGB and Depth paired images is then input into the corresponding encoder e after undergoing the same data augmentation process as in step 3.1. θ With momentum encoder e ξ Encode: f′=e θ (t′(x′)), f″=e ξ (t″(x″)), where f′ and f″ are the feature maps generated after encoding;

[0048] Momentum encoder e ξ The parameter update rules and encoder e θ Different, e θ The loss function described above is used for updating, and e ξ To maintain the stability of the negative sample pool's sample distribution, a moving average method is used for parameter updates, defined as follows:

[0049]

[0050] Where ξ is e ξ The parameter set, For e θ The parameter set.

[0051] Step S33: Feature extraction is performed on the feature map generated in step S32 using a spatially aware local-global coupling module: f is used to obtain a global feature map y through a global pooling layer, and local feature maps F1 and F2 are obtained through two local pooling layers. Then, the three feature maps y, F1, and F2 are unfolded into vectors, and finally these vectors are merged to form a one-dimensional representation y containing multi-scale features. * =concat(y, F1, F2);

[0052] Step S34: The encoders for RGB and Depth modalities need to use images of their own modalities as anchor samples q to calculate the loss function in order to update the corresponding modal encoders. Cross-modal comparative learning pre-training is performed using one-dimensional vectors generated by different modalities. The total loss function is the cross-modal loss function:

[0053]

[0054] In the formula, q r As anchor samples for updating the RGB modality encoder parameters, q d The anchor samples are used to update the parameters of the Depth mode encoder. and These are negative sample pools for two different modalities.

[0055] The basic form of the loss function is the contrastive loss function, which is defined as follows:

[0056]

[0057] In the formula, q and k + They are positive sample pairs, q and k - Negative sample pairs This represents the negative sample vector of the negative sample pool, and τ represents the temperature coefficient, used to adjust the degree of influence of the loss function.

[0058] In this embodiment, the parameters of the pre-training framework are set as follows: number of negative sample pool samples. The value is 65536, and the momentum coefficient m of the momentum encoder is 0.999. The pre-training dataset of this invention is much smaller than ImageNet, with 200 training epochs and 8 mini-batch sample inputs on a single GPU, and the temperature hyperparameter τ is set to 0.05.

[0059] Step S4: Supervised pre-training for the RGB-D salient object detection task: Extract the parameters from the RGB and Depth encoders in Step S3 and input them into the encoder of the downstream test network as initialization parameters for the downstream task. Select an existing RGB-D salient object detection model as the downstream network model for supervised training, keeping the training parameters consistent with the original downstream network.

[0060] Step S5: Prediction result output: Based on different RGB-D tasks, use the downstream RGB-D dense prediction task network model trained in step S4 to complete the inference and output the prediction result.

[0061] Figure 3 This demonstrates the effectiveness of the present invention on two common dense object detection tasks. The first and second columns show the input RGB-depth image pairs, the third column shows the ground truth results for salient object detection and semantic segmentation, and the fourth and fifth columns show the results of the baseline method and the present invention, respectively. Figure 3 As shown, the "baseline model" based on traditional contrastive learning focuses more on learning global representations and is not good at capturing local details. This invention, on the other hand, achieves excellent results in both global localization and capturing object details, and achieves performance comparable to ImageNet supervised pre-training schemes.

[0062] Test case

[0063] Performance testing of the self-supervised pre-trained model: To verify the generalization of the pre-trained parameters of this invention, this test case uses three existing excellent RGB-D salient object detection downstream networks (SPNet, CMINet, and CFIDNet) for performance testing.

[0064] The pre-training method used for comparison is as follows:

[0065] Random: Initializes the downstream encoder with random parameters;

[0066] MoCo: Train the encoder for each modality separately using the RGB-D data in this example, and initialize the downstream encoder with that parameter;

[0067] Supervised: Initializes the downstream backbone network with parameters pre-trained using ImageNet data.

[0068] The evaluation metrics for the RGB-D salient object detection task are as follows:

[0069]

[0070] The MAE metric measures the absolute error between the predicted map and the label map; a lower value indicates a better model.

[0071] S α =α*S o +(1-α)*S r

[0072] S α It is a well-known metric for measuring the extraction of foreground and background from an image by a model. It is also applicable to salient object detection tasks and is usually set to 0.5 as an adjustment parameter.

[0073]

[0074] F β β is a common metric for measuring the accuracy and recall of a comprehensive model. In salient object detection tasks, where more attention is paid to the accuracy of the prediction map, the β value is set to 0.3.

[0075]

[0076] E φ It is often used to capture pixel-level data and matching information of local pixels, while Φ FM It is an enhanced alignment matrix. Unlike MAE, higher values ​​for the last three metrics indicate a better model.

[0077] The comparison results of the self-supervised pre-training method of this invention with three other initialization strategies in quantitative analysis are shown in the table below:

[0078]

[0079] As can be seen from the table above, the pre-training framework of this invention far surpasses the strong baseline model MoCo, demonstrating performance comparable to the ImageNet pre-trained model across all three downstream networks. Furthermore, on the DES, DUT, and COME-E datasets, CFIDNet pre-trained using the framework of this invention outperforms ImageNet pre-trained variants on multiple metrics.

[0080] Therefore, this invention fully utilizes the spatial alignment within RGB-D pairs to design our RGB-D cross-modal self-supervised framework, a spatially aware multimodal contrastive learning framework tailored for downstream dense prediction tasks. Through a carefully designed local-global coupled projection module, the model can be well pre-trained with multi-scale features and various cross-modal correlations. Extensive experiments demonstrate that this model outperforms other pre-training methods, showcasing the effectiveness of the proposed cross-modal contrastive learning strategy and local-global coupled projection, as well as its strong generalization ability across various downstream multimodal dense prediction tasks.

[0081] The cross-modal contrastive learning method for dense RGB-D prediction visual tasks provided by this invention can widely offer high-performance pre-training parameters for downstream models such as pixel-level segmentation and detection. Specifically, in industry, companies with a large number of unlabeled RGB-D images can maximize the advantages of this contrastive learning framework to obtain models that meet their specific needs and surpass supervised pre-training methods.

[0082] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. A cross-modal contrastive learning method for dense prediction tasks of RGB-D images, characterized in that... A self-supervised RGB-D cross-modal framework is constructed to pre-train the encoder. The pre-trained encoder parameters are then input into the network model of the downstream RGB-D dense prediction task. The network model is then trained in a supervised manner to obtain the trained network model of the downstream RGB-D dense prediction task, which completes the inference and outputs the prediction results. The RGB-D cross-modal self-supervised framework includes at least a local-global coupling module and a cross-modal training paradigm. The local-global coupling module contains a spatially aware cross-modal contrast loss for a multi-scale projector, which can simultaneously learn features with global understanding, local context, and multi-scale cross-modal relevance during the pre-training phase. The cross-modal training paradigm uses only cross-modal consistency as the multi-modal contrast loss. Specifically, it includes the following steps: S1, Dataset Acquisition: Acquire RGB-D cross-modal image datasets, which are used for pre-training the encoder in the RGB-D cross-modal self-supervised framework and training the network model for the downstream RGB-D dense prediction task, respectively. S2, Encoder selection: ResNet50 network structure is adopted as the encoder; S3, Construct a self-supervised framework: Construct an RGB-D cross-modal self-supervised framework to pre-train the encoder selected in step S2. The pre-training provides RGB-D domain-specific model initialization parameters for downstream RGB-D dense prediction tasks. S31: Images of each training sample It will be done through two different data augmentation operations Transform to obtain This results in different perspectives on the same sample, among which This indicates digital augmentation operations such as random cropping, splicing, rotation, scaling, and flipping. S32: Each pair of RGB and Depth paired images is then input into the corresponding encoder after undergoing the same data augmentation process as in step S31. With momentum encoder Encode: ,in and The feature map generated after encoding; S33: Feature extraction is performed on the feature map generated in step S32 using a spatially aware local-global coupling module: f is used to obtain the global feature map y through a global pooling layer, and local feature maps F1 and F2 are obtained through two local pooling layers. The three feature maps y, F1, and F2 are unfolded into vectors, and the vectors are merged to make them a one-dimensional representation containing multi-scale features. ; S34: Use one-dimensional vectors generated from different modalities for cross-modal comparative learning pre-training, with the total loss function being the cross-modal loss function: ; In the formula, Used as anchor samples for updating RGB modal encoder parameters. The anchor samples are used to update the parameters of the Depth mode encoder. and These are negative sample pools for two different modalities; S4, Network model training for the prediction task: Input the self-supervised pre-trained encoder parameters obtained in step S3 into the network model of the downstream RGB-D dense prediction task, and perform supervised training on the network model. S5, Prediction Result Output: Based on different RGB-D tasks, the downstream RGB-D dense prediction task network model trained in step S4 is used to complete the inference and output the prediction result.

2. The cross-modal contrastive learning method for dense prediction tasks of RGB-D images as described in claim 1, characterized in that: In the data set of step S1, different RGB-D unlabeled image datasets are collected as training sets for the pre-training of the encoder by the RGB-D cross-modal self-supervised framework; and the publicly labeled datasets of the corresponding tasks are used as training sets for supervised learning of the downstream RGB-D dense prediction task.

3. The cross-modal contrastive learning method for dense prediction tasks of RGB-D images as described in claim 2, characterized in that: In step S32 The moving average method is used for parameter updates, and the formula is defined as follows: ; in, for The parameter set, for The parameter set.

4. The cross-modal contrastive learning method for dense prediction of RGB-D images as described in claim 2, characterized in that, The loss function in step S34 is a contrastive loss function, specifically defined as follows: ; In the formula, and They are positive sample pairs. and Negative sample pairs This represents the negative sample vector of the negative sample pool. This represents the temperature coefficient.

5. A cross-modal contrastive learning system for dense prediction tasks of RGB-D images, comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-4 above.

Citation Information

Patent Citations

  • Full-view medical picture region segmentation method based on self-supervised learning

    CN115760875A

  • Visual-semantic representation learning via multi-modal contrastive training

    US20220284321A1