System and method for cross-modal knowledge transfer without using task-related source data
SOCKET addresses the challenge of cross-modal knowledge transfer without task-related data by using task-irrelevant data and batch normalization statistics, achieving superior performance in adapting target models across different modalities.
Patent Information
- Application Number
- JP2025519306
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-01-16
- Filing Date
- 2023-06-02
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2043-06-02
AI Technical Summary
Existing cross-modal knowledge transfer methods require access to task-related paired data or source data, which is not feasible due to memory or privacy concerns, especially when transferring knowledge from one source modality to a different target modality.
The SOCKET (Source-free Cross-modal KnowledgE Transfer) method uses task-irrelevant paired data and matches the mean and variance of target features to the batch normalization statistics in the source model, enabling effective knowledge transfer without access to source data.
SOCKET significantly outperforms existing source-free methods by up to 12% in classification tasks, effectively reducing the modality gap and adapting target models without labeled target data.
Smart Images

Figure 2025518979000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a system and method for cross-modal knowledge transfer without using task-related source data.
Background Art
[0002] The cross-modal knowledge distillation (CMKD) method aims to learn the rich representation of a certain modality without a large number of labeled data from a large labeled dataset of another modality. The above method has been used in various practical computer vision tasks such as action recognition and image recognition. Most of the research along this line is based on the premise that task-related paired data across different modalities can be accessed. The recent research line has relaxed this premise in the sense of domain generalization, that is, task-related paired data on the target domain cannot be accessed, but task-related paired data on the source domain can be accessed. As an example, the above method considers Uniform Database Access (UDA) across different domains where the target domain has unlabeled RGB-D pairs instead of a single modality. All of the above research either uses task-related paired data for cross-modal knowledge transfer or regards cross-modal paired data as a domain. There is also research on zero-shot domain adaptation that uses paired data not related to external tasks but requires access to source data. To address the memory or privacy issues regarding source data, a new research line named Hypothesis Transfer Learning (HTL) has recently emerged, where only the trained source model can be accessed instead of source data. In this situation, people have explored the adaptation of target domain data with limited labels or no labels in the presence of both single-source, i.e., Source-Free Domain Adaptation (SFDA), or multi-source models, i.e., Multiple Source-Free Domain Adaptation (MSFDA). The above method does not work well in a regime where the unlabeled target set comes from a modality different from the source.There is a need for a novel cross-modal knowledge transfer method that can perform effective knowledge transfer without accessing task-related data used to train the source model, considering different source modalities and target modalities.
Summary of the Invention
[0003] The present disclosure relates to systems and methods for cross-modal knowledge transfer systems and methods that do not use task-related source data.
[0004] Some embodiments of the present invention have realized cost-effective depth sensors and infrared sensors as alternatives to conventional RGB sensors, and explain that the advantages of these sensors over RGB in areas such as autonomous navigation and remote sensing are clearly understood. Therefore, it is very important to build computer vision and deep learning systems for depth and infrared data. However, large labeled datasets for these modalities are still lacking. In such cases, it is very valuable to transfer knowledge from a neural network trained on a large, well-labeled dataset in the source modality (RGB) to a neural network that functions on the target modality (depth, infrared, etc.). There may be cases where access to the source data is not possible due to reasons such as memory and privacy, and knowledge transfer needs to function only with the source model. We will explain SOCKET: source-free Cross-modal KnowledgE Transfer, which is an effective solution to this difficult task of transferring knowledge from one source modality to a different target modality without access to the source data. The framework reduces the modality gap by using paired, task-irrelevant data and matching the mean and variance of the target features to the batch normalization statistics present in the source model. We demonstrate through large-scale experiments that our method outperforms existing source-free methods for classification tasks that do not consider the modality gap by up to 12% in some cases.
[0005] According to some embodiments of the present invention, a cross-modal knowledge transfer system for adapting one or more source model networks to one or more target model networks is provided. The cross-modal knowledge transfer system includes a task-irrelevant (TI) paired dataset, an unlabeled task-related (TR) dataset, one or more source model networks including a batch normalization (BN) layer, a feature encoder, a convolutional neural network layer (CNN layer), and a classifier, one or more target model networks including a BN layer, a feature encoder, a CNN layer, and a classifier, a memory configured to store a cross-modal knowledge transfer method implemented by a computer having instructions, and at least one processor configured to execute steps of the cross-modal knowledge transfer method implemented by the computer according to the instructions. The steps include extracting TI source features and TR source moments from one or more source model networks by sending a TI source paired dataset through the one or more source model networks, wherein the CNN layer and the classifier of the one or more source model networks are frozen. The steps further include extracting per-batch TI target features and TR target moments from one or more target model networks by sending the TI paired dataset and the unlabeled TR dataset through the one or more target model networks, wherein the classifier of the one or more target model networks is frozen. The steps further include calculating a modality-independent loss function based on the extracted TR target features of the one or more target model networks, jointly training the feature encoders of the one or more target model networks by minimizing the calculated modality-independent loss function, and generating a final target model network by combining the trained one or more target model networks.
[0006] Furthermore, some embodiments of the present invention provide a cross-modal knowledge transfer method implemented by a computer having instructions, and the cross-modal knowledge transfer method uses at least one processor and at least one memory. In this case, the instructions include extracting TI source features and TR source moments from one or more source model networks by sending a TI source pair dataset through the one or more source model networks, and the CNN layers and classifiers of the one or more source model networks are frozen. The instructions further include extracting per-batch TI target features and TR target moments from one or more target model networks by sending a TI pair dataset and an unlabeled TR dataset through the one or more target model networks, and the classifier of the one or more target model networks is frozen. The instructions further include calculating a modality-independent loss function based on the extracted TR target features of the one or more target model networks, jointly training the feature encoders of the one or more target model networks with mixing weights by minimizing the calculated modality-independent loss function, and generating a final target model network by combining the trained one or more target model networks.
[0007] The accompanying drawings included for a better understanding of the present invention illustrate embodiments of the present invention and, together with the description of the specification, serve to explain the principles of the present invention.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12A
Figure 12B
Figure 13
Best Mode for Carrying Out the Invention
[0009] Hereinafter, various embodiments of the present invention will be described with reference to the drawings. Note that the drawings are not drawn to the correct scale, and elements having the same structure or function are denoted by the same reference numerals throughout the drawings. Also note that the drawings are only intended to facilitate the description of specific embodiments of the present invention. These are not intended to comprehensively describe the present invention or to limit the scope of the present invention. In addition, aspects described in relation to a particular embodiment of the present invention are not necessarily limited to that embodiment and can be implemented in any other embodiment of the present invention.
[0010] FIG. 1 shows a schematic diagram for explaining a Source-free Cross-modal Knowledge Transfer (SOCKET) method according to some embodiments of the present invention. The problem of single / multi-source cross-modality knowledge transfer without using data to train the source model is described. In order to effectively perform knowledge transfer, the cross-modal features are aligned for pair data unrelated to the task in the feature space, and the distribution of the label-free task-related features and the source features is matched to minimize the modality gap.
[0011] Figure 2 shows related research and compares them with SOCKET. In this figure, the existing problem settings in the literature regarding knowledge transfer across different domains and modalities are compared with our research. The comparison target settings described in this figure are respectively: (1) UDA (Unsupervised Domain Adaptation), DT (Domain Translation), (2) MSDA (Multi-source domain adaptation), (3) SFDA (Source-free single-source DA), (4) MSFDA (Source-free multi-source DA), (5) CMKD (Cross-modal knowledge distillation), and (6) ZDDA (Zero-shot DA). Note that only the system according to the present invention enables cross-modal knowledge transfer from multiple sources without any access to the relevant source training data for the target dataset without labels of different modalities.
[0012] Figure 3 shows source models 1 to n (program modules 1 to n) formed by one or more neural networks based on the cross-modal knowledge transfer system 1100 (shown in FIG. 13). As an example, source model 310 includes a convolutional neural network (CNN) layer 311 that outputs an intermediate feature map 312, a batch normalization (BN) layer 313 in feature encoder 314, and a classification layer (classifier). The output of feature encoder 315 is sent to fully connected classifier layer 316. In some cases, the source model (corresponding to the source modality) may be called the source model network. Our framework can be divided into the following two parts. (i) Before knowledge transfer (upper side): The cross-modal knowledge transfer system activates the BN layer of the source model and the source feature encoder, freezes the other parts of the source model including the CNN and classifier, and extracts TI source features and source moments by sending task-irrelevant (TI) source data through the activated BN layer and source feature encoder. The term "freeze" means that the parameters in the frozen layer remain fixed and are not updated. Since task-relevant (TR) source feature maps are not available, the cross-modal knowledge transfer system extracts the moments of its distribution from the BN layer. (ii) During knowledge transfer (lower side): The cross-modal knowledge transfer system freezes only the classification layer (classifier), provides TI and label-free TR target data to the target model (corresponding to the target modality), and uses the mixing weight ζ k to obtain, for each batch, the combined TI target features 324 and TR target moments 322 from all models respectively. In some cases, the target model may be called the target model network.
[0013] The cross-modal knowledge transfer system matches them with source features 321 pre-extracted using the TI feature matching loss 320 and matches (combines) them with source moments 317 using the distribution matching loss 318. Both the TI feature matching loss and the distribution matching loss are modality-specific losses. The TR target features 323 are used to calculate a modality-independent loss (modality-independent loss function) 319. The combination of the modality-specific loss function and the modality-independent loss function is minimized to jointly train all feature encoder parameters with the mixing weight ζ k together. The final target model is the optimal linear combination of the updated source model (corresponding to the trained target model network).
[0014] Depth sensors such as Kinect and RealSense, LIDAR for directly measuring point clouds, or high-resolution infrared sensors such as FLIR enable the expansion of the scope of computer vision applications compared to using only visible wavelengths. By directly detecting depth, an approximate 3D image of the scene can be provided, thereby improving the performance of applications such as autonomous navigation, while detection in the infrared wavelength enables easier pedestrian detection or more advanced object detection in adverse atmospheric conditions such as rain, fog, and smoke. These are just a few examples.
[0015] For modalities such as depth and infrared, building computer vision applications using current explicit teacher deep learning approaches requires large amounts of diverse labeled data. However, such large and diverse datasets do not exist for these modalities, and the cost of constructing such datasets can be prohibitively high. In such cases, researchers have developed methods such as knowledge distillation to transfer knowledge from models trained on modalities such as RGB for which large amounts of labeled data are available to target modalities such as depth.
[0016] Different from previous studies, we address a new and challenging problem related to cross-modal knowledge transfer. Assume that (a) there is a source model trained for the task of interest (TOI), and (b) we only have access to unlabeled data in the target modality where a model for the same TOI needs to be constructed. An important aspect is that the cross-modal knowledge transfer system does not require access to any data from the source modality of the TOI. Such a problem setting is important when memory and privacy considerations do not allow sharing of training data from the source modality, and only the trained model can be shared.
[0017] Some embodiments provide SOCKET, i.e., source-free cross-modal knowledge transfer, as an effective solution to this problem to bridge the gap between the source and target modalities. To that end, (1) it is shown that the use of an external dataset of source-target modality pairs unrelated to the TOI, called task-irrelevant (TI) data, can help learn an effective target model by bringing the features of the two modalities closer. In addition to the use of TI data, we encourage matching the statistics of the features of the unlabeled target data, defined as task-relevant (TR), with the statistics of the source data available from the normalization layers present in the trained source model.
[0018] We provide important empirical evidence showing that modality shift from a source modality such as RGB to a target modality such as depth can be much more difficult than domain shift from one RGB dataset to another RGB dataset. This indicates that the proposed framework is necessary to help minimize the modality gap for more effective knowledge transfer. Based on the above idea, we show that we can improve existing prior art methods devised only for cross-domain settings in the same modality. The main features of this disclosure are summarized below.
[0019] 1. Formulate a new problem for knowledge transfer from a model trained for a source modality to a different target modality when there is no access to any task-related source data and the target data is unlabeled.
[0020] 2. To bridge the gap between modalities, propose SOCKET as a new framework for cross-modal knowledge transfer without accessing source data by (a) using an external pair dataset unrelated to the task and (b) matching the moments obtained from the normalization layers in the source model with the moments calculated on the unlabeled target data.
[0021] 3. Large-scale experiments on multiple datasets for both knowledge transfer from RGB to depth and knowledge transfer from RGB to IR, and for both single-source and multi-source cases, show that SOCKET is useful for reducing the modality gap in the feature space and significantly outperforms existing source-free domain adaptation baselines that do not consider the modality differences between the source and target modalities (improvements of up to 12% in some cases).
[0022] 4. Also, empirically, for the target dataset, the knowledge transfer problem between modalities such as RGB and depth is more difficult than the mere domain shift within the same modality such as sensor changes and viewpoint shifts. (Problem setting and notation)
[0023]
Number
[0024]
Number
[0025]
Number
[0026]
Number
[0027]
Number
[0028]
Number
[0029]
Number
[0030]
Number
[0031] In the matching of features not related to the task, the TI features of two modalities in the feature space are matched. Even if this captures some class-independent cross-modal mapping between the source modality and the target modality, there is no information regarding the TR class-conditioned cross-modal mapping. Using this term, when the relevant class is given, the cross-modal relationship between the source and the target is referred to. Assuming that the marginal distribution of the source features across batches can be modeled as Gaussian, such feature statistics can be fully characterized by its mean and variance. We propose to match the feature statistics spanning the source and the target in order to further reduce the modality gap.
[0032]
Number
[0033]
Number
[0034] The above two proposed methods help reduce the modality gap between the source and the target without accessing task-related source data. In addition to these, unlabeled target data is directly used for knowledge transfer. Specifically, information maximization is performed along with the minimization of the self-teaching pseudo-label loss.
[0035] Information Maximization (IM): IM is essentially a task of maximizing the amount of mutual information between the distribution of the target data and its label predicted by the source model. This amount of mutual information is a combination of the conditional entropy and the marginal entropy of the target label distribution.
[0036]
Number
[0037] [Number]
[0038] Calculate the comprehensive objective function as the sum of the modality-independent loss and the modality-specific loss, and optimize the weights in the feature encoder by minimizing the following objective function, also called the loss function, using the algorithm shown in Figure 4. This figure shows a sufficient explanation of the SOCKET method in the form of an algorithm according to an embodiment of the present invention. [Number] (Experiment)
[0039] First, explain the details of the dataset, baseline, and experiment to be used. Next, show the results of single-source and multi-source cross-modal transfer demonstrating the effectiveness of our method. Also, experimentally demonstrate that source-free cross-modal is a much more difficult problem compared to cross-domain knowledge transfer. End the experiment by performing an analysis of different hyperparameters. (Details of Dataset, Baseline, and Experiment)
[0040] Dataset: To demonstrate the effectiveness of our method, we conduct extensive tests on publicly available cross-modal datasets. Results for two RGB-D (RGB and depth) datasets, namely SUN RGB-D and DIML RGB+D, as well as RGB-NIR scene (RGB and near-infrared) datasets are shown. Figure 5 shows the main features of the datasets used in the experiment according to an embodiment of the present invention, summarizing the statistics of the datasets.
[0041] SUN RGB-D: A scene understanding benchmark dataset containing 10,335 RGB-D image pairs of indoor scenes. This dataset has images acquired from four different sensors named Kinect version 1 (kv1), Kinect version 2 (kv2), Intel® RealSense, and Asus Xtion. These four sensors are treated as four different domains. All of these images are distributed among 45 classes, 17 of which are common to all domains. These common classes are called TR classes, and the remaining 28 classes are called TI classes. To train four source models, one for each domain, RGB images from the TR classes specific to that particular domain are used. The TR depth images from each domain are treated as the target modality dataset. The goal here is to perform classification among the TR scene classes by adapting the RGB source model to unlabeled target data in the depth modality.
[0042] DIML RGB+D: This publicly available dataset consists of over 200 indoor / outdoor scenes. Instead of a complete dataset with 1500 / 500 RGB-D pairs for training / testing distributed among 18 scene classes, a smaller sample dataset is used. The training pairs are split into RGB and depth, and these two are treated as source and target respectively. Following Figure 5, these images are further split into TR images and TI images. Synchronized RGB-D frames are captured using Kinect v2 and Zed stereo cameras.
[0043] RGB-NIR Scene: This publicly available dataset consists of 477 images from 9 scene categories captured in RGB and near-infrared (NIR). These images are captured using separate exposures from a modified SLR camera using visible and NIR. By designating 6 of the categories as TR and the remaining 3 categories as TI, single-source knowledge transfer is performed on this dataset. For this dataset, two experiments were conducted: adaptation from RGB to NIR and vice versa. (Baseline Method)
[0044] The problem description noted in this specification is novel and has not been considered in past literature. Therefore, there is no direct baseline for our method. However, the most relevant research is source-free cross-domain knowledge transfer methods that work for both single-source and multi-source cases. SHOT and DECISION are the most promising and well-known research on single-source and multi-source SFDA respectively, and we compare with these two methods.
[0045] Unlike SOCKET, none of these baselines adopt strategies to overcome modality differences and use only the modality-independent loss L ma to train the target model. Using scene classification as the target task, it is shown that SOCKET outperforms these baselines for cross-modal knowledge transfer without using access to task-related source data. (Network Architecture)
[0046] In our experiments, we adopt the well-known Resnet50 model pre-trained on ImageNet as the backbone architecture for training the source model. According to the architecture, the last fully connected (FC) layer is replaced with a bottleneck layer containing 256 units, and a batch normalization (BN) layer is added at the end of the FC layer. A task-specific FC layer with weight normalization is added at the end of the bottleneck layer. (Implementation of Knowledge Transfer)
[0047] It should be recalled that the target model is initialized with source weights and the classifier layers are frozen. The weights in the feature encoder and the source mixing weight parameter (ζ k ) in the case of multi-source are optimization parameters. For all the following experiments, λ pl is set to 0.3. The regularization parameters λ TI and λ d for modality-specific losses are set to be equal. Empirically, these parameters are selected to balance the modality-independent loss and ensure that no loss component exceeds the others by a large margin. Empirically, the range (0.1, 0.5) has been found to work best. All values within this range have better performance than the baseline, and we report the best accuracy among them. For images from modalities other than RGB, namely depth and NIR, single-channel images are repeated to 3-channel images so that they can be fed through the feature encoder initialized from the source model trained on RGB images. A batch size of 32 is used for all experiments. We run our method 3 times for all experiments using 3 random seeds of PyTorch and report their average accuracy. (Results Regarding SUN RGB-D Dataset)
[0048] Our method is general enough to handle any number of sources and demonstrates both single-source and multi-source knowledge transfer. Figure 6 shows the results for the SUN RGB-D dataset for a single-source cross-modal knowledge transfer task from RGB to depth modality without using access to task-related source data, for all pairs of domains according to an embodiment of the present invention. In this figure, single-source RGB-depth results are shown for all four domains. The unlabeled depth data of each domain is treated as the target and these are adapted using a source model trained on RGB data from each of the four domains. The rows represent the RGB domains on which the source models are trained. The columns represent the knowledge transfer results for the depth domains for three methods, and "unadapted" shows the results for the unadapted source, SHOT, and SOCKET. What is readily apparent from Figure 6 is that for the target domains Kinect V1, Kinect V2, Realsense, and Xtion, SOCKET consistently outperforms the baseline by a sufficient margin of 6.7%, 4.5%, 2.3%, and 3.8% respectively, thereby demonstrating the effectiveness of SOCKET in a source-free cross-modal setting. In some cases, SOCKET outperforms the baseline by a very large margin of 12.4% (from Realsense-RGB to Kinect V1-depth) or 9.0% (from Xtion-RGB to Xtion-depth).
[0049] Figure 7 shows the results for the SUN RGB-D dataset for a multi-source cross-modal knowledge transfer task from RGB to depth modality without using access to task-related source data, according to an embodiment of the present invention.
[0050] In this case, this figure shows the results of 2-source RGB-depth adaptation. For the four domains, six 2-source combinations are obtained, and each of them is used for adaptation to depth data from all of the four domains. The columns represent the knowledge transfer results regarding domain-specific depth data for DECISION and SOCKET. Also in this case, on average, SOCKET was found to have performance that is sufficiently better by a margin than the baselines for all four target domains. Following the trend of single-source adaptation, SOCKET shows some very good improvements in several individual cases, such as a 12.2% improvement for Kinect v1 depth from (Kinect v1+Xtion)-RGB and a 10.4% improvement for Kinect v2 depth from (Kinect v2+Realsense)-RGB. (Results regarding the DIML RGB+D dataset)
[0051] For this dataset, single-source adaptation experiments were conducted by reconstructing the dataset according to FIG. 5. FIG. 8 shows the results regarding the DIML dataset when different datasets not related to the task are used by SOCKET in comparison with the unadapted source and SHOT, according to some embodiments of the present invention. In FIG. 8, TI data from both the DIML RGB+D dataset and the SUN RGB-D dataset are used in two separate columns, and the TI data of SUN RGB+D is the same as that used in the experiments regarding the SUN RGB-D dataset. By doing so, it is shown that SOCKET can function well even with TI data from completely different datasets, and it was found that SOCKET has relative gains of 4.7% and 11.8% with respect to the baseline for these two TI data settings, respectively. (Results regarding the RGB-NIR scene dataset)
[0052] Next, even when the modalities are RGB and NIR using the RGB-NIR dataset, SOCKET shows better performance than the baseline. Here, we follow the split shown in Figure 5. Experiments are conducted for both RGB to NIR and vice versa. Figure 9 shows the results for the RGB-NIR dataset for single-source cross-modal knowledge transfer tasks from RGB to NIR and vice versa, without using task-related source data, according to some embodiments of the present invention. The columns represent the knowledge transfer results on the depth domain for three methods, and "unadapted" shows the results on the unadapted source, SHOT, and SOCKET. For the transfer from RGB to NIR, SOCKET shows a 3.5% improvement, while for the transfer from NIR to RGB, it shows a 0.5% improvement over competing methods. (Cross-modal vs. cross-domain)
[0053] To show the importance of the new problem we are considering, we compare the single-source knowledge transfer results for the SUN RGB-D dataset for modality change vs. domain shift. Figure 10 shows the difference in the results of cross-modal and cross-domain knowledge transfer for SUN RGB-D scene classification using SHOT, according to some embodiments of the present invention. SHOT, which is a source-free UDA method for this experiment, is used. All domain-specific source models are trained on RGB images. For domain shift, the target is all RGB images of the remaining three domains, and the average is reported. Domain shift involves changes in sensor configuration, viewpoint, etc. In the case of modality change, the target data is depth images from the same domain. The scene is the same as the RGB source, except that it is captured using a depth sensor. This figure clearly shows that when transferring knowledge between different modalities rather than between domains of the same modality, the accuracy drops by a large margin of 12.5%. This indicates that cross-modal knowledge transfer is not the same as DA, and a framework like SOCKET is necessary to reduce the modality gap in effective cross-modal knowledge transfer. (Ablation and Sensitivity Analysis) (Contribution of Loss Components)
[0054] Figure 11 shows the impact on accuracy of the proposed new loss according to some embodiments of the present invention. The first accuracy column (a) corresponds to single-source adaptation from RGB to depth on the Kinect V2 domain, and the second column (b) shows the results of multi-source RGB-depth adaptation from Kinect V1 Xtion of the SUN RGB-D dataset to the Kinect V1 domain. The first row is the result for only the modality-independent loss L ma alone, and the second and third rows show the individual impacts of the proposed modality-specific losses along with L ma . In both cases, SOCKET outperforms the baseline. The final row showing both of the proposed losses along with L ma yields the best results. Only in parentheses is shown the accuracy gain for using only the modality-independent loss (L ma ). (Effect of the Number of TI Images)
[0055] Figure 12A shows the effect of the number of TI data on the SOCKET results for the SUN RGB-D dataset according to some embodiments of the present invention. Knowledge transfer from Kinect v1 RGB to unlabeled depth data is performed. Six random TI classes are used, and the number of TI images per class is varied in 20 steps from 0 to 60. This figure clearly shows that by increasing the samples of TI data per class, the scene classification accuracy for RGB-to-depth transfer for the SUN RGB-D dataset is improved. In short, for a fixed number of TI classes, the higher the number of TI images per class, the higher the performance of SOCKET. (Effect of the Regularization Parameter)
[0056] Figure 12B shows the effect of the regularization hyperparameter on the test accuracy for a novel loss proposed as part of SOCKET according to some embodiments of the present invention. For the SUN RGB-D dataset, perform a transfer from Kinect v1 and Kinect v2 RGB to Kinect v1 depth. λ TI and λ d are kept equal to each other for values between 0 and 1. Using the value 0 is equivalent to using SHOT. From this figure, it can be seen that as the value of the parameter increases, the accuracy also increases up to a certain point and then begins to decrease.
[0057] We identify a novel and difficult problem of cross-modal knowledge transfer that does not use access to task-related data from the source modality. For effective knowledge transfer to the target modality when there is only unlabeled data, some embodiments of the present invention can provide a framework SOCKET that includes devising a loss function that helps bridge the gap between the two modalities in the feature space. The results of both the experiment from RGB to depth and the experiment from RGB to NIR show that SOCKET outperforms the baseline designed for source-free teacherless domain adaptation, which does not function well in modality shift.
[0058] According to some embodiments of the present invention, a cross-modal knowledge transfer system and a cross-modal knowledge transfer method implemented by a computer can reduce the memory size of storage, solve the privacy problem regarding source data, and shorten the training period of the target model network. Therefore, the cross-modal knowledge transfer system of the present invention and the cross-modal knowledge transfer method implemented by a computer can improve the function of a computer system (processor) and reduce the energy consumption of the computer system.
[0059] According to some features of the present invention, each source model network includes a BN (batch normalization) layer and receives an unannotated / unlabeled set of data of the target modality to be classified. Optionally, an adapted target model (trained target network) is used to perform computer vision tasks on the target modality data.
[0060] In the case of a set of "task-irrelevant" data sets, each data point is a pair of images having corresponding source and target modality images and is utilized to assist the knowledge transfer procedure by reducing the source-target modality gap.
[0061] In a cross-modal knowledge transfer system, statistics from one or more batch normalization layers are matched against the batch-wise statistics of the features of the unlabeled target modality data to assist the knowledge transfer procedure by reducing the source-target modality gap.
[0062] Furthermore, the adapted target model can be obtained by using the source model as an initialization and adjusting the parameters of this model by minimizing one or more loss functions. In this case, the combination of loss functions can include entropy, pseudo-labeling, and diversity, which are defined in the neural network output when unlabeled target data is provided as input.
[0063] Optionally, the combination of loss functions can include the feature distance between the source modality image and the target modality image from task-irrelevant data that helps reduce the modality gap between the source features and the target features.
[0064] Furthermore, the combination of loss functions can include the difference between the source feature statistics obtained from the batch normalization layer of the source model and the statistics of the features of the unlabeled target data set.
[0065] In some cases, the source-target modality pairs may be, respectively, RGB-depth point cloud, RGB-infrared point cloud, RGB-LIDAR point cloud, or vice versa, or other combinations of such modalities. The dataset may be in the form of images taken in a single snapshot or in the form of videos taken over a longer period of time. The tasks performed on the unlabeled target dataset may be computer vision tasks such as image recognition, object recognition, and scene recognition.
[0066] FIG. 13 shows a cross-modal knowledge transfer system 1100 according to some embodiments of the present invention. The system 1100 may be a neural network module trained to provide a layout of a device. The cross-modal knowledge transfer system 1100 includes a TI / TR interface circuit interface 150, at least one processor 120, a storage 130, and a memory 140. The storage 130 and the memory 140 can be integrated into one circuit and can also be referred to as a memory. The storage 130 includes a pair data set 131 not related to a task, an unlabeled target data set 132, a source model network (source model) 133, a target model network (target model) 134, and a cross-modal knowledge transfer program (a cross-modal knowledge transfer method implemented by a computer) 135. The interface 150 is configured to communicate between the memory 140, the storage 130, and at least one processor 120. In some cases, the interface 150 may receive the pair data set 131 not related to the task, the unlabeled target data set 132, the source model network (source model) 133, and the target model network (target model) 134 from an external database device (server) 195 of the system 100 via a communication network 190 including a wireless communication network, a wired network, the Internet, or a combination thereof.
[0067] The above embodiments of the present invention can be realized in any of a number of ways. For example, the embodiments may be realized using hardware, software, or a combination thereof. When realized in software, the software code may be provided on a single computer or distributed among multiple computers and may be executed on any suitable processor or collection of processors. Such a processor may be realized as an integrated circuit having one or more processors as components of the integrated circuit. However, the processor may be realized using circuitry in any suitable format.
[0068] Also, embodiments of the present invention may be implemented as a method, and examples thereof are provided. The order of operations performed as part of this method may be determined in any suitable manner. Accordingly, the embodiments may be configured such that operations are performed in an order different from the illustrated order, which may include performing some operations simultaneously that are shown as a series of operations in the illustrated embodiments.
[0069] In the claims, terms such as "first," "second," which modify an element of a claim, do not themselves imply any superiority, precedence, or order of one element of a claim over another, or any temporal order of performing the acts of a method, but are merely used as labels to distinguish one element of a claim having a particular name (where no ordinal terms are used) from another element of the same claim having the same name.
[0070] It should be understood that although this disclosure has been described by way of example of preferred embodiments, various other adaptations and modifications can be made within the spirit and scope of the present invention.
[0071] Accordingly, the purpose of the appended claims is to cover all such variations and modifications that fall within the true spirit and scope of the present invention.
Claims
1. A cross-modal knowledge transfer system for adapting one or more source model networks to one or more target model networks, wherein the cross-modal knowledge transfer system comprises: having a memory, the memory comprising: A task-irrelevant (TI) pair dataset; An unlabeled task-relevant (TR) dataset; The one or more source model networks including a batch normalization (BN) layer, a feature encoder, a convolutional neural network layer (CNN layer), and a classifier; The one or more target model networks including the BN layer, the feature encoder, the CNN layer, and the classifier; A cross-modal knowledge transfer method implemented by a computer having instructions, and is configured to store, and the cross-modal knowledge transfer system further comprises: At least one processor configured to execute steps of the cross-modal knowledge transfer method implemented by the computer according to the instructions, the steps comprising: Extracting TI source features and TR source moments from the one or more source model networks by sending the TI source pair dataset through the one or more source model networks, wherein the CNN layer and the classifier of the one or more source model networks are frozen, and the steps further comprise: Extracting per-batch TI target features and TR target moments from the one or more target model networks by sending the TI pair dataset and the unlabeled TR dataset through the one or more target model networks, wherein the classifier of the one or more target model networks is frozen, and the steps further comprise: Calculating a modality - independent loss function based on the extracted TR target features of the one or more target model networks; Calculating a modality - specific loss function by calculating the distance between the extracted TI target features and the TI source features, and the distance between the extracted TR target moments and the TI source moments; Jointly training the feature encoders of the one or more target model networks with the mixing weights by minimizing the calculated modality - independent loss function and the modality - specific loss function; Generating a final target model network by combining the trained one or more target model networks, a cross - modality knowledge transfer system.
2. Further comprising a TI / TR dataset interface configured to receive the TI paired dataset and the unlabeled TR dataset via a communication network, The TI / TR dataset interface is configured to store the received TI paired dataset and the unlabeled TR dataset in the memory, the cross - modality knowledge transfer system according to claim 1.
3. The at least one processor is further configured to transmit the model parameters of the generated final target model network to another one or more un - trained target model networks via a network, the cross - modality knowledge transfer system according to claim 1.
4. The pair of the task - unrelated (TI) dataset and the unlabeled task - related (TR) dataset is an RGB image and a depth image, an RGB image and an infrared image, or an RGB image and a LiDAR point cloud, the cross - modality knowledge transfer system according to claim 1.
5. The label-free task-related (TR) dataset is the cross-modal knowledge transfer system according to claim 4, which is used for computer vision tasks including image recognition, object recognition, and scene recognition.
6. The TI source features are extracted from the feature encoder of the one or more source model networks, and the TR source moments are extracted from the BN layers of the one or more source model networks. The per-batch TI target features are extracted from the feature encoder of the one or more target model networks, and the TR target moments are extracted from the BN layers of the one or more target model networks. The cross-modal knowledge transfer system according to claim 1.
7. The per-batch TI target features and the TR target moments extracted from the one or more target model networks are combined using mixing weights respectively. The cross-modal knowledge transfer system according to claim 1.
8. The modality-independent loss function and the modality-specific loss function are used to bridge the gap between the one or more source model networks and the one or more target model networks in the feature space. The cross-modal knowledge transfer system according to claim 1.
9. The combinations of the modality-independent loss functions include entropy, pseudo-labeling, and diversity. The entropy, pseudo-labeling, and diversity are defined by the outputs of the one or more source model networks. The cross-modal knowledge transfer system according to claim 8.
10. The combination of the modality-specific loss functions includes the feature distance between the extracted TI source features and the TI target features, and the distance between the TI source moments and the TR target moments, for the cross-modal knowledge transfer system according to claim 8.
11. A cross-modal knowledge transfer method implemented by a computer having instructions, the cross-modal knowledge transfer method using at least one processor and at least one memory, the instructions being including extracting TI source features and TR source moments from the one or more source model networks by sending a TI source pair dataset through the one or more source model networks, where the CNN layers and classifiers of the one or more source model networks are frozen, and the instructions further including extracting per-batch TI target features and TR target moments from the one or more target model networks by sending the TI pair dataset and the unlabeled TR dataset through the one or more target model networks, where the classifier of the one or more target model networks is frozen, and the instructions further calculating a modality-independent loss function based on the extracted TR target features of the one or more target model networks; calculating a modality-specific loss function by calculating the distance between the extracted TI target features and the TI source features, and the distance between the extracted TR target moments and the TI source moments; jointly training the feature encoders of the one or more target model networks with mixing weights by minimizing the calculated modality-independent loss function and the modality-specific loss function; A computer-implemented cross-modal knowledge transfer method, including generating a final target model network by combining the one or more trained target model networks.
12. Using the TI / TR dataset interface to further receive the TI pair dataset and the unlabeled TR dataset via a communication network, The TI / TR dataset interface is configured to store the received TI pair dataset and the unlabeled TR dataset in the at least one memory. The cross-modal knowledge transfer method implemented by a computer according to claim 11.
13. The at least one processor is further configured to transmit, via a network, the model parameters of the generated final target model network to one or more other untrained target model networks. The cross-modal knowledge transfer method implemented by a computer according to claim 11.
14. The pair of task-unrelated (TI) datasets and the unlabeled task-related (TR) dataset are RGB images and depth images, RGB images and infrared images, or RGB images and LIDAR point clouds. The cross-modal knowledge transfer method implemented by a computer according to claim 11.
15. The unlabeled task-related (TR) dataset is used for computer vision tasks including image recognition, object recognition, and scene recognition. The cross-modal knowledge transfer method implemented by a computer according to claim 14.
16. The TI source features are extracted from the feature encoders of the one or more source model networks, The TR source moments are extracted from the BN layers of the one or more source model networks, The TI target features for each batch are extracted from the feature encoder of the one or more target model networks, The TR target moment is extracted from the BN layer of the one or more target model networks, and the cross-modal knowledge transfer method realized by the computer according to claim 11.
17. The TI target features for each batch and the TR target moments extracted from the one or more target model networks are combined using mixing weights respectively, and the cross-modal knowledge transfer method realized by the computer according to claim 11.
18. The modality-independent loss function and the modality-specific loss function are used to bridge the gap between the one or more source model networks and the one or more target model networks in the feature space, and the cross-modal knowledge transfer method realized by the computer according to claim 11.
19. The combination of the modality-independent loss functions includes entropy, pseudo-labeling, and diversity, The entropy, pseudo-labeling, and diversity are defined by the outputs of the one or more source model networks, and the cross-modal knowledge transfer method realized by the computer according to claim 18.
20. The combination of the modality-independent loss functions includes the feature distance between the extracted TI source features and the TI target features, and the distance between the TI source moment and the TR target moment, and the cross-modal knowledge transfer method realized by the computer according to claim 18.
Citation Information
Patent Citations
Domain adaptation and fusion using task-irrelevant paired data in sequential form
WO2020256732A1