Method and system for comparing cross-modal similarity on basis of contrastive learning
The contrast learning-based method addresses the challenge of comparing similarity across different modal data formats by using a backbone network to extract and contrast features, enhancing object recognition and search capabilities in 3D virtual environments.
Patent Information
- Application Number
- PCT/KR2023/017077
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-31
- Filing Date
- 2023-10-31
- Publication Date
- 2025-05-08
AI Technical Summary
Existing technologies struggle to efficiently compare the similarity between objects represented in different modal data formats, such as images, meshes, and point clouds, in 3D virtual environments, which is crucial for tasks like object recognition, classification, and search.
A contrast learning-based method that uses a backbone network to extract features from multiple modal data types and then contrasts these features to perform cross-modal similarity comparisons, enhancing the ability to recognize, classify, and search objects across different modal representations.
This approach enables effective cross-modal similarity comparisons, improving object recognition, classification, and search capabilities in 3D virtual environments, while being robust to changes in hyperparameters and data quantity, and applicable to various modal data types.
Smart Images

Figure KR2023017077_08052025_PF_FP_ABST
Abstract
Description
Cross-modal similarity comparison method and system based on contrastive learning
[0001] The present invention relates to deep learning-based computer vision technology, and more particularly, to a method for comparing similarity between modal data such as images, meshes, point clouds, etc. based on deep learning.
[0002] With the advancement of semiconductor technologies such as AI-based deep learning and GPUs, technologies for representing 3D objects in 3D virtual spaces, such as the metaverse and digital twins, are gaining attention. Consequently, the need for object search, detection, and classification technologies for randomly acquired objects for the following purposes is emerging.
[0003] 1. Check for copyright infringement
[0004] 2. Search for objects similar to 3D objects
[0005] 3. Check for overlap with the underlying object
[0006] Therefore, it is necessary to compare the similarity of objects for the above purpose for randomly acquired objects, and cross-similarity comparison is required for heterogeneous modal data rather than homogeneous modal data.
[0007] The present invention has been devised to solve the above problems, and the purpose of the present invention is to provide a method for comparing cross-modal similarity between modal objects expressed in multiple modal forms such as images, meshes, and point clouds, in order to enable object recognition, classification, detection, and search by performing cross-modal similarity comparison on objects of multiple modal forms randomly acquired in the process of implementing and configuring a 3D virtual environment.
[0008] In order to achieve the above object, a backbone network learning method according to one embodiment of the present invention includes the steps of: extracting a first feature from first modal data using a first backbone network; extracting a second feature from second modal data having a different modal type from the first modal data using a second backbone network; and performing contrastive learning on the first backbone network and the second backbone network based on classes of the first modal data and the second modal data.
[0009] The learning step may be to perform contrastive learning on the first backbone network and the second backbone network so that the similarity between the first feature and the second feature increases when the first modal data and the second modal data are of the same class, and so that the similarity between the first feature and the second feature increases when the first modal data and the second modal data are of different classes.
[0010] The learning step may be to perform contrastive learning on the first backbone network and the second backbone network so that the similarity between the first feature and the second feature is lower when the first modal data and the second modal data are of different classes, and so that the similarity between the first feature and the second feature is lower when the first modal data and the second modal data are of the same class.
[0011] The backbone network learning method according to the present invention may further include a step of augmenting first modal data to add first' modal data; a step of extracting first' features from the first' modal data using the first backbone network; and a step of performing contrastive learning on the first backbone network and the second backbone network so that the similarity between the first feature and the first' feature is higher than the similarity between the first feature and the second feature.
[0012] The backbone network learning method according to the present invention may further include a step of augmenting second modal data to add second' modal data; a step of extracting second' features from the second' modal data using the second backbone network; and a step of performing contrastive learning on the first backbone network and the second backbone network so that the similarity between the second feature and the second feature is higher than the similarity between the first feature and the second feature.
[0013] The backbone network learning method according to the present invention may further include a step of adding first" modal data by augmenting the first modal data with an augmentation technique different from the augmentation technique applied to add the first' modal data; a step of extracting first" features from the first" modal data using the first backbone network; and a step of performing contrastive learning on the first backbone network and the second backbone network so that the similarity between the first feature and the first" feature is higher than the similarity between the first feature and the second feature.
[0014] The first modal data may be one of image data, mesh data, point cloud data, volume data, and text data, and the second modal data may be another one of image data, mesh data, point cloud data, volume data, and text data.
[0015] The backbone network learning method according to the present invention may further include a step of extracting features from first inference target modal data using a first backbone network; a step of extracting features from second inference target modal data using a second backbone network; and a step of calculating similarity between the extracted features to confirm similarity between the first inference target modal data and the second inference target modal data.
[0016] The backbone network learning method according to the present invention may further include a step of analyzing first inference target modal data and second inference target modal data based on the confirmed similarity.
[0017] According to another aspect of the present invention, a backbone network learning system is provided, comprising: a processor for extracting a first feature from first modal data using a first backbone network, extracting a second feature from second modal data having a different modal type from the first modal data using a second backbone network, and performing contrastive learning on the first backbone network and the second backbone network based on classes of the first modal data and the second modal data; and a storage unit for providing storage space required by the processor.
[0018] According to another aspect of the present invention, a cross-modal similarity comparison method is provided, including: a step of extracting features from first inference target modal data using a first backbone network; a step of extracting features from second inference target modal data using a second backbone network; and a step of calculating similarities between the extracted features to confirm similarities between the first inference target modal data and the second inference target modal data; wherein the first backbone network and the second backbone network are contrastively learned based on classes of first modal data from which the first backbone network extracts features and second modal data from which the second backbone network extracts features.
[0019] According to another aspect of the present invention, a cross-modal similarity comparison system is provided, comprising: a processor for extracting features from first inference target modal data using a first backbone network, extracting features from second inference target modal data using a second backbone network, and calculating similarities between the extracted features to determine similarity between the first inference target modal data and the second inference target modal data; and a storage unit for providing storage space required by the processor; wherein the first backbone network and the second backbone network are comparatively learned based on classes of first modal data from which the first backbone network extracts features and second modal data from which the second backbone network extracts features.
[0020] As described above, according to embodiments of the present invention, by performing cross-modal similarity comparisons for objects expressed in multiple modal forms, such as images, meshes, and point clouds, similarity comparisons are performed on objects of multiple modal forms randomly acquired during the process of implementing and configuring a 3D virtual environment, thereby enabling extensive object recognition, classification, detection, and search.
[0021] In addition, according to embodiments of the present invention, by unifying the form of embedding features for data acquired in a specific modality in a multi-modal environment, advanced deep learning techniques can be applied regardless of the modality, and development and application of deep learning-based computer vision techniques independent of a specific modality are enabled.
[0022] Figure 1 shows the Center loss method,
[0023] Figure 2 is the CLF method,
[0024] Figure 3 shows a contrastive learning method in a supervised learning environment.
[0025] Figure 4 is a cross-modal similarity comparison method according to one embodiment of the present invention;
[0026] Figure 5 is a cross-modal similarity comparison method according to another embodiment of the present invention;
[0027] Figure 6 is a cross-modal similarity comparison system according to another embodiment of the present invention.
[0028] Hereinafter, the present invention will be described in more detail with reference to the drawings.
[0029] Cross-Modal Similarity Measurement (CMSM) methods that enable cross-modal similarity comparison include methods based on Cross-Modal Center Loss and RONO (RObust discriminative learning with NOisy labels for 2D-3D cross-modal retrieval).
[0030] Figure 1 illustrates the Center loss method. As illustrated in Figure 1, Center loss is a learning method in which features (embedded features) extracted through a backbone network in a single-modal environment are trained to resemble the Center feature, which represents the representative feature of each class. The Center feature is also continuously updated during the learning process to represent the characteristic representing each class. The Center loss method, with these characteristics, is primarily used in image classification and recognition.
[0031] Regarding this Center loss, Figure 2 illustrates an example of the CLF method, which can be applied to Center loss in a multi-modal environment. This method ensures that the features extracted by the Center loss method converge to a single center feature in all modalities, regardless of modality. This method enables object-to-object search in a cross-modal environment. Later, to address the issue of inaccurate labeling (noisy labels) that arises during the actual evaluation and execution stages of CLF, the RONO method was introduced, which measures label reliability and applies a modal search method accordingly.
[0032] CLF and RONO methods fundamentally place a significant emphasis on a method called Center loss. However, since Center loss requires the Center feature to best represent the characteristics of a given class, it tends to be highly dependent on fluctuations in the Center feature. Consequently, Center loss is highly sensitive to changes in hyperparameters used in the training process, such as batch size and learning rate. Furthermore, since Center feature updates are determined by the amount of data used for training, the long-tailed issues inherent in traditional image classification are more pronounced in classes with limited data.
[0033] Accordingly, in an embodiment of the present invention, a cross-modal similarity comparison method based on contrastive learning, which deviates from the center loss-based flow, is proposed.
[0034] Figure 3 illustrates Supervised Contrastive Learning, which classifies images using a contrastive learning-based method in a supervised learning environment where ground truth labels for the data exist. As illustrated in Figure 3, each data is trained as a positive sample that expresses similar characteristics for objects of the same class, and a negative sample that expresses different characteristics for objects of different classes. The ultimate goal of this method is to represent positive samples closely and negative samples farther apart.
[0035] Figure 4 is a diagram illustrating a cross-modal similarity comparison method according to one embodiment of the present invention. It extends contrastive learning within supervised learning to a multi-modal environment.
[0036] Figure 4 shows the results of extending the method of representing results of the same class as positive samples and modals of different classes as negative samples to a multi-modal environment. Applying modality-independent contrastive learning to results extracted from multi-modal sources, such as images, meshes, and point clouds, as shown in Figure 4, offers the following advantages.
[0037] 1. It eliminates the tendency to rely on the center feature and maximizes the distance between classes. Unlike the existing center loss, contrastive learning trains similar components closer together and dissimilar components further apart. Therefore, unlike center loss, which is vulnerable to hyperparameters such as batch size and learning rate, contrastive learning is relatively robust to changes in hyperparameters.
[0038] 2. Center loss updates the center feature based on the amount of data, making it unreliable for classes with limited data. In contrast, contrastive learning differentiates classes based on relative features, making it more robust to data volume than the center feature. Furthermore, data augmentation techniques can be used to generate additional data for contrastive learning, making it even more advantageous.
[0039] This method can be applied to various applications such as object search, classification, and recognition by performing feature similarity comparison regardless of inter-modal gaps in a multi-modal environment.
[0040] A cross-modal similarity comparison method according to an embodiment of the present invention is described in detail below with reference to FIG. 4.
[0041] First, labeled training data are acquired to train an image backbone network (e.g., ResNet) that extracts features from image data, a mesh backbone network (e.g., MeshNet) that extracts features from mesh data, and a point cloud backbone network (e.g., DGCNN) that extracts features from point cloud data. In Fig. 4, modal data corresponding to the 'bed' class and modal data corresponding to the 'chair' class are exemplified as training data.
[0042] Features are extracted through backbone networks appropriate for each modality. Specifically, the image backbone network extracts features from bed and chair image data, the mesh backbone network extracts features from bed and chair mesh data, and the point cloud backbone network extracts features from bed and chair point cloud data. There are no restrictions on the dimensionality of the extracted embedding features.
[0043] Afterwards, the extracted features are compared for feature similarity between positive and negative samples, and the objective function is calculated based on the results, and training is performed on the backbone networks through backpropagation. This is a process of comparing and learning similarity by classifying samples of the same class as positive and samples of different classes as negative, regardless of modality.
[0044] Specifically, backbone networks are trained so that the similarity between features extracted from heterogeneous modal data of the same class increases, and the similarity between features extracted from heterogeneous modal data of different classes decreases.
[0045] Once learning is complete, features can be extracted from heterogeneous modal data of the inference target using backbone networks, and the similarity between the extracted features can be calculated to determine the degree of similarity between heterogeneous modal data.
[0046] Figure 5 is a diagram illustrating a cross-modal similarity comparison method in another embodiment of the present invention. This method augments training data prior to performing contrastive training on backbone networks. Utilizing data augmentation can result in better performance because it can robustly respond to data changes during the actual evaluation and execution stages, separate from the learning process.
[0047] Data augmentation can be applied to image data, mesh data, and point cloud data. Multiple augmentation methods can be applied to each modality to generate multiple sets of augmented data. For example, for training data, augmented data can be generated using data augmentation technique #1 and then again using augmentation technique #2. There is no limit to the number of data augmentations.
[0048] The following method extracts features through a backbone network suitable for each modality, compares the feature similarity between the extracted features for positive and negative samples, and then performs learning through backpropagation by calculating the objective function based on the result of the feature similarity comparison.
[0049] However, in similarity comparisons, the backbone networks must be trained to enhance the similarity between features extracted from the training data and the augmented training data. Naturally, the backbone networks should be trained to enhance the similarity between features extracted from heterogeneous modal data within the same class, while reducing the similarity between features extracted from heterogeneous modal data within different classes.
[0050] FIG. 6 is a diagram illustrating the configuration of a cross-modal similarity comparison system according to another embodiment of the present invention. As illustrated, the cross-modal similarity comparison system according to an embodiment of the present invention can be implemented as a computing system comprising a communication unit (110), an output unit (120), a processor (130), an input unit (140), and a storage unit (150).
[0051] The communication unit (110) is a communication interface for connection with an external network or external device, and acquires heterogeneous modal data for learning and heterogeneous modal data to be inferred. The output unit (120) is an output means for displaying the results of calculations performed by the processor (130), and the input unit (140) is a user interface for receiving user commands and transmitting them to the processor (130).
[0052] The processor (130) augments learning data, performs backbone network contrastive learning, and determines cross-similarity between heterogeneous modal data according to the procedures illustrated in FIGS. 4 and 5 described above. The storage unit (150) provides the storage space necessary for the processor (130) to function and operate.
[0053] So far, a preferred embodiment of a cross-modal similarity comparison method and system based on contrastive learning has been described in detail.
[0054] In an embodiment of the present invention, unlike existing methods based on Center Loss, we propose a cross-modal similarity comparison method that is relatively robust to performance changes due to hyperparameters and the amount of training data by applying a comparison method based on cross-modal similarity criteria contrastive learning. Furthermore, we apply data augmentation techniques to enable feature estimation that is robust to various changes required in real-world environments.
[0055] In the above embodiments, the image data, mesh data, and point cloud data mentioned as modal data are merely exemplary. The technical concepts of the present invention can also be applied to cases where other modal data, such as volume data or text data, are added or replaced.
[0056] Meanwhile, the similarity between modal data determined by the cross-modal similarity comparison method and system according to an embodiment of the present invention can be applied to various applications such as image search, classification, and recognition.
[0057] Meanwhile, it goes without saying that the technical idea of the present invention can also be applied to a computer-readable recording medium containing a computer program that performs the functions of the device and method according to the present embodiment. In addition, the technical idea according to various embodiments of the present invention can be implemented in the form of computer-readable code recorded on a computer-readable recording medium. The computer-readable recording medium can be any data storage device that can be read by a computer and store data. For example, the computer-readable recording medium can be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical disk, a hard disk drive, etc. In addition, the computer-readable code or program stored on the computer-readable recording medium can be transmitted through a network connected between computers.
[0058] In addition, although the preferred embodiments of the present invention have been illustrated and described above, the present invention is not limited to the specific embodiments described above, and various modifications can be made by a person having ordinary skill in the art to which the present invention pertains without departing from the gist of the present invention as claimed in the claims, and such modifications should not be understood individually from the technical idea or prospect of the present invention.
Claims
1. A step of extracting a first feature from the first modal data using a first backbone network; A step of extracting a second feature from second modal data having a different modal type from the first modal data using a second backbone network; A backbone network learning method, characterized by including a step of performing contrastive learning on a first backbone network and a second backbone network based on classes of first modal data and second modal data.
2. In claim 1, The learning phase is, A backbone network learning method characterized in that the first backbone network and the second backbone network are contrastively trained so that the similarity between the first feature and the second feature increases when the first modal data and the second modal data are of the same class, and the similarity between the first feature and the second feature increases when the first modal data and the second modal data are of different classes.
3. In claim 1, The learning phase is, A backbone network learning method characterized in that the first backbone network and the second backbone network are contrastively trained so that the similarity between the first feature and the second feature is lower when the first modal data and the second modal data are of different classes, and the similarity between the first feature and the second feature is lower when the first modal data and the second modal data are of the same class.
4. In claim 1, A step of augmenting the first modal data and adding the first' modal data; A step of extracting a first feature from the first modal data using a first backbone network; A backbone network learning method, further comprising: a step of performing contrastive learning on a first backbone network and a second backbone network so that the similarity between the first feature and the first' feature is higher than the similarity between the first feature and the second feature.
5. In claim 4, A step of adding second' modal data by augmenting the second modal data; A step of extracting second features from second modal data using a second backbone network; A backbone network learning method, characterized in that it further includes a step of performing contrastive learning on a first backbone network and a second backbone network so that the similarity between the second feature and the second' feature is higher than the similarity between the first feature and the second feature.
6. In claim 5, A step of adding first" modal data by augmenting the first modal data with an augmentation technique different from the augmentation technique applied to add the first' modal data; A step of extracting a first feature from the first modal data using a first backbone network; A backbone network learning method, further comprising: a step of performing contrastive learning on a first backbone network and a second backbone network so that the similarity between the first feature and the first" feature is higher than the similarity between the first feature and the second feature.
7. In claim 1, The first modal data is, One of image data, mesh data, point cloud data, volume data, and text data, The second modal data is, A backbone network learning method characterized by having at least one of image data, mesh data, point cloud data, volume data, and text data.
8. In claim 1, A step of extracting features from the first inference target modal data using the first backbone network; A step of extracting features from the second inference target modal data using a second backbone network; A backbone network learning method, characterized in that it further includes a step of calculating the similarity between the extracted features and confirming the similarity between the first inference target modal data and the second inference target modal data.
9. In claim 8, A backbone network learning method, characterized in that it further includes a step of analyzing first inference target modal data and second inference target modal data based on the confirmed similarity.
10. A processor that extracts a first feature from the first modal data using a first backbone network, extracts a second feature from the second modal data having a different modal type from the first modal data using a second backbone network, and performs comparative learning on the first backbone network and the second backbone network based on the classes of the first modal data and the second modal data; and A backbone network learning system characterized by including a storage unit that provides storage space required by a processor.
11. A step of extracting features from the first inference target modal data using the first backbone network; A step of extracting features from the second inference target modal data using a second backbone network; A step of calculating the similarity between the extracted features and confirming the similarity between the first inference target modal data and the second inference target modal data; The first backbone network and the second backbone network are, A cross-modal similarity comparison method characterized in that a first backbone network is contrastively trained based on the classes of first modal data from which features are extracted and a second backbone network is contrastively trained based on the classes of second modal data from which features are extracted.
12. A processor that extracts features from the first inference target modal data using the first backbone network, extracts features from the second inference target modal data using the second backbone network, and calculates similarity between the extracted features to confirm similarity between the first inference target modal data and the second inference target modal data; and A storage unit that provides storage space required by the processor; The first backbone network and the second backbone network are, A cross-modal similarity comparison system characterized in that a first backbone network is contrastively trained based on the classes of first modal data from which features are extracted and a second backbone network is contrastively trained based on the classes of second modal data from which features are extracted.
Citation Information
Patent Citations
A Data Feature Learning Method for Modal Incomplete Alignment
CN113033438B
Manufacturing system of display device and manufacturing method of display device using the same
KR1020250031985A
Method for training multi-modal data matching degree calculation model, method for calculating multi-modal data matching degree, and related apparatuses
US20230215136A1
Cross-modal retrieval method and related device
WO2023168997A1