Generating a descriptor of second modality data associated with first modality data

By using encoder and converter to convert data descriptors of different modes in autonomous vehicles, the scenario matching problem of unpaired data is solved, and effective data retrieval and operation efficiency are improved.

JP2025515381AActive Publication Date: 2025-05-14オクサ オートノミー リミテッド
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024564621
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-03
Filing Date
2023-05-02
Publication Date
2025-05-14
Estimated Expiration
2043-05-02

AI Technical Summary

Technical Problem

In autonomous vehicles, when data of different sensor modes are not paired, it is difficult to determine which data corresponds to which scenario, thereby affecting the effective retrieval and use of data.

Method used

By receiving data from different modalities and using trained encoders and converters, one modal descriptor is converted into another modal descriptor, and then stored and retrieved in the database.

Benefits of technology

Comparability and matching between data of different modality is achieved, the scenario matching problem of unmatched data is solved, and the data retrieval and operation efficiency of autonomous driving cars is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025515381000001_ABST
    Figure 2025515381000001_ABST
Patent Text Reader

Abstract

The present invention relates to a computer-implemented method of generating a descriptor associated with data of a first modality, the method comprising the steps of receiving first data associated with a first modality and second data related to a second modality, the first and second modalities being different, generating respective first and second descriptors corresponding to the respective first and second modalities by encoding the first and second data using respective first and second encoders, the first and second encoders being trained based on first training data of the first modality and second training data of the second modality, respectively, converting the first descriptor into a third descriptor, the third descriptor corresponding to the second modality, and storing the third descriptor in a database.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The subject matter of the present disclosure relates to a computer-implemented method for generating a descriptor for data of a first modality, a computer-implemented method for training a descriptor generator for generating a descriptor for data of a first modality, a computer-implemented method for retrieving data of a first modality, and a transitory or non-transitory computer-readable medium. [Background technology]

[0002] An autonomous vehicle (AV) includes a variety of sensors of different sensor types. Each sensor type records data of a different modality. Different modalities can be paired when they are captured simultaneously by sensors on a particular AV. For example, paired data can be data of two or more modalities that captured the same scene. The paired data can be labeled as such. The paired data can be retrieved for various purposes, such as for comparison with current data to control the AV. Summary of the Invention [Problem to be solved by the invention]

[0003] However, some AVs may not capture paired data, in which case data extraction becomes problematic as it is not possible to identify which data corresponds to which scene if the data from different modalities are individually labeled.

[0004] It is an object of the present disclosure to address such problems and improve upon the prior art. [Means for solving the problem]

[0005] According to one aspect, a computer-implemented method for generating a descriptor associated with data of a first modality is provided, the method including receiving first data associated with a first modality and second data related to a second modality, the first and second modalities being different, generating respective first and second descriptors corresponding to the respective first and second modalities by encoding the first and second data with respective first and second encoders, the first and second encoders being trained based on first training data of the first modality and second training data of the second modality, respectively, converting the first descriptor into a third descriptor, the third descriptor corresponding to the second modality, and storing the third descriptor in a database.

[0006] The foregoing aspects may also be expressed differently as a computer-implemented method of generating a descriptor associated with first modality data recorded by a first modality sensor, wherein the autonomous vehicle comprises a first modality sensor and a second modality sensor, the method including recording first data associated with the first modality by the first modality sensor and recording second data related to the second modality by the second modality sensor, the first and second modalities being different; and recording the first and second data, respectively. generating respective first and second descriptors corresponding to the respective first and second modalities by encoding using first and second encoders, the first and second encoders being trained based on first training data of the first modality and second training data of the second modality, respectively; converting the first descriptors into third descriptors, the third descriptor corresponding to the second modality; and storing the third descriptors in a database.

[0007] In one embodiment, the converting step may include decoding the first descriptor into decoded third data corresponding to the second modality using a decoder, the decoder being trained to generate data of the second modality, and encoding the third data to generate the third descriptor corresponding to the second modality using a third encoder.

[0008] In one embodiment, the first modality and the second modality may each correspond to the modality of a sensor used to record the respective data.

[0009] In one embodiment, the sensor may comprise one of a LiDAR sensor, a RADAR sensor, and a camera. In other words, the sensor of the first modality and the sensor of the second modality may each comprise one of a LiDAR sensor, a RADAR sensor, and a camera.

[0010] In one embodiment, the first encoder and the second encoder may be from respective autoencoders, and the respective first descriptor and second descriptor may be representations corresponding to the bottlenecks of the respective autoencoders.

[0011] In one embodiment, the autoencoder may be a variational autoencoder.

[0012] According to another aspect of the invention, there is provided a computer-implemented method of training a descriptor generator for generating a descriptor associated with data of a first modality, the descriptor generator comprising a first encoder, a second encoder, and a transformer, the method including: training the first encoder to generate a first descriptor corresponding to the first data of the first modality using first training data of the first modality, training the second encoder to generate a second descriptor corresponding to the second data of the second modality using second training data of the second modality, the first and second modalities being different, and training the transformer to transform the first descriptor into a third descriptor corresponding to the second modality using contrastive learning based on the first and second training data.

[0013] In one embodiment, training the transformer using contrastive learning may include determining a distance between the second and third descriptors, and modifying the transformer to reduce the distance.

[0014] In one embodiment, the modifying step may include modifying the transformer to minimize the distance.

[0015] In one embodiment, training the transformer using contrastive learning may include comparing the second descriptor and the second descriptor using a discriminator, and determining whether the third descriptor is real or fake using the discriminator.

[0016] The descriptor generator may further comprise an inverse transformer, and the method may further include the steps of transforming the third descriptor into a fourth descriptor using the inverse transformer, calculating a distance between the first and fourth descriptors, and modifying the inverse transformer to reduce the distance.

[0017] In one embodiment, the first modality and the second modality may each correspond to a modality used to record the respective data.

[0018] In one embodiment, the sensor may comprise one of a LiDAR sensor, a RADAR sensor, and a camera.

[0019] In one embodiment, the first encoder and the second encoder may each be an autoencoder.

[0020] In one embodiment, the autoencoder may be a variational autoencoder.

[0021] According to another aspect of the invention, there is provided a computer-implemented method of retrieving data of a first modality, the method including retrieving a descriptor from a database of descriptors associated with data of a second modality, transforming the retrieved descriptor into a transformed descriptor using an inverse transformer, the transformed descriptor being associated with the first modality, the first and second modalities being different, and retrieving data of the first modality associated with the transformed descriptor from the database.

[0022] The above aspects may also be expressed differently as a computer-implemented method of operating and moving an autonomous vehicle, the method including obtaining real-time data from a sensor of a first modality; retrieving a descriptor from a database of descriptors associated with data of a second modality; converting the retrieved descriptor into a transformed descriptor using an inverse transformer, the transformed descriptor being associated with the first modality, the first and second modalities being different; decoding the transformed descriptor of the first modality with a decoder to generate data of the first modality; comparing the real-time data with the generated data of the first modality; and operating and moving the autonomous vehicle based on the comparison.

[0023] In one embodiment, the first and second modalities may each correspond to a modality of a sensor used to record the respective data.

[0024] In one embodiment, the sensor may be one of a LiDAR sensor, a RADAR sensor, or a camera.

[0025] In one embodiment, the method may further include obtaining real-time data from a sensor of the first modality, comparing the real-time data to the retrieved data, and operating and moving the autonomous vehicle based on the comparison.

[0026] According to another aspect of the invention, there is provided a transitory or non-transitory computer readable medium having stored thereon instructions which, when executed by a processor, cause the processor to perform a method according to any one of the preceding claims.

[0027] The subject matter of the present disclosure is best understood with reference to the accompanying drawings. [Brief description of the drawings]

[0028] [Figure 1]FIG. 1 is a schematic diagram of an AV according to one or more embodiments. [Diagram 2] A block diagram of variational autoencoders, one for each modality. [Diagram 3] FIG. 3 is a more detailed block diagram of the variational autoencoder of FIG. 2. [Figure 4] FIG. 3 is a block diagram of a classifier used to train an encoder from the variational autoencoder of FIG. 2 using paired data. [Diagram 5] FIG. 3 is a block diagram of a distance comparator used to train an autoencoder from the variational autoencoder of FIG. 2 using paired data. [Figure 6] FIG. 3 is a block diagram illustrating training of a decoder from the variational autoencoder of FIG. [Figure 7] FIG. 3 is a block diagram for training a transformer to transform a descriptor from the bottleneck of the variational autoencoder of FIG. 2 and training an inverse transformer to undo the transformation of the descriptor using unpaired data, in accordance with one or more embodiments. [Figure 8] FIG. 8 is a block diagram of the transformer of FIG. 7. [Figure 9] FIG. 8 is a block diagram of the inverse transformer of FIG. [Figure 10] FIG. 3 is a block diagram for training a transformer to transform a descriptor from the bottleneck of the variational autoencoder of FIG. 2 and training an inverse transformer to undo the transformation of the descriptor using paired data, in accordance with one or more embodiments. [Figure 11] FIG. 1 is a block diagram for storing a translated descriptor and updating the index of the closest descriptor with the translated descriptor, according to one or more embodiments. [Figure 12] FIG. 2 is a block diagram illustrating retrieval of data using translated descriptors in accordance with one or more embodiments. [Figure 13] 1 is a flowchart of a computer-implemented method for generating a descriptor associated with data of a first modality. [Figure 14] 1 is a flowchart of a computer-implemented method for training a descriptor generator to generate descriptors associated with data of a first modality. [Figure 15] 1 is a flowchart of a computer-implemented method for retrieving data of a first modality. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0029] The embodiments described herein are embodied as a set of instructions stored as electronic data on one or more storage media. Specifically, these instructions may be provided on a transitory or non-transitory computer-readable medium. When executed by a processor, the processor is configured to perform various methods described in the following embodiments. As such, these methods may be computer-implemented methods. Specifically, the processor and the storage containing the instructions may be incorporated into a vehicle. The vehicle may be an AV.

[0030] The following embodiments provide specific illustrative examples, which should not be construed as limiting, and the scope of protection is defined by the claims. Features of certain embodiments may be used in combination with features of other embodiments without extending the subject matter beyond the content of the present disclosure.

[0031] With reference to FIG. 1 , the AV 10 may include a number of sensors 12. The sensors 12 may be mounted on the roof of the AV 10. The sensors 12 may be communicatively connected to a computer 14. The computer 14 may be on-board the AV 10. The computer 14 may include a processor 16 and a memory 18. The memory may include the non-transitory computer-readable medium described above. Alternatively, the non-transitory computer-readable medium may be located remotely and communicatively linked to the computer 14 via a cloud 20. The computer 14 may be communicatively linked to one or more actuators 22 to control the one or more actuators 22 to move the AV 10. The actuators may include, for example, a motor, a braking system, a power steering system, and the like.

[0032] The sensors 12 may include a variety of sensor types. Examples of sensor types include LiDAR sensors, RADAR sensors, and cameras. Each sensor type may be referred to as a sensor modality. Each sensor type may record data related to the sensor modality. For example, a LiDAR sensor may record LiDAR modality data.

[0033] The data may capture various scenes encountered by the AV 10. For example, a scene may be the visible scene surrounding the AV 10 and may include roads, buildings, weather, objects (e.g., other vehicles, pedestrians, animals, etc.), etc.

[0034] Referring to FIG. 2, for each modality, an autoencoder 24 is trained. The autoencoder may be a variational autoencoder (VAE). The VAE 24 may be configured to receive packets 26 of data of a particular modality, such as image data or LiDAR data. A scene may be captured in each data packet 26. The VAE 24 may include an encoder 28 and a decoder 30. The encoder may reduce the dimensionality of the data packet 26 to a distribution of a bottleneck 32. This distribution may be referred to as a descriptor 34. The decoder 30 may reconstruct the data packet 26 by increasing the dimensionality of the descriptor 34 to the same size as the original data packet 26.

[0035] To train the VAE 24, the error between the reconstructed data packets and the original data packets is calculated. Hyperparameters of the encoder and decoder may be optimized using backpropagation.

[0036] The descriptor 34 may be stored in a database. The descriptor 34 may be used to regenerate a scene corresponding to the data packet associated with the descriptor 34. The regenerated scene may be used by the AV 10 (FIG. 1) when moving the AV 10. For example, the regenerated scene may be associated with a particular operational mode. Current sensor data may be compared to the regenerated scene, and based on how closely the current scene data matches the regenerated scene data, the AV 10 may be controlled to operate in an associated operational mode. The regenerated scene data may also be used in a simulator to enrich the scenarios that the AV 10 is trained to handle in real time.

[0037] Referring to FIG. 3, the VAE 24 may receive data packets 26 from sensors. Additionally, the VAE 24 may receive synthesized data packets 26′. The synthesized data packets 26′ may be data packets in which features in the original sensor data packets have been augmented, or data from a simulation. For example, additional objects such as pedestrians or vehicles, or weather conditions such as rain or fog, may be added using a generative adversarial network. In this way, the VAE 24 may have a larger data train from which to train.

[0038] Typically, the VAE 24 may generate a descriptor with a particular distribution, which may be a Gaussian distribution. However, in one or more embodiments, the distribution may not be a Gaussian distribution. Depending on the features in the scene of the data packet, the distribution may be different. For example, there may be a distribution where there are two dynamic objects in the scene, a particular distribution where there is a stationary vehicle in the scene, and a particular distribution where there is a particular weather condition such as fog. A look-up table, called a codebook 36, may be provided. The codebook includes all the distributions. A hidden layer in the encoder 28 may be extracted, which includes the features in the scene extracted from the original data packet 26, 26'. The features in the hidden layer may be used to identify one or more distributions from the look-up table. If multiple features are present, the multiple distributions may be statistically combined to obtain a distribution for a particular scene.

[0039] Since each VAE 24 is trained independently, the descriptors are modality-specific for each modality of the VAE 24. Thus, there is no correlation between the descriptors. In other words, it is difficult to compare descriptors derived from data of different modalities. For example, it is possible to compare a descriptor derived from image data with a descriptor derived from LiDAR data. However, using, for example, cosine similarity or Euclidean distance, matching scenes may have a relatively short or non-zero distance.

[0040] With reference to Figures 4 and 5, contrastive learning can be used to reduce the discrepancy between the descriptors of different modalities.

[0041] With reference to FIG. 4, contrastive learning may be implemented by training each encoder 28 using a classifier 38 as a supervisor. For example, a first modality scene 26A, such as an image, may be input to a first encoder 28A. A second modality scene 26B, such as LiDAR data, may be input to a second encoder 28B. The second modality scene may also be input to a third encoder 28C. The inputs to the second and third encoders 28A, 28B may differ in that the input of the second encoder may positively match the first modality scene, while the scene input to the third encoder may not match or negatively match the first modality scene. The positive and negative matches are known because the first and second modality data are paired. By "paired" we mean that they represent the same scene and may be synchronized in time when captured.

[0042] The classifier 38 is used to train the parameters of the first and second encoders 28A, 28B, respectively. Training means that the first and second descriptors 34A, 34B derived from the first and second encoders 28A, 28B, respectively, categorize the scene in the same way. Furthermore, a third descriptor 34C derived from a third encoder 28C categorizes the scene differently.

[0043] Referring to FIG. 5, contrastive learning may be performed using first and second distance matchers 40, 42.

[0044] The first distance matcher 40 may calculate the distance between the first and second descriptors 34A, 34B. The first and second encoders 28A, 28B are trained to derive the first and second descriptors 34A, 34B that have zero or at least a very small or negligible distance. The first distance matcher 40 may calculate the distance using Euclidean distance, cosine similarity, a trained matcher (e.g., a neural network), or other distance metric.

[0045] The second distance matcher 42 may calculate the distance between the second and third descriptors 34B, 34C. The second and third encoders 28B, 28C are trained to derive the second and third descriptors 34B, 34C with a very large distance that may tend to infinity. The second distance matcher 42 may calculate the distance using Euclidean distance, cosine similarity, a trained matcher (e.g., a neural network), or other distance metric.

[0046] FIG. 6 shows a block diagram illustrating how each decoder is trained in addition to training the encoder using the contrastive learning shown in FIG.

[0047] For example, there are first to third decoders 30A to 30C to be trained. Each decoder 30A to 30C corresponds to a first to third encoder 28A to 28C, respectively. Each decoder is trained by comparing a reconstructed scene of a given modality with an original scene 26 of that modality and modifying the hyper-parameters of the respective decoder to reduce or minimize the error.

[0048] The algorithms described with reference to Figures 4-6 work well when the data is paired, however when the data is unpaired, e.g. from a single modality, it may not be possible to make the descriptors comparable when obtained from different modalities.

[0049] To alleviate this problem, an embodiment of a descriptor generator 45 according to FIG. 7 is provided.

[0050] Referring to Fig. 7, the descriptor generator 45 includes a first encoder 28A, a second encoder 28B, a transformer 46, a discriminator 48, an inverse transformer 50, and a distance matcher 52. The first training data 26A (e.g., sample scenes) and the second training data 26B (e.g., sample scenes) of the first and second modalities 26A, 26B, respectively, are input to the first and second encoders 28A, 28B. The first and second encoders 28A, 28B are trained as described above with reference to Figs. 2 and 3. In other words, the first and second encoders 28A, 28B are trained independently. The first encoder 28A uses the first training data of the first modality to generate a first descriptor 34_1 corresponding to the first data 26A of the first modality. The second encoder 28B uses the second training data of the second modality to generate a second descriptor 34_2 corresponding to the second data 26B of the second modality. The first modality and the second modality may differ in that they are acquired by different sensor types. For example, the first modality may be LiDAR sensor data and the second modality may be camera image data.

[0051] The transformer 46 may be trained to transform the first descriptor 34_1 into a third descriptor 34_3, which may correspond to a second modality. In other words, the transformer 46 is configured to transform the descriptor from the first modality to the second modality such that the third descriptor 34_3 appears as if it was generated by the first encoder 28A instead of the second encoder 28B.

[0052] With brief reference to Fig. 8, the transformer 46 includes a decoder 53 and an encoder 54. The decoder 53 is configured to generate third data of a first modality, e.g. a camera image. The encoder 54 is configured to generate a third descriptor 34_3. The third descriptor 34_3 thus corresponds to the second modality.

[0053] With further reference to FIG. 7, the descriptor generator 45 uses a discriminator 48 to compare the second and third descriptors 34_2 and 34_3. The discriminator 48 is configured to determine whether the third descriptor 34_3 is genuine or fake. For example, a genuine determination means that the transformer 46 has successfully generated a third descriptor 34_3 that the discriminator 48 believes is derived directly from the second modality data sample 26B, rather than a transformed form of the first modality data sample 26A. A fake determination means that the transformer 46 has generated a third descriptor 34_3 that the discriminator 48 does not believe is derived directly from the second modality data 26B. The discriminator 48 may be a neural network.

[0054] If the discriminator 48 determines that the third descriptor 34_3 is fake, the transformer 46 is changed. Specifically, the parameters of the transformer 46 are modified until the third descriptor 34_3 is determined to be a genuine descriptor.

[0055] The inverse transformer 50 may transform the third descriptor 34_3 into a fourth descriptor 34_4.

[0056] 9, the inverse transformer 50 may include a decoder 56 and an encoder 58. The decoder 56 may generate data of a first modality, such as LiDAR data, using a third descriptor corresponding to a second modality, such as a camera image. The encoder 58 may generate a fourth descriptor 34_4 from the generated data of the second modality.

[0057] With further reference to FIG. 7, the distance matcher 52 is configured to calculate the distance between the first and fourth descriptors. The distance matcher 52 may perform distance matching using any of Euclidean distance, cosine similarity, a trained matcher (e.g., neural network), or any other suitable distance metric. The distance may be interpreted as an error value. The inverse transformer 50 may be modified to reduce or minimize the distance. In certain circumstances, the transformer 46 may also be modified to reduce or minimize the distance.

[0058] One reason for using the distance matcher 52 is to avoid the situation where the transformer 46 produces a third descriptor 34_3 that is deemed authentic but has different characteristics than the original data 26A. For example, the original image may contain three pedestrians and four moving vehicles. If the third descriptor 34_3 produces an image that contains no pedestrians and only two animals and a stationary car, then the third descriptor 34_3 is deemed authentic in that it is derived from the second modality.

[0059] In this manner, data of a first modality can be indexed using a descriptor of a target modality, such as a second modality different from the first modality. Thus, all modalities may be indexed using the descriptors translated to the target modality. As a result, data does not need to be paired data if derived from different sensor types in order to be indexed and compared to data of a similar or identical scene.

[0060] With reference to Figure 10, one embodiment of a descriptor generator 145 is provided. The descriptor generator 145 is the same as that of Figure 7, except that the discriminator 48 has been replaced by a distance matcher 160. To avoid repetition, only the differences between the embodiment of Figure 7 and the embodiment of Figure 10 will be described, with similar features being prefixed with a 1.

[0061] The third descriptor 134_3 is compared to the second descriptor 134_2 by a distance matcher 160. The distance matcher 160 may calculate the distance using Euclidean distance, cosine similarity, a trained matcher (e.g., a neural network), or any suitable matching metric.

[0062] The transformer 146 may be modified to shorten or minimize the distance based on the distance between the second and third descriptors 134_2, 134_3. For example, parameters of the transformer 146 may be modified.

[0063] Referring to FIG. 11, a descriptor generator 45 is shown generating a descriptor for indexing the first modality data.

[0064] First data 26A associated with a first modality (e.g., LIDAR sensor data) may be used as input to a first encoder 28A. Similar to the above embodiment, the first encoder 28A generates a first descriptor 34_1. A transformer 46 converts the first descriptor into a third descriptor 34_3, which corresponds to a second modality. The third descriptor 34_3 is stored in a database 62. The database 62 also includes other descriptors, all of which correspond to the second modality, either directly derived from the sensor data of the second modality or derived by transformation from another modality to the second modality.

[0065] A distance matcher 64 receives the third descriptor directly. The third descriptor 34_3 is compared to the descriptors in the database 62 by the distance matcher 64. The comparison may be performed using techniques such as Euclidean distance, cosine similarity, trained matchers such as neural networks, etc. The result of the distance matcher 64 is input to the index 66 of the closest descriptor.

[0066] 12, the requesting system 68 may request a scene in a first modality. The requesting system may identify the scene using, for example, the index shown in FIG. 11. The requesting system 68 may retrieve a third descriptor 34_3 from the database of descriptors 62. The third descriptor 34_3 is associated with data of a second modality, for example camera image data.

[0067] The inverse transformer 50 may transform the third descriptor 34_3 into a fourth descriptor 34_4 of the first modality. The fourth descriptor 34_4 may be referred to as a transformed descriptor. The transformed descriptor 34_4 may be of the first modality. The fourth descriptor 34_4 may be input to a decoder 56, which may be the decoder of FIG. 9. The decoder 56 may generate a reconstructed image 70A of the first modality (e.g., LiDAR data).

[0068] During run time, computer 14 (FIG. 1) may obtain real-time data from a sensor 12 of a first modality (e.g., a LiDAR sensor). Computer 14 may compare the real-time data to the retrieved data. In this manner, computer 14 may be a requesting system 68. Computer 14 may operate and move AV 10 based on the comparison. For example, AV 10 may have previously passed through a very similar scene, so may use the previously used trajectory to construct a trajectory for navigating the current scene.

[0069] While the above description provides details regarding the embodiments, the subject matter of the present disclosure may be summarized with reference to FIGS.

[0070] 13, a computer-implemented method for generating a descriptor associated with data of a first modality is provided, the method including receiving (step S100) first data associated with a first modality and second data associated with a second modality, the first and second modalities being different, generating (step S102) respective first and second descriptors corresponding to the respective first and second modalities by encoding the first and second data using respective first and second encoders, the first and second encoders being trained based on first training data of the first modality and second training data of the second modality, respectively, converting (step S104) the first descriptor into a third descriptor, the third descriptor corresponding to the second modality, and storing (step S106) the third descriptor in a database.

[0071] 14, a computer-implemented method for training a descriptor generator for generating a descriptor for indexing data of a first modality is provided, the descriptor generator comprising a first encoder, a second encoder, and a transformer, the method including: training the first encoder to generate a first descriptor corresponding to a first data of the first modality using first training data of the first modality (step S200); training the second encoder to generate a second descriptor corresponding to a second data of the second modality using second training data of the second modality (step S202), the first and second modalities being different; and training the transformer to transform the first descriptor into a third descriptor corresponding to the second modality using contrastive learning based on the first and second training data (step S204).

[0072] 15, a computer-implemented method for retrieving data of a first modality is provided, the method including retrieving a descriptor from a database of descriptors associated with data of a second modality (step S300), converting the retrieved descriptor into a transformed descriptor using an inverse transformer (step S302), the transformed descriptor being associated with the first modality, the first and second modalities being different, and retrieving data of the first modality associated with the transformed descriptor from the database (step S304).

[0073] The foregoing embodiments are described to illustrate the subject matter of the present disclosure, but the features of the embodiments should not be construed as limiting the scope of protection, which, for the avoidance of doubt, is defined by the following claims.

Claims

1. 1. A computer-implemented method for generating a descriptor associated with data of a first modality, comprising: receiving first data associated with a first modality and second data associated with a second modality, the first and second modalities being different; generating respective first and second descriptors corresponding to respective first and second modalities by encoding the first and second data using respective first and second encoders, the first and second encoders being trained based on first training data of the first modality and second training data of the second modality, respectively; transforming the first descriptor into a third descriptor, the third descriptor corresponding to the second modality; storing said third descriptor in a database; 4. A computer-implemented method comprising:

2. The converting step includes: decoding the first descriptor into decoded third data corresponding to the second modality using a decoder, the decoder being trained to generate data of the second modality; encoding the third data to generate the third descriptor corresponding to the second modality using a third encoder; The computer-implemented method of claim 1 , comprising:

3. 3. The computer-implemented method of claim 1 or claim 2, wherein the first modality and the second modality respectively correspond to modalities of sensors used to record the respective data.

4. The computer-implemented method of claim 3 , wherein the sensor comprises one of a LiDAR sensor, a RADAR sensor, and a camera.

5. 5. The computer-implemented method of claim 1, wherein the first encoder and the second encoder are from respective autoencoders, and the respective first and second descriptors are representations corresponding to a bottleneck of the respective autoencoder.

6. The computer-implemented method of claim 5 , wherein the autoencoder is a variational autoencoder.

7. 1. A computer-implemented method of training a descriptor generator for generating a descriptor associated with data of a first modality, the descriptor generator comprising a first encoder, a second encoder, and a transformer, the method comprising: training the first encoder to generate a first descriptor corresponding to first data of a first modality using first training data of a first modality; training the second encoder to generate a second descriptor corresponding to second data of a second modality using second training data of a second modality, the first and second modalities being different; training the transformer to transform the first descriptor into a third descriptor corresponding to the second modality using contrastive learning based on the first and second training data; 4. A computer-implemented method comprising:

8. The step of training the transformer using contrastive learning comprises: determining a distance between the second and third descriptors; modifying the transformer to reduce the distance; The computer-implemented method of claim 7 , comprising:

9. The step of training the transformer using contrastive learning comprises: comparing the first descriptor and the second descriptor using a discriminator; using the discriminator to determine whether the third descriptor is genuine or fake; The computer-implemented method of claim 7 , comprising:

10. The descriptor generator further comprises an inverse transformer, the method comprising: transforming the third descriptor into a fourth descriptor using the inverse transformer; calculating a distance between the first and fourth descriptors; modifying the inverse transformer to reduce the distance; The computer-implemented method of any one of claims 7 to 9, further comprising:

11. The computer-implemented method of any one of claims 7 to 10, wherein the first modality and the second modality correspond to modalities used to record the respective data.

12. 12. The computer-implemented method of claim 11, wherein the sensor comprises one of a LiDAR sensor, a RADAR sensor, and a camera.

13. The computer-implemented method of any one of claims 7 to 12, wherein the first encoder and the second encoder are each an autoencoder encoder.

14. The computer-implemented method of claim 13 , wherein the autoencoder is a variational autoencoder.

15. 1. A computer-implemented method for retrieving data of a first modality, comprising: Retrieving a descriptor from a database of descriptors associated with the second modality data; transforming the retrieved descriptor into a transformed descriptor using an inverse transformer, the transformed descriptor being associated with the first modality, the first and second modalities being different; retrieving the first modality data associated with the transformed descriptor from a database; 4. A computer-implemented method comprising:

16. 16. The computer-implemented method of claim 15, wherein the first and second modalities respectively correspond to modalities of sensors used to record the respective data.

17. 20. The computer-implemented method of claim 16, wherein the sensor is one of a LiDAR sensor, a RADAR sensor, or a camera.

18. acquiring real-time data from a sensor of the first modality; comparing the real-time data with the retrieved data; operating and moving the autonomous vehicle based on the comparison; The computer-implemented method of any one of claims 15 to 17, further comprising:

19. A transitory or non-transitory computer readable medium having stored thereon instructions which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 18.

Citation Information

Patent Citations

  • Method and system for case-based computer-aided diagnosis using cross-modality

    JP2011508917A

  • Object recognition apparatus, object recognition system, learning method of object recognition apparatus, object recognition method of object recognition apparatus, learning program of object recognition apparatus, and object recognition program of object recognition apparatus

    JP2022066879A

  • Generating cross-domain data using variational mapping between embedding spaces

    US20190318040A1