Method and system for accessing data set
By using embedded networks to embed and train and query embedded networks in large data sets related to autonomous driving systems, the data set accessibility problem is solved, and the data set is quickly accessed and full-text search functions are realized.
Patent Information
- Application Number
- CN202411893250.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-20
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art is difficult to effectively improve the accessibility of large data sets related to autonomous driving systems, especially in the case of multi-data sources and complex data models.
Full-text search and retrieval of the data set is achieved by embedding sensor data and auxiliary data in a multidimensional vector space using an embedded network and training the query embedding network to associate it with the sensor data embedding network.
It realizes fast, convenient and effective access to large data sets related to autonomous driving systems, and can extract data related to specific scenarios from multiple data sources, improving the searchability and understanding of data.
Smart Images

Figure CN120179688A_ABST
Abstract
Description
Technical Field
[0001] The disclosed technology relates to methods and systems for accessing a dataset that includes multiple data samples from multiple data sources. Specifically but not exclusively, the disclosed technology relates to methods and systems for improving the accessibility and understanding of large datasets having data samples collected by an autonomous driving system (ADS) and associated auxiliary data samples. Background Art
[0002] Machine learning (ML) algorithms and neural networks (NN) have gained a strong position in solving complex problems in various fields such as classification, detection, recognition, and segmentation tasks. The ability of these algorithms to perform complex and multi-dimensional tasks involving almost infinite data variables and combinations makes these models well-suited for today's expanding big data applications. A particular area where neural networks and deep learning models have presented pioneering applications is the emergence of vehicle autonomous driving systems.
[0003] In the past few years, the number of research and development activities related to autonomous vehicles has surged, and many different methods are being explored. An increasing number of modern vehicles are equipped with advanced driver assistance systems (ADAS) to improve vehicle safety and more broadly road safety. ADAS is an electronic system that can provide assistance to a vehicle's driver. For example, it can be represented by adaptive cruise control (ACC), collision avoidance systems, forward collision warning systems, etc. Nowadays, many technical fields related to the ADAS and autonomous driving (AD) areas are under research and development. In this article, ADAS and AD are collectively referred to as an autonomous driving system (ADS), which corresponds to all levels in different automation levels defined, for example, by the SAE J3016 levels (0 - 5) of driving automation.
[0004] ADS solutions have found their way into most new cars on the market, and their application prospects will only rise in the future. ADS can be understood as a complex combination of various components and a system for introducing automation into road traffic, where the components can be defined as systems that perform or cooperate with a human driver to perform the vehicle's perception, decision-making, and operation by electronic and mechanical means instead of a human driver. This includes the vehicle's maneuvering, destination, and perception of the surrounding environment. While having control of the vehicle, the automation system allows the human operator to retain full or at least partial responsibility for the system. ADS typically combines various sensors (such as radar, lidar, sonar, cameras, navigation systems (such as GPS), odometers, and / or inertial measurement units (IMU)) to sense the vehicle's surrounding environment, and an advanced control system can interpret the sensed information based on the surrounding environment to identify suitable navigation paths as well as obstacles, free space areas, and / or relevant signs.
[0005] An important aspect of achieving reliable ADS functionality for a vehicle of interest is obtaining a comprehensive understanding of the scenarios occurring in the vehicle's surrounding environment. The unpredictable dynamic scenarios of the situations, events, or objects in the vehicle's surrounding environment, including those on the road on which the vehicle is traveling, can involve almost endless variety and complexity. In other words, to achieve reliable autonomous driving functionality, a large amount of data (sensor data recorded by the vehicle, output data of various ADS functions, metadata, etc.) is required.
[0006] As a result of this need for large amounts of data for the development and validation of ADS functionality, there is ultimately an inevitable need for large databases, which are generally difficult to process effectively and often rely on the user's expertise and knowledge of the data models in the databases for accessing relevant information. Accordingly, there is a need for a data model that allows for improved accessibility to the data of large databases used for ADS development and validation. SUMMARY OF THE INVENTION
[0007] The techniques disclosed herein are directed to alleviating, mitigating, or eliminating one or more of the deficiencies and drawbacks identified in the prior art above, to address various issues related to the accessibility of large data sets associated with autonomous driving systems.
[0008] Aspects and embodiments of the disclosed techniques are defined below and in the independent and dependent claims.
[0009] The first aspect of the disclosed technology includes: a method for accessing a dataset that includes multiple data samples from multiple data sources. The multiple data samples include sensor data samples captured by one or more vehicles, where the sensor data samples include information about the vehicle's surrounding environment, and each sensor data sample is represented by a corresponding sensor data embedding, which is generated by processing the sensor data sample via a sensor data embedding network that has been trained to process the sensor data sample and output a corresponding sensor data embedding for each sensor data sample in a multi-dimensional vector space. The multiple data samples further include auxiliary data samples, where each auxiliary data sample is represented by a corresponding auxiliary data embedding, which is generated by processing the auxiliary data sample via an auxiliary data embedding network that has been trained to process the auxiliary data sample and output a corresponding auxiliary data embedding in the multi-dimensional vector space. Additionally, the auxiliary data embedding network has been trained in association with the sensor data embedding network such that the auxiliary embedding of the auxiliary data sample associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample. The method includes, in response to obtaining a query embedding, identifying one or more embeddings in the multi-dimensional vector space based on proximity to the obtained query embedding. The query embedding is generated by processing the query via a query embedding network that has been trained to process the query and output a corresponding query embedding for each query in the multi-dimensional vector space. Additionally, the query embedding network has been trained in association with the sensor data embedding network such that the query embedding of the query associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample. The method further includes outputting one or more data samples within the dataset represented by the identified embeddings.
[0010] The second aspect of the disclosed technology includes a computer program product comprising instructions that, when executed by a computing device, cause the computing device to perform the method according to any of the embodiments of the first aspect disclosed herein. For this aspect of the disclosed technology, there are similar advantages and preferred features as the other aspects.
[0011] The third aspect of the disclosed technology includes a (non-transitory) computer-readable storage medium comprising instructions that, when executed by a computing device, cause the computing device to perform the method according to any of the embodiments of the first aspect disclosed herein. For this aspect of the disclosed technology, there are similar advantages and preferred features as the other aspects.
[0012] As used herein, the term "non-transitory" is intended to describe a computer-readable storage medium (or "memory"), excluding propagating electromagnetic signals, but is not intended to otherwise limit the type of physical computer-readable storage device covered by the terms computer-readable medium or memory. For example, the term "non-transitory computer-readable medium" or "tangible memory" is intended to cover types of storage devices that do not necessarily store information permanently, including, for example, random access memory (RAM). Program instructions and data stored in a non-transitory form on a tangible computer-accessible storage medium can further be transmitted by a transmission medium or by a signal (e.g., an electrical, electromagnetic, or digital signal) that can be transmitted via a communication medium such as a network and / or a wireless link. Thus, as used herein, the term "non-transitory" is a limitation on the medium itself (i.e., tangible, rather than a signal), and does not limit data storage persistence (e.g., RAM versus ROM).
[0013] A fourth aspect of the disclosed technology includes: a system for accessing a dataset that includes multiple data samples from multiple data sources. The multiple data samples include sensor data samples captured by one or more vehicles, where the sensor data samples include information about the vehicle's surrounding environment, and each sensor data sample is represented by a corresponding sensor data embedding, which is generated by processing the sensor data sample via a sensor data embedding network that has been trained to process sensor data samples and output a corresponding sensor data embedding for each sensor data sample in a multi-dimensional vector space. The multiple data samples further include auxiliary data samples, each auxiliary data sample being represented by a corresponding auxiliary data embedding, which is generated by processing the auxiliary data sample via an auxiliary data embedding network that has been trained to process auxiliary data samples and output a corresponding auxiliary data embedding in the multi-dimensional vector space. Additionally, the auxiliary data embedding network has been trained in association with the sensor data embedding network such that the auxiliary embedding of an auxiliary data sample associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample. The system includes control circuitry configured to, in response to obtaining a query embedding, identify one or more embeddings in the multi-dimensional vector space based on proximity to the obtained query embedding within the multi-dimensional vector space. The query embedding is generated by processing a query via a query embedding network that has been trained to process queries and output a corresponding query embedding for each query in the multi-dimensional vector space. Additionally, the query embedding network has been trained in association with the sensor data embedding network such that the query embedding of a query associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample. The control circuitry is further configured to output one or more data samples within the dataset represented by the identified embeddings. For this aspect of the disclosed technology, there are similar advantages and preferred features to other aspects.
[0014] The disclosed aspects and preferred embodiments may be appropriately combined with each other in any manner that would be obvious to a person of ordinary skill in the art, such that one or more features or embodiments disclosed in relation to one aspect may also be considered to be disclosed in relation to another aspect or an embodiment of another aspect.
[0015] An advantage of some embodiments is that relevant data samples in a huge dataset collected from vehicles can be accessed in a convenient and fast manner.
[0016] An advantage of some embodiments is that data related to a particular scenario or situation, or useful in development, testing, and validation, can be extracted from a large dataset in an effective manner.
[0017] Further embodiments are defined in the dependent claims. It should be emphasized that, when used in this specification, the terms "comprising / including" are used to specify the presence of the stated features, integers, steps or components. It does not preclude the presence or addition of one or more other features, integers, steps, components or groups thereof.
[0018] These and other features and advantages of the disclosed technology will be further clarified below with reference to the embodiments described in the following text. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above aspects, features and advantages of the disclosed technology will be more fully understood by reference to the following illustrative and non - limiting detailed description of example embodiments of the present disclosure in conjunction with the accompanying drawings, in which:
[0020] Figure 1 is a schematic block diagram representation of a system for accessing a data set including multiple data samples from multiple data sources according to some embodiments.
[0021] Figure 2 is a schematic diagram of a system for accessing a data set including multiple data samples from multiple data sources according to some embodiments, a server including such a system, and a schematic diagram of a cloud environment including multiple servers.
[0022] Figure 3 is a schematic flowchart representation of a method for accessing a data set including multiple data samples from multiple data sources according to some embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] The present technology will now be described in detail with reference to the accompanying drawings, in which some example embodiments of the disclosed technology are shown. However, the disclosed technology may be embodied in other forms and should not be construed as limited to the disclosed example embodiments. The purpose of providing the disclosed example embodiments is to fully convey the scope of the disclosed technology to those skilled in the art. Those skilled in the art will understand that the steps, services and functions explained herein can be implemented using separate hardware circuits, using software working in conjunction with a programmable microprocessor or a general - purpose computer, using one or more application - specific integrated circuits (ASICs), using one or more field - programmable gate arrays (FPGAs) and / or using one or more digital signal processors (DSPs).
[0024] It should also be understood that when the present disclosure is described from a method perspective, it can also be embodied in an apparatus including one or more processors and one or more memories coupled to the one or more processors that load computer code for implementing the method. For example, in some embodiments, the one or more memories may store one or more computer programs that, when executed by the one or more processors, cause the apparatus to perform the steps, services, and functions disclosed herein.
[0025] It should also be understood that the terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting. It should be noted that, as used in the specification and claims, the articles "a", "the", and "" are intended to mean that there is one or more of the elements, unless the context clearly dictates otherwise. Thus, for example, in certain contexts, reference to "a unit" or "the unit" may refer to more than one unit, and so on. Additionally, the terms "including", "comprising" do not exclude other elements or steps. It should be emphasized that when used in this specification, the term "including / comprising" is used to specify the presence of the stated features, integers, steps, or components. It does not exclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. The term "and / or" should be interpreted as meaning both simultaneously and each as an alternative.
[0026] It should also be understood that although the terms first, second, etc. may be used herein to describe various elements or features, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first signal may be referred to as a second signal, and similarly, a second signal may be referred to as a first signal, without departing from the scope of the embodiments. The first signal and the second signal are both signals, but they are not the same signal.
[0027] Overview
[0028] As previously mentioned, an important aspect of achieving reliable ADS functionality for future vehicles is obtaining a comprehensive understanding of the scenarios occurring in the vehicle's surrounding environment. The unpredictable dynamic scenarios, including the situations, events, or objects in the vehicle's surrounding environment and on the road on which the vehicle is traveling, can involve almost endless varieties and complexities. In other words, to achieve reliable autonomous driving functionality, a large amount of data (sensor data recorded by the vehicle and its associated metadata) is required.
[0029] In addition, currently, these large amounts of sensor data are typically collected by a fleet of dedicated test vehicles (which can also be referred to as development vehicles). However, as production vehicles become more and more capable, they also provide large amounts of sensor data. In either case, large sensor data sets are collected for the development, validation, and / or testing of existing and new functions of the ADS. However, when developing, validating, and / or testing the existing and new functions of the ADS, not only is the sensor data collected from the vehicle useful, but also auxiliary data sets (weather data, map data, ADS data, DMS data, and related metadata) are required in order to be able to fully understand the field test data and improve and / or test the ADS functions.
[0030] Typically, these huge data sets are stored in traditional databases that can be searched based on static metadata searches. However, the data models used in these traditional databases are complex, which reduces the usability of the data because searching the database and extracting relevant data samples heavily rely on the expertise of the user. In addition, the task of constructing these data models and correctly storing the data samples is a difficult task, especially in the field of autonomous driving systems, because these data sets are generated from different data sources and have endless variations. More specifically, the accessibility of these data sets typically highly depends on the relevant metadata of each data sample, which is static and defined by one or more users according to some defined guidelines. However, in terms of data accessibility and retrieval, it is difficult to foresee future requirements, resulting in the risk that previous settings become obsolete and incompatible with future requirements.
[0031] To this end, the technology disclosed in this article proposes a solution to unlock at least a part of the full value of the available data sets and use embeddings to make them searchable. More specifically, it is proposed in this article to use embeddings as carriers of information regardless of the data source, where the embeddings are generated for the collected data samples and stored in the database, along with links back to the actual data samples to track their exact data sources. Thus, it can enable full-text free text and / or image searches of the entire data set.
[0032] Accordingly, some embodiments herein propose an architecture where one embedding network serves as a basis for training subsequent embedding networks. More specifically, once the sensor data embedding network is trained to generate sensor data embeddings, the query embedding network can be trained relying on the sensor data embedding network such that the generated query embeddings associated with a specific sensor data sample point to the same point as the corresponding sensor data embeddings of that specific sensor data sample in a multi-dimensional vector space. For example, this can be done by creating training examples with pairs of queries and corresponding sensor data samples, which is known in the field of machine learning. For example, a training pair could be the query "stop sign", then retrieving a sensor data sample of a stop sign (e.g., a camera image), and using its corresponding embedding to form the ground truth for the query embedding network. Thereby, a connection between the query and the sensor data sample is obtained, and it becomes possible to search for specific sensor data samples using, for example, free text queries.
[0033] Then, for any subsequent embedding networks ("auxiliary embedding networks") of other data sources (e.g., ADS output samples, DMS output samples, map data, etc.), a similar training process is performed. In other words, to add map data samples to the database and link them to the correct sensor data samples, the map data embedding network is trained relying on the sensor data embedding network such that the generated map data embeddings associated with a specific sensor data sample point to the same point as the corresponding sensor data embeddings of that specific sensor data sample in a multi-dimensional vector space.
[0034] By further adding these auxiliary data samples from other data sources, a more refined and powerful search of relevant data samples in the database can be achieved. For example, by simply adding map data samples to the already connected sensor data samples and queries, due to the connection between the sensor data samples (e.g., an image of a pedestrian) and the map data (geographical location and road type), it becomes possible for a user to retrieve relevant data samples for a query in the form of, for example, "pedestrians on a rural road in Germany". Without the connection between different data sources, such a query would either be impossible or would only yield wrong / incorrect results.
[0035] In some embodiments, the sensor data embedding network is a first sensor data embedding network configured to generate sensor data embeddings for sensor samples captured by a first sensor. More specifically, the first sensor data embedding can be a camera image embedding network, and correspondingly, the first sensor can be a camera. For any other sensor type (e.g., lidar, radar, etc.), additional embedding networks can be provided and trained in a manner similar to that described previously for the query embedding network and the auxiliary embedding network. Thus, the "base embedding network" can be a camera image embedding network.
[0036] Definition
[0037] In this context, an "Autonomous Driving System" ("ADS") refers to a complex combination of hardware and software components designed to control and operate a vehicle without direct human intervention. ADS technology aims to automate all aspects of driving (e.g., steering, accelerating, decelerating, and monitoring the surrounding environment). The primary goal of ADS is to improve the safety, efficiency, and convenience of transportation. Depending on its level of automation (classified according to standards such as SAE J3016), the scope of ADS can range from basic driver assistance systems to advanced autonomous driving systems. These systems use various sensors, cameras, radars, lidars, and powerful computer algorithms to perceive the environment and make driving decisions. The specific capabilities and features / functions of ADS can vary widely from systems that provide limited assistance to systems that can independently handle complex driving tasks under specific conditions.
[0038] Advanced Driver Assistance Systems (ADAS) are technologies that assist the driver during the driving process, but they do not necessarily provide full autonomy. ADAS features are often used as building blocks for ADS. Examples include Adaptive Cruise Control Systems, Lane Keeping Assistance Systems, Automatic Emergency Braking Systems, and Parking Assistance Systems. They enhance safety and convenience but typically require a certain degree of human supervision and intervention. On the other hand, Autonomous Driving (AD) is a technology designed to control and navigate a vehicle without human supervision. Accordingly, it can be said that the difference between ADAS and AD lies in the level of autonomy and control. ADAS systems are set up to help and support the driver, while AD aims to fully control the vehicle without the need for continuous human monitoring. Accordingly, AD aims to achieve a higher level of autonomy (such as Levels 4 and 5 according to SAE International standards), i.e., the vehicle can operate independently in most or all driving scenarios without human intervention. As described previously, the term "ADS" is used in this document as an umbrella term encompassing both ADAS and AD. In this context, ADS functions or ADS features can be understood as specific functions or features of the entire ADS stack, such as highway driver features, traffic jam driver features, path planning features, and so on.
[0039] In this context, a "sensor" or "sensor device" refers to a dedicated component or system designed to capture and collect information from the vehicle's surrounding environment. These sensors play a crucial role in enabling ADS to sense and understand its environment, make informed decisions, and navigate safely. Sensors are typically integrated into the hardware and software systems of autonomous vehicles to provide real-time data for various tasks such as obstacle detection, localization, road model estimation, and object recognition. Common types of sensors used in autonomous driving include lidar (Light Detection and Ranging), radar, cameras, and ultrasonic sensors. Lidar sensors use laser beams to measure distances and create a high-resolution 3D map of the vehicle's surrounding environment. Radar sensors use radio waves to determine the distance and relative speed of objects around the vehicle. Camera sensors capture visual data, enabling the vehicle's computer system to identify traffic signs, lane markings, pedestrians, and other vehicles. Ultrasonic sensors use sound waves to measure the proximity of objects. Various machine learning algorithms (e.g., artificial neural networks) can be employed to process the output of the sensors to understand the environment.
[0040] In this context, the term "data sample" refers to a subset of data obtained from a larger dataset or statistical population. Specifically, a "data sample" can be collected by sampling data collected by a fleet of vehicles equipped with ADS or by sampling data from other data sources external to the vehicle (e.g., simulation data). For a given dataset size, the data can be sampled at an appropriate sampling rate. For example, the data can be sampled at intervals of 1 second, 5 seconds, 10 seconds, etc. In some examples, data samples include "sensor data samples" and "auxiliary data samples".
[0041] The term "sensor data sample" can be interpreted as a specific instance or collection of data collected by sensors mounted on a vehicle equipped with ADS at a particular moment or within a particular time range. Sensor data samples typically include various types of information captured by the sensors, such as camera images, lidar outputs, radar outputs, GPS coordinates, accelerometer readings, and other sensor-generated data. "Sensor data samples" can include attached metadata (e.g., timestamps, location information, vehicle information, log duration, etc.).
[0042] The term "auxiliary data sample" can be interpreted as data sampled from a dataset that is not generated by the vehicle's sensors. Some examples of auxiliary data include map data, ADS data, driver monitoring system (DMS) data, and weather data. Map data samples can include, for example, HD map data samples (e.g., geographical regions defined in an HD map). ADS data samples can include, for example, perception outputs (e.g., object detection, semantic segmentation, detected lane markings, detected lanes, object classes, free space estimation, etc.), path planning outputs (e.g., candidate paths, executed paths, etc.), trajectory planning outputs (e.g., candidate trajectories, executed trajectories, etc.), hardware error logs, software error logs, and so on. In some embodiments, ADS data samples include outputs from one or more ADS functions (e.g., path planner, trajectory planner, perception system, decision and control functions, safety functions, etc.). For example, DMS data samples can include the driver's state (e.g., focused, sleepy, fatigued, distracted, etc.). For example, weather data samples can include weather data (e.g., precipitation, temperature, visibility, etc.). "Auxiliary data samples" can also include attached metadata (e.g., timestamps, location information, vehicle information, log duration, etc.).
[0043] A Driver Monitoring System (DMS) can be understood as a system including one or more cameras that focus on a vehicle driver to capture an image of the driver's face in order to determine various facial features of the driver, including the position, orientation, and movement of the driver's eyes, face, and head. Additionally, the DMS can be further configured to derive the driver's state based on the determined facial features, such as whether the driver is in a focused state or a non-focused state, whether the driver is fatigued, whether the driver is sleepy, and so on.
[0044] The term "embedding network" ("embedding neural network" or "embedding artificial neural network") refers to a collection of computational models or techniques used to enable a computer to generate embeddings for input data samples, where an "embedding" can be understood as a mathematical representation of data. More specifically, an "embedding network" is used to transform high-dimensional data into a low-dimensional space (a multi-dimensional vector space) while preserving the meaningful relationships between the input data points.
[0045] For example, embedding networks are used in tasks such as natural language processing (NLP) and computer vision. These networks take in raw input data (such as words in a sentence or an image) and transform them into fixed-size numerical vectors (embeddings) that capture the essential characteristics or features of the input data. More specifically, in NLP, an embedding network transforms words into numerical vectors where words with similar meanings or contextual usages are represented closer to each other in the embedding space (a multi-dimensional vector space). Similarly, in computer vision, an embedding network transforms an image into a numerical vector such that the network can understand visual similarities, for example, similar objects or scenes are grouped closer together in the embedding space (a multi-dimensional vector space).
[0046] A neural network or artificial neural network mimics a computational system inspired by the biological neural networks found in animal brains. These systems exhibit learning capabilities that gradually improve their performance without the need for specific task programming. For example, in image recognition, a neural network can be trained to detect specific objects in an image by analyzing labeled example images. Once the correlation between an object and its name is learned, the neural network can apply this knowledge to identify similar objects in unlabeled images.
[0047] Fundamentally, a neural network consists of interconnected units known as neurons, which are connected by synapses that transmit signals of different strengths. These signals propagate unidirectionally and activate the receiving neuron based on the strength of these connections. When the combined input signal from multiple transmitting neurons reaches a certain threshold, the receiving neuron activates and transmits a signal to downstream neurons. This activation strength becomes a key parameter that controls the propagation of signals in the network.
[0048] In addition, during the training of a neural network architecture, regression -- which consists of statistical processes used to understand variable relationships -- may involve minimizing a cost function. This function is used to measure the performance of the network in terms of accurately linking training examples to their expected outputs. If the value of this cost function exceeds a predetermined range based on known training data during training, a technique called backpropagation is employed. Backpropagation is a widely used method for training artificial neural networks and it is combined with optimization methods like Stochastic Gradient Descent (SGD).
[0049] In addition, the use of backpropagation can include propagation and weight updates. Backpropagation consists of two key steps: propagation and weight adjustment. When an input enters the neural network, it moves forward through each layer until it reaches the output layer. Here, the cost function is used to measure the output of the neural network relative to the desired output, generating an error value for each output node. These error values then flow back from the output layer, assigning error values to each node based on its contribution to the final output. These error values are crucial -- they help calculate the gradient of the cost function with respect to the weights of the neural network. This gradient guides the chosen optimization technique to adjust the weights to minimize the cost function.
[0050] Accordingly, an embedding network itself includes the various layers of a neural network architecture and typically employs techniques like convolutional layers, recurrent layers, or fully connected layers to learn and extract meaningful patterns from input data. The embedding network can be trained through processes like supervised learning, unsupervised learning, or self-supervised learning to optimize the embeddings for a specific downstream task (such as classification, clustering, or recommendation). In some embodiments, various embedding networks are trained to generate embeddings in the same embedding space (the same multi-dimensional vector space) such that embeddings that are contextually, spatially, and / or temporally related (generated by different embedding networks) point to the same point within the multi-dimensional vector space.
[0051] This can be done, for example, by training a first embedding network to generate an embedding in a multi-dimensional vector space based on input data from a first data source. Then, each of the other embedding networks is trained "in dependence on" or in association with the first embedding network such that the embeddings of the other networks that are contextually, spatially, and / or temporally related to the embedding of the first embedding network point to the same point within the multi-dimensional vector space as the relevant embedding of the first embedding network. For example, if the first embedding network is trained to generate an image embedding for a camera image and a second embedding network is intended to generate an embedding for lidar data, the second embedding network can be trained by feeding it the lidar data of a scene, where the corresponding image embedding of (that scene) will be used as a basis for forming the ground truth (desired output). By performing this process for each subsequent embedding network, a set of embedding networks can be obtained that are capable of ingesting outputs from various data sources and outputting corresponding embeddings, where the relationships in context, space, and / or time are represented by the proximity or directional similarity of the embeddings (vectors) within the multi-dimensional vector space.
[0052] Embodiment
[0053] Figure 1 is a schematic block diagram representation of a system 10 for accessing a data set including a plurality of data samples from a plurality of data sources 31a, 31b, 32 - 36. The system 10 includes control circuitry 11 (e.g., one or more processors - see Figure 2 ), and the control circuitry 11 is configured to perform the functions of the method S100 disclosed herein, where the functions may be included in a non-transitory computer-readable storage medium 12 or in other computer program products configured to be executed by the control circuitry 11. In other words, the system 10 includes one or more memory storage areas 12 containing program code, and the one or more memory storage areas 12 and the program code are configured to, together with one or more processors 11, cause the system 10 to perform the method S100 according to any one of the embodiments disclosed herein. However, for better elucidation of the embodiments disclosed herein, the control circuitry is represented in Figure 1 as various components or blocks, each of which is linked to one or more specific functions of the control circuitry.
[0054] More specifically, Figure 1An example architecture of system 10 according to some embodiments is outlined. Fleet data (also referred to as “probe data,” where test vehicles and / or production vehicles act as “probes”) 21 is sampled S201 at a certain predetermined sampling rate based on the amount of data 21 provided. For example, with 60,000 hours of available data, a sampling rate of one data sample every 10 seconds would result in a dataset comprising 216 million data samples. In some embodiments, the probe data is collected at a fixed frequency (e.g., during a specific time of day) or in a more dynamic manner (where the collection of the probe data is triggered based on conditions detected by the vehicle).
[0055] Further, the sampled data is divided into two data bins, a sensor data bin 22 and an auxiliary data bin 23. This is done mainly for illustrative purposes, and all the data samples can be stored in a “common bin” or divided into additional data bins depending on the specific implementation and requirements.
[0056] Accordingly, a plurality of data samples includes sensor data samples 31a that capture information about the vehicle's surrounding environment by one or more vehicles. Here, each sensor data sample 31a is represented by a corresponding sensor data embedding, which is generated by processing the sensor data sample 31a via a sensor data embedding network 41a that has been trained to process the sensor data sample 31a and output a corresponding sensor data embedding for each sensor data sample 31a in a multi-dimensional vector space. As Figure 1 shown, system 10 may include additional sensor data embedding networks 41b. Specifically, system 10 may include sensor data embedding networks 41a, 41b, 42 for each sensor 31a, 31b, 32 of the vehicle, or one sensor data embedding network 41a, 41b, 42 for each sensor type / modality. For example, vehicle state information may include, for instance, the speed, position, and / or angular velocity of the vehicle output by the vehicle's inertial measurement unit (IMU) and / or global navigation satellite system (GNSS).
[0057] The plurality of data samples further includes auxiliary data samples 23, where each auxiliary data sample is represented by a corresponding auxiliary data embedding, and the auxiliary data embedding is generated by processing the auxiliary data sample via an auxiliary data embedding network 43-46, and the auxiliary data embedding network 43-46 has been trained to process the auxiliary data sample and output the corresponding auxiliary data embedding in a multi-dimensional vector space. In addition, the auxiliary data embedding network 43-46 has been trained in association with the sensor data embedding network 41a such that the auxiliary embedding of the auxiliary data sample associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample. Similar to the sensor data embedding networks 41a, 41b, 42, the system 10 may include an auxiliary data embedding network 43-46 for each auxiliary data source 33-36. For example, the system 10 may include a map data embedding network 43, an autonomous driving system (ADS) data embedding network 44, a driver monitoring system (DMS) data embedding network 45, and a weather data embedding network 46. Naturally, additional auxiliary data embedding networks may be added to the system 10 depending on the availability of the auxiliary data samples.
[0058] Accordingly, the auxiliary data sample 23 may include a map data sample 33. Then, each map data sample 33 is represented by a corresponding map data embedding, and the map data embedding is generated by processing the map data sample 33 via a map data embedding network 43, and the map data embedding network 43 has been trained to process the map data sample 33 and output the corresponding map data embedding in a multi-dimensional vector space. In addition, the map data embedding network 43 has been trained in association with the sensor data embedding network 41a such that the map embedding of the map data sample 33 associated with a particular sensor data sample 31a points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample 31a. The map data sample 33 may be, for example, an HD map data sample. The map data sample 33 and the sensor data sample 31a may be linked, for example, by the associated metadata of the sensor data sample 31a, which indicates the location where the particular sensor data sample was initially captured / generated.
[0059] In addition, the auxiliary data sample 23 may include an ADS data sample 34. Each ADS data sample 34 is represented by a corresponding ADS data embedding, which is generated by processing the ADS data sample 34 via an ADS data embedding network 44 that has been trained to process the ADS data sample 34 and output the corresponding ADS data embedding in a multi-dimensional vector space. In addition, the ADS data embedding network 44 has been trained in association with the sensor data embedding network 31a such that the ADS embedding of the ADS data sample 34 associated with a particular sensor data sample 31a points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample 31a. The ADS data sample 34 and the sensor data sample 31a may be linked, for example, by associated metadata of the sensor data sample 31a and the ADS data sample 34 that indicates the time at which the particular sensor data sample 31a was initially captured / generated and the time and location at which the particular ADS data sample 34 was output / generated.
[0060] Still further, the auxiliary data sample 23 may include a driver monitoring system (DMS) data sample 35. Each DMS data sample 35 is represented by a corresponding DMS data embedding, which is generated by processing the DMS data sample 35 via a DMS data embedding network 45 that has been trained to process the DMS data sample 35 and output the corresponding DMS data embedding in a multi-dimensional vector space. In addition, the DMS data embedding network 45 has been trained in association with the sensor data embedding network 31a such that the DMS embedding of the DMS data sample 35 associated with a particular sensor data sample 31a points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample 31a. The DMS data sample 35 and the sensor data sample 31a may be linked, for example, by associated metadata of the sensor data sample 31a and the DMS data sample 35 that indicates the time and location at which the particular sensor data sample 31a was initially captured / generated and the time and location at which the particular DMS data sample 35 was output / generated.
[0061] Further, the auxiliary data samples can include weather data samples 36. Each weather data sample 36 is represented by a corresponding weather data embedding, which is generated by processing the weather data sample 36 via a weather data embedding network 46 that has been trained to process the weather data sample 36 and output a corresponding weather data embedding in a multi-dimensional vector space. Additionally, the weather data embedding network 46 has been trained in association with the sensor data embedding network 31a such that the weather embedding of the weather data sample 36 associated with a particular sensor data sample 31a points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample 31a. The weather data sample 36 and the sensor data sample 31a can be linked, for example, by associated metadata of the sensor data sample 31a and the weather data sample 36, which indicates the time and location at which the particular sensor data sample 31a was initially captured / generated and the time and location at which the particular weather data sample 36 was output / generated.
[0062] Accordingly, the system has a collection of data samples collected from the fleet data 21, where each data source 31a, 31b, 32 - 36 is associated with a corresponding embedding network 41, 41b, 42 - 46, which are trained to output embeddings of the input data samples. The embeddings are then stored in a suitable data store 50, which can further include links to the respective data samples.
[0063] Moreover, the system 10 includes control circuitry 11 configured to identify one or more embeddings 50 in the multi-dimensional vector space based on proximity to an acquired query embedding in the multi-dimensional vector space in response to acquiring the query embedding.
[0064] Here, the query embedding is generated by processing the query 60 via a query embedding network 61 that has been trained to process the query 60 and output a corresponding query embedding for each query in the multi-dimensional vector space. Additionally, the query embedding network 61 has been trained in association with the sensor data embedding network 31a such that the query embedding of the query 60 associated with a particular sensor data sample 31a points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample 31a. The query 60 can be in the form of a text query, an image query, or a combination thereof.
[0065] The system 10 can include the query embedding network 61, and accordingly, the control circuitry 11 can be configured to receive a query from a client device and use the query embedding network 60 to process the query 60 to generate a query embedding. However, in some embodiments, the query embedding network 61 is external to the system 10 and the query embedding is received from the client device.
[0066] The control circuit 11 may be configured to identify one or more sensor data and / or auxiliary data embeddings within a multi-dimensional vector space by identifying one or more sensor data and / or auxiliary data embeddings within a certain distance value from the obtained query embedding. The term "within a certain distance value" may be interpreted as "satisfying a distance metric". This is represented by Figure 1 the scoring algorithm block 71 of the system 10 in
[0067] wherein the scoring algorithm is configured to score the embeddings based on the relevance defined by the distance value from the query embedding. Some suitable distance metrics that can be used are Euclidean distance, Manhattan distance, or cosine distance.
[0068] Simply go to Figure 2 , Figure 2 is a schematic diagram of the system 10 for accessing a data set including multiple data samples from multiple data sources according to some embodiments, a schematic diagram of the server 401 including such a system, and a schematic diagram of the cloud environment 402 including multiple servers 401.
[0069] As described, the system 10 includes a control circuit (such as one or more processors) 11, and the control circuit 11 is configured to perform the functions of the method S100 disclosed herein, where the functions may be included in the non-transitory computer-readable storage medium 12 or in other computer program products configured to be executed by the control circuit 11. In other words, the system 10 includes one or more memory storage areas 12 containing program code, and the one or more memory storage areas 12 and the program code are configured to cause the system 10 to execute a method according to any one of the embodiments disclosed herein in conjunction with one or more processors 11.
[0070] The control circuit 11 may physically comprise a single circuit device. Alternatively, the control circuit 11 may be distributed over a number of circuit devices. The control circuit 11 may include one or more processors, such as a central processing unit (CPU), a microcontroller, or a microprocessor. The one or more processors may be configured to execute program code stored in the memory 12 to implement various functions and operations in addition to the methods disclosed herein. The processor 11 may be or may include any number of hardware components for performing data or signal processing or for executing computer code stored in the memory 12. The memory 12 optionally includes high-speed random access memory (such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices), and optionally includes non-volatile memory (such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices). The memory 12 may include database components, object code components, script components, or any other type of information structure for supporting the various activities of this specification.
[0071] Figure 3 is a schematic flowchart representation of a method S100 for accessing a data set including multiple data samples from multiple data sources according to some embodiments. The method S100 is preferably a computer-implemented method S100, executed by a processing system of a computer. The processing system may, for example, include one or more processors and one or more memories coupled to the one or more processors, wherein the one or more memories store one or more programs that, when executed by the one or more processors, perform the steps, services, and functions of the method 100 disclosed herein.
[0072] In some embodiments, the method S100 includes obtaining a data set including embeddings of multiple data samples from multiple data sources.
[0073] As previously mentioned, the multiple data samples include sensor data samples captured by one or more vehicles and including information about the surrounding environment of the vehicle. Each sensor data sample is represented by a corresponding sensor data embedding, which is generated by processing the sensor data sample via a sensor data embedding network that has been trained to process the sensor data sample and output a corresponding sensor data embedding for each sensor data sample in a multi-dimensional vector space.
[0074] In addition, the plurality of data samples include auxiliary data samples, where each auxiliary data sample is represented by a corresponding auxiliary data embedding, and the auxiliary data embedding is generated by processing the auxiliary data sample via an auxiliary data embedding network that has been trained to process the auxiliary data sample and output the corresponding auxiliary data embedding in a multi-dimensional vector space. The auxiliary data embedding network has been trained in association with the sensor data embedding network such that the auxiliary embedding of the auxiliary data sample associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample.
[0075] Method S100 includes obtaining S102 a query embedding. In some embodiments, the query embedding is received from an external entity (e.g., a client device - e.g., a general-purpose computer). However, in some embodiments, method S100 includes receiving a query from a client device and generating a query embedding by processing the received query using a query embedding network. The term "obtaining" is broadly understood herein and encompasses receiving, retrieving, collecting, acquiring, etc., directly and / or indirectly between two entities configured to communicate with each other or further with other external entities. However, in some embodiments, the term "obtaining" is interpreted as determining, deriving, forming, calculating, etc. Thus, as used herein, "obtaining" may indicate that a parameter is received at a first entity / unit from a second entity / unit, or that a parameter is determined at a first entity / unit, e.g., based on data received from another entity / unit.
[0076] Further, in response to obtaining S102 the query embedding, method S100 includes identifying S104 one or more embeddings in the multi-dimensional vector space based on proximity to the obtained query embedding within the multi-dimensional vector space. Here, the query embedding is generated by processing the query via a query embedding network that has been trained to process the query and output a corresponding query embedding in the multi-dimensional vector space for each query. In addition, the query embedding network has been trained in association with the sensor data embedding network such that the query embedding of the query associated with a particular sensor data sample points to the same point in the multi-dimensional vector space as the sensor data embedding of that sensor data sample.
[0077] Further, method S100 includes outputting S105 one or more data samples within the data set represented by the identified sensor data embeddings and / or auxiliary data embeddings. The output S105 of the one or more data samples may include sending the one or more data samples to a client device.
[0078] In addition, in some embodiments, the identification S104 of one or more embeddings within the multi-dimensional vector space includes identifying one or more embeddings within the multi-dimensional vector space that are within a certain distance value from the obtained query embedding. As previously mentioned, the term "within a certain distance value" can be interpreted as "satisfying a distance metric", where the distance metric can be that the selected embedding is an embedding within the Euclidean / Manhattan / cosine distance value (e.g., a distance threshold) from the query embedding. Thus, in some embodiments, the method S100 may include comparing the obtained S102 query embedding with the sensor data embeddings and the auxiliary data embeddings within the multi-dimensional vector space in terms of the distance metric to identify relevant sensor data embeddings and / or auxiliary data embeddings.
[0079] The executable instructions for performing these functions are optionally included in a non-transitory computer-readable storage medium or included in other computer program products configured to be executed by one or more processors.
[0080] The present invention has been presented with reference to specific embodiments above. However, other embodiments are possible and within the scope of the present invention. Method steps for performing the methods described above by hardware or software different from those described above can be provided within the scope of the present invention. Thus, according to an exemplary embodiment, there is provided a non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of a computing system, the one or more programs including instructions for performing the method according to any one of the above embodiments. Alternatively, according to another exemplary embodiment, a cloud computing system can be configured to execute any one of the methods presented herein. The cloud computing system can include distributed cloud computing resources that jointly execute the methods presented herein under the control of one or more computer program products.
[0081] Generally speaking, a computer-accessible medium can include any tangible or non-transitory storage medium or memory medium (such as electronic media, magnetic media, or optical media - for example, a disk or CD / DVD-ROM) coupled to a computer system via a bus. As used herein, the terms "tangible" and "non-transitory" are intended to describe computer-readable storage media (or "memory") that do not include propagating electromagnetic signals, but are not intended to otherwise limit the types of physical computer-readable storage devices covered by the term computer-readable medium or memory. For example, the term "non-transitory computer-readable medium" or "tangible memory" is intended to cover storage device types that do not necessarily store information permanently, including, for example, random access memory (RAM). Program instructions and data stored in a tangible computer-accessible storage medium in non-transitory form can be further transmitted by a transmission medium or by signals (such as electrical, electromagnetic, or digital signals) that can be transmitted via a communication medium such as a network and / or a wireless link.
[0082] (Associated with system 10) Processor 11 can be or can include any number of hardware components for performing data or signal processing or for executing computer code stored in memory 12. Device 10 has an associated memory 12, and memory 12 can be one or more devices for storing data and / or computer code for completing or facilitating the various methods described in this specification. The memory can include volatile memory or non-volatile memory. Memory 12 can include database components, object code components, script components, or any other type of information structure for supporting the various activities of this specification. According to an exemplary embodiment, any distributed or local storage device can be used with the systems and methods of this specification. According to some embodiments, memory 12 is communicatively coupled to processor 11, for example, via circuitry or any other wired, wireless, or network connection, and includes computer code for executing one or more of the processes described herein.
[0083] It should be noted that any reference numerals do not limit the scope of the claims, and the present invention can be implemented at least in part in hardware and software, and several "means" or "units" can be represented by the same hardware item.
[0084] Although the figures may show a particular order of method steps, the order of the steps may be different from that depicted. In addition, more than two steps may be performed simultaneously or partially simultaneously. For example, the steps of receiving a signal including information about motion and information about the current road scene may be interchanged based on a particular implementation. Such variations will depend on the software and hardware systems selected and the choices of the designer. All such variations are within the scope of the present invention. Similarly, software implementations can be accomplished with standard programming techniques having rule-based logic and other logic for performing the various acquisition, comparison, identification, and output steps. The embodiments mentioned and described above are given only as examples and should not be a limitation on the present invention. Other solutions, uses, objectives, and functions within the scope of the present invention claimed in the patent claims should be obvious to those skilled in the art.
Claims
1. A method (S100) for accessing a data set comprising a plurality of data samples from a plurality of data sources, wherein: The plurality of data samples include: sensor data samples captured by one or more vehicles, including information about the surroundings of the vehicles, wherein each sensor data sample is represented by a corresponding sensor data embedding generated by processing the sensor data sample through a sensor data embedding network that has been trained to process the sensor data samples and output a corresponding sensor data embedding for each sensor data sample in a multi-dimensional vector space; auxiliary data samples, wherein each auxiliary data sample is represented by a corresponding auxiliary data embedding generated by processing the auxiliary data sample through an auxiliary data embedding network that has been trained to process the auxiliary data samples and output a corresponding auxiliary data embedding in the multidimensional vector space, and wherein the auxiliary data embedding network has been trained in association with the sensor data embedding network such that the auxiliary embedding of the auxiliary data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of the sensor data sample, Wherein, the method (S100) comprises: In response to obtaining (S102) a query embedding, identifying (S104) one or more embeddings in the multidimensional vector space based on proximity to the obtained query embedding in the multidimensional vector space, wherein the query embedding is generated by processing a query via a query embedding network that has been trained to process queries and output a corresponding query embedding in the multidimensional vector space for each query, and wherein the query embedding network has been trained in association with the sensor data embedding network such that the query embedding for a query associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding for the sensor data sample; and One or more data samples within the data set represented by the identified embeddings are output (S105).
2. The method (S100) according to claim 1, wherein: The auxiliary data samples include: map data samples, wherein each map data sample is represented by a corresponding map data embedding, the map data embedding being generated by processing the map data samples through a map data embedding network, the map data embedding network having been trained to process the map data samples and output a corresponding map data embedding in the multidimensional vector space, and wherein the map data embedding network has been trained in association with the sensor data embedding network such that the map embedding of the map data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of the sensor data sample.
3. The method (S100) according to claim 1, wherein: The auxiliary data samples include: ADS data samples, wherein each ADS data sample is represented by a corresponding ADS data embedding, the ADS data embedding being generated by processing the ADS data sample via an ADS data embedding network, the ADS data embedding network having been trained to process the ADS data samples and output a corresponding ADS data embedding in the multidimensional vector space, and wherein the ADS data embedding network has been trained in association with the sensor data embedding network such that the ADS embedding of an ADS data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of that sensor data sample.
4. The method (S100) according to claim 1, wherein: The auxiliary data samples include: DMS data samples, wherein each DMS data sample is represented by a corresponding DMS data embedding, the DMS data embedding being generated by processing the DMS data samples via a DMS data embedding network, the DMS data embedding network having been trained to process the DMS data samples and output a corresponding DMS data embedding in the multidimensional vector space, and wherein the DMS data embedding network has been trained in association with the sensor data embedding network such that the DMS embedding of a DMS data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of the sensor data sample.
5. The method (S100) according to claim 1, wherein: The auxiliary data samples include: Weather data samples, wherein each weather data sample is represented by a corresponding weather data embedding, the weather data embedding being generated by processing the weather data sample through a weather data embedding network that has been trained to process the weather data samples and output a corresponding weather data embedding in the multidimensional vector space, and wherein the weather data embedding network has been trained in association with the sensor data embedding network such that the weather embedding of the weather data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of the sensor data sample.
6. The method (S100) according to claim 1, further comprising: receiving the query from a client device; The query embedding is generated by processing the received query using the query embedding network.
7. The method (S100) according to claim 1, wherein: The identifying (S104) of one or more embeddings in the multi-dimensional vector space comprises: The one or more embeddings within the multidimensional vector space that are within a distance value from the obtained query embedding are identified.
8. The method (S100) according to claim 1, wherein: The query is a text query or an image query.
9. A computer program product storing instructions, which, when executed by a computer, cause the computer to perform the method (S100) according to claim 1.
10. A non-transitory computer-readable storage medium comprising instructions, which, when executed by a computer, cause the computer to perform the method (S100) according to claim 1.
11. A system (10) for accessing a data set comprising a plurality of data samples from a plurality of data sources, wherein: The plurality of data samples include: Sensor data samples (22) captured by one or more vehicles, including information about the surrounding environment of the vehicles, wherein each sensor data sample (22) is represented by a corresponding sensor data embedding, the sensor data embedding being generated by processing the sensor data sample via a sensor data embedding network (41a, 41b), the sensor data embedding network (41a, 41b) having been trained to process the sensor data samples and output a corresponding sensor data embedding for each sensor data sample in a multi-dimensional vector space; auxiliary data samples (23), wherein each auxiliary data sample is represented by a corresponding auxiliary data embedding, the auxiliary data embedding being generated by processing the auxiliary data sample through an auxiliary data embedding network (43, 44, 45, 46), the auxiliary data embedding network (43, 44, 45, 46) having been trained to process the auxiliary data samples and output a corresponding auxiliary data embedding in the multidimensional vector space, and wherein the auxiliary data embedding network has been trained in association with the sensor data embedding network (41a, 41b) such that the auxiliary embedding of the auxiliary data sample associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding of that sensor data sample, The system (10) comprises a control circuit (11), wherein the control circuit (11) is configured to: In response to obtaining a query embedding, identifying one or more embeddings in the multidimensional vector space based on proximity to the obtained query embedding in the multidimensional vector space, wherein the query embedding is generated by processing the query via a query embedding network (61), the query embedding network (61) having been trained to process the queries and output a corresponding query embedding in the multidimensional vector space for each query, and wherein the query embedding network has been trained in association with the sensor data embedding network (41a, 41b) such that the query embedding for a query associated with a particular sensor data sample points to the same point in the multidimensional vector space as the sensor data embedding for the sensor data sample; One or more data samples within the data set represented by the identified embeddings are output (80).
12. The system (10) according to claim 11, wherein: The control circuit (11) is further configured to: receiving the query from a client device (60); The query embedding is generated by processing the received query using the query embedding network (61).
13. The system (10) of claim 11, wherein: The identifying of one or more embeddings within the multi-dimensional vector space comprises: The one or more embeddings within the multidimensional vector space that are within a distance value from the obtained query embedding are identified.
14. A server (401) comprising the system according to any one of claims 11-13.
15. A cloud (402) environment comprising one or more servers according to claim 14.