Techniques for sensor-agnostic task head training
By pre-training the task head on optical sensor data and utilizing latent representations and generative models, the computational resource and time requirements for model retraining after sensor replacement are addressed, achieving efficient sensor adaptive training.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2026-02-02
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies require significant computational resources and time to retrain the model after sensor replacement, and face the challenges of limited available training data and ground truth.
By pre-training the task head on optical sensor data, sensor-agnostic training is performed using latent representations, and combined with generative models and renderers, the requirements for computing resources and time are reduced.
This enables the task head to be trained without retraining after sensor replacement, reducing computational resource and time requirements and improving training efficiency and model adaptability.
Smart Images

Figure CN122491389A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a technique for pre-training a task head for a perception task and for performing the perception task, specifically including a method, a task head, a neural network system including the task head, and a computer program product. Background Technology
[0002] Computer-implemented models for performing perception tasks often face challenges such as limited available training data and relevant ground truth, as well as sensor replacement issues. For example, in automotive and robotic sensors, newer models of vehicles or robots are often equipped with newer versions of optical sensors, and the models are not trained on these newer versions of optical sensors.
[0003] There are conventional methods such as those proposed in [1], which involve joint embedding learning with unsupervised, self-supervised and / or domain adversarial losses; style transfer methods that map target domain data to source domain data, as proposed in [2]; and label preservation or label conditional data synthesis methods for target domain data, as proposed in [3].
[0004] However, all these conventional methods have limited applicability and typically require significant computational resources and training time.
[0005] Therefore, one object of the present invention is to provide a solution for transferring a model trained on labeled source domain data for certain tasks to unlabeled target domain data. Alternatively or additionally, an object is to provide at least partially sensor-agnostic training, and / or to provide training techniques that reduce the demand for computational resources and / or reduce the time required for retraining (e.g., after replacing sensors for a newer model). Summary of the Invention
[0006] This objective is achieved by a method for pre-training a task head for a perception task, a method for performing a perception task, a task head, a neural network system, a computer program (and / or a computer program product), and / or a computer-readable storage medium, as described in the appended independent claims. Advantageous aspects, features, embodiments, and advantages are described in the dependent claims and the following description.
[0007] In the following description, the solutions according to the invention are described with respect to the claimed method and the claimed task head and neural network system. Features, advantages, or alternative embodiments described herein can be assigned to other claimed objects (e.g., computer programs or computer program products), and vice versa. In other words, the claims for the task head and neural network system can be improved using features described or claimed in the context of the method. In this case, the functional features of the method are embodied by the structural units of the system, and vice versa.
[0008] Regarding the first method, a method (particularly computer-implemented) for pre-training a task head for a perception task is provided. The method includes the step of receiving multiple latent representations at the input layer of the task head. Each latent representation is associated with a first training dataset included in a first training database. Each first training dataset includes optical sensor data acquired by at least one optical sensor. Each first training dataset also includes a perception task label. The method further includes the step of transforming each of the multiple received latent representations into a perception task result by one or more hidden layers of the task head. The method further includes the step of outputting the perception task result of each of the multiple received latent representations at the output layer of the task head. The method further includes the step of pre-training the weights of the task head (e.g., the weights of at least one or more hidden layers) by optimizing a task-specific loss function. Optimizing the task-specific loss function involves comparing the output perception task result with the perception task label for each of the multiple received latent representations.
[0009] A task head is trained (also called pre-trained) on a latent representation (also called world representation and / or scene representation) of optical sensor data (and potentially other sensor data, such as motion sensor data), where the latent representation is sensor-agnostic. Thus, the pre-trained task head can be combined with various encoders (also called converters) that can output such a latent representation. Encoders may differ, for example, in the type of optical sensor data they were trained on and / or on at least one optical sensor acquiring the optical sensor data. For example, when an optical sensor (e.g., an older camera) is replaced by a newer one (e.g., a newer camera), the pre-trained task head may be combined with the updated encoder. Alternatively or additionally, sensor parameters (e.g., resolution) may differ between optical sensors. Encoders may differ in terms of the sensor parameters (also called sensor-specific parameters) for which they were pre-trained.
[0010] If the type of optical sensor data and / or the optical sensors acquiring it changes, it is possible to pre-train a task head for the perception task on a finite number of first training datasets (also known as: source domain datasets), each of which includes labels related to the perception task (also known as: perception task labels and / or ground truths), by only needing to pre-train a second encoder (or retrain the first encoder used to pre-train the task head). The pre-trained task head is then transferred to a neural network system capable of performing the perception task on an additional dataset (also known as: target domain dataset) that differs in at least one optical sensor (and / or sensor parameters) from which the optical sensor data was acquired. The neural network system used to perform the perception task may include a second encoder trained on a second training dataset without labels (also known as: target domain dataset). In particular, retraining the task head is not required, which reduces the amount of computational resources and time required for retraining in the event of changes in the acquisition of optical sensor data.
[0011] The modalities (also referred to as types) of the optical sensors used to acquire the first training dataset (for a labeled first training dataset) and the second training dataset (for a unlabeled second training dataset) can be the same. For example, at least one camera and at least one LiDAR sensor can be used to acquire training datasets within the first and second training datasets.
[0012] At least one sensor may include at least one sensor of a predetermined type (also referred to as: modal, measurement principle and / or acquisition principle; for example, camera or LiDAR).
[0013] At least one sensor may include at least two sensors of two different predetermined types.
[0014] By progressively pre-training a task head (particularly sensor-agnostic) and (e.g., a second) sensor-specific encoder with output sensor-agnostic and / or task-agnostic representations, model complexity and / or the size of the sensor-specific encoder can be reduced. Alternatively or additionally, less training data is required. The less training data required can be applied to training the task head with less labeled training data (also referred to as the first training dataset). Alternatively or additionally, the less training data required can be applied to training the second encoder with less unlabeled training data (also referred to as the second training dataset).
[0015] Applications of perception tasks can include autonomous driving and robotics (e.g., for manufacturing and / or for home automation), such as for perception and / or sensor fusion, and / or for control actuators.
[0016] According to the Society of Automotive Engineers (SAE) classification system, autonomous driving can include any of Level 1 (driver assistance), Level 2 (partial automation), Level 3 (conditional automation), Level 4 (high automation), and / or Level 5 (full automation).
[0017] At least one optical sensor may be configured to capture the scene and / or environment of a vehicle (also known as a self-driving car) and / or robot (e.g., in a manufacturing site and / or at home). Alternatively or additionally, at least one optical sensor may be mounted on the self-driving car and / or robot.
[0018] At least one optical sensor may be included in the first group of sensors. The first group of sensors may include one or more sensors of predefined sensor types (also referred to as: modal, acquisition principle, and / or measurement type), such as cameras and LiDAR sensors.
[0019] At least one optical sensor may include multiple optical sensors and / or sensors of different types. Thus, the potential representation associated with optical sensor data (also known as: sensor signal) may be truly sensor-agnostic.
[0020] The latent representation can be task-agnostic in the sense that it can be used to pre-train multiple task heads for different perception tasks (also known as downstream tasks). For example, the first task head can be pre-trained for object detection, such as detecting traffic signs near a vehicle. The second task head can be pre-trained for map estimation, such as estimating the shape of road lanes near a vehicle.
[0021] The latent representation can be task-agnostic and / or sensor-agnostic.
[0022] Sensor-agnostic (also known as: sensor-specific opposite, and / or non-sensor-specific) can correspond to a potential representation that is not in the sensor coordinate system.
[0023] The latent representation can have an internal structure, such as a single bird's-eye view (BEV) representation around the vehicle. The components of the latent representation can represent scene dynamics (e.g., scene flow and / or object flow) and / or scene appearance.
[0024] The latent representation can be modular, e.g., static versus temporal, and / or dense versus sparse. This can improve the transparency and / or interpretability of the task head's pre-training. Alternatively or additionally, several auxiliary SSL tasks (SSL: Self-Supervised Learning) can be executed in parallel, thereby improving the overall pre-training of the task head.
[0025] The latent representation may include the current scene and optionally include a generative model.
[0026] Alternatively or additionally, the first encoder and / or any other encoder may include (and / or may be provided with) a generative model. Further alternatively or additionally, the task head and / or renderer may include (and / or may be provided with) a generative model.
[0027] Generative models can encode the temporal evolution of optical sensor data. Alternatively or additionally, generative models can generate sensor inputs for a given sensor parameterization, for example, to allow for temporal evolution.
[0028] (For example, the first and / or second) encoders can be trained using the reconstruction loss of the generative model and optional auxiliary SSL loss, particularly in both the source and target domains.
[0029] Encoder architecture can, in principle, depend on the sensor modality (also known as: sensor type or sensor type). For example, for camera input, certain convolutional networks (e.g., ResNet, VovNet, etc.) or visual transformer networks (ViT) can be chosen and / or beneficial. For radar input, other encoder models (e.g., PillarNet) may be more suitable.
[0030] The generative model G can depend on the sensor parameterization p. The fact that G can take p as a separate input G(z, p) makes G independent of the choice of p. G(z, p) depends on the latent state z (and / or the general scene representation; also known as the latent representation) which can be mapped using different sensor parameterizations p_1, p_2, for example, to two different camera views of the same scene: x_1=G(z, p_1), x2=G(z, p_2). In this sense, the scene representation z itself, as well as the generative model, can be sensor-agnostic.
[0031] The generative model G does not need to be part of the encoder (e.g., the first and / or the second). Alternatively or additionally, the generative model G may operate on the output z of the encoder (e.g., the first and / or the second) and generate sensor states x=G(z, p) given sensor parameterization.
[0032] (e.g., dense and / or sparse) latent representations can be the output of (e.g., first and / or second) encoders as well as inputs to both the renderer (and / or the generative model) and the task head.
[0033] The first training database may include a first training dataset that uses a combination of one or more optical sensors. For example, the Cityscapes database (available at https: / / www.cityscapes-dataset.com / ) may provide a first training dataset that is camera-based (and / or acquired via one or more cameras as optical sensors). The first training database may alternatively or additionally include a first training dataset provided by nuScenes (available at https: / / www.nuscenes.org / nuscenes). The dataset in the nuScenes database was acquired via up to six cameras, one LiDAR sensor, and / or five radar sensors as optical sensors. The dataset in the nuScenes database also includes sensor data from motion sensors and / or navigation sensors (e.g., inertial measurement units (IMUs) and / or global positioning satellite (GPS) sensors).
[0034] The results of pre-training the task head can include or may include the weights of the task head.
[0035] In some embodiments, the first training dataset may include 3D object bounding box labels for predetermined categories, and optionally include additional labels for other categories. This can improve the sensor agnosticness and / or task agnosticness of the pre-trained task head.
[0036] The method may include the step of receiving a perceptual task label associated with each of a plurality of received latent representations at the loss function of the task head.
[0037] The perception task labels can be received along with the optical sensor data from the first training dataset.
[0038] At least one optical sensor may include a camera, a LiDAR sensor, a radar sensor, an ultrasonic sensor, and / or a thermal sensor.
[0039] Any optical sensor can acquire digital sensor data and / or can output optical sensor data in digital form.
[0040] The camera can be a fisheye camera or a stereo camera. Alternatively or additionally, the camera can be a video camera.
[0041] At least one optical sensor (e.g., a video camera) may be adapted to acquire sensor data for performing regression and / or determining one or more continuous values, such as distance (e.g., distance to a vehicle or landmark in a scene (e.g., an intersection), speed, and / or tracking of items (e.g., objects). The regression and / or determination of one or more continuous values may be performed based on low-level features.
[0042] The first training database may include a first training dataset comprising optical sensor data from two or more different optical sensors and / or from two or more different optical sensor types. Therefore, a task head may be pre-trained to perform a perception task for any combination of optical sensor types, acquiring optical sensor data from the first training database.
[0043] At least one subset of the first training dataset may also include motion sensor data, position sensor data, and / or navigation sensor data.
[0044] Perception tasks can be performed based on a combination of optical sensor data and motion-related (e.g., including motion, position, and / or navigation) sensor data. Thus, the motion of a vehicle or robot can be combined with the scene and / or environment captured by optical sensors.
[0045] The temporal evolution of optical sensor data can be encoded and / or predicted by using motion-related sensor data from the same scene and / or environment.
[0046] Motion sensor data can be acquired using motion sensors. Alternatively or additionally, position sensor data and / or navigation sensor data can be acquired using satellite-based navigation systems.
[0047] By combining multiple sensor types, such as motion sensors and optical sensors, the accuracy of representing scenes and / or environments through sensor data can be improved.
[0048] Multiple latent representations can be obtained through a pre-trained first encoder.
[0049] The first encoder can be pre-trained using the first training database by combining the first encoder with the renderer (also known as the decoder) and optimizing the reconstruction loss.
[0050] The renderer can be a white-box renderer or a gray-box renderer. Alternatively or additionally, the renderer can be trained to reconstruct optical sensor data from the latent representation.
[0051] White-box (also known as: glass-box, transparent-box, and / or open-box) renderers can be non-trainable renderers, and / or the weights of white-box renderers can be immutable.
[0052] The black-box renderer can be (e.g., fully) trainable, and / or the (e.g., all) weights of the black-box renderer can be variable. The training of the black-box renderer can depend on the type of sensor and / or sensor parameters.
[0053] A gray-box renderer can be an intermediary between a white-box renderer and a black-box renderer. For example, the renderer can be trainable based on a subset of sensor parameters, such as the sensor's position and / or range (e.g., the field of view of FOV).
[0054] For example, a renderer could be a Gaussian sputter renderer. A Gaussian sputter renderer is a renderer that creates a 3D scene from multiple 2D images using Gaussian sputtering. It uses Gaussian splats to represent the scene; Gaussian splats are small, soft specks that capture identifier parameters (including color, position, and / or opacity) that are blended together.
[0055] The method may include the step of receiving optical sensor data comprising at least one subset of each first training dataset at the input layer of a first encoder. The method may also include the step of encoding each received optical sensor data of the first training dataset within a subset into a latent representation by one or more hidden layers of the first encoder. The method may further include the step of outputting the latent representation at the output layer of the first encoder.
[0056] The method may include the step of receiving an output latent representation at the input layer of a renderer. The method may also include the step of reconstructing optical sensor data by one or more hidden layers of the renderer based on the received latent representation.
[0057] This method may include pre-training the weights of a first encoder (e.g., at least one or more hidden layers) by optimizing a reconstruction loss. Optimizing the reconstruction loss may include comparing the reconstructed optical sensor data with the received sensor data.
[0058] By pre-training the first encoder separately from the task head, i.e. by using optimizations to the renderer and reconstruction loss, the pre-training of the first encoder can be performed in a time- and resource-efficient manner.
[0059] During the pre-training of the task head, the first encoder can be optionally fine-tuned.
[0060] The first encoder can be pre-trained for a specific type of optical sensor and / or sensor parameters associated with optical sensor data included in the first training dataset.
[0061] A camera's sensor parameters can include extrinsic camera parameters (hereinafter referred to as extrinsic parameters) and / or intrinsic camera parameters (hereinafter referred to as intrinsic parameters). Extrinsic parameters can represent the camera's position in a 3D scene. Alternatively or additionally, intrinsic parameters can represent the camera's optical center and / or focal length. Extrinsic parameters can be used to transform world points into camera coordinates.
[0062] Sensor calibration can refer to the process of measuring the correspondence between the sensor's output and the actual data measured by the sensor.
[0063] Sensor parameters for a LiDAR sensor may include laser wavelength, detection range, field of view (FOV), angular resolution, point rate, and / or beam count. Alternatively or additionally, sensor parameters for a LiDAR sensor may include ranging accuracy, safety level, output parameters, IP rating, power, supply voltage, laser emission mode (e.g., mechanical / solid-state), and / or lifetime.
[0064] The sensor parameters of a radar sensor may include frequency band and / or modulation.
[0065] The sensor parameters of an ultrasonic sensor may include frequency range, detection area, and / or response speed.
[0066] The sensor parameters of a thermal sensor may include resolution, pixel pitch, fill factor, magnification, FoV and / or frame rate.
[0067] In some embodiments, the renderer may be untrained (and / or may be untrainable). In other embodiments, the renderer may be trained based on at least a subset of sensor parameters.
[0068] The renderer can employ distinguishable rendering techniques. Alternatively or additionally, the renderer can apply Gaussian sputtering. Gaussian sputtering may include using the position and / or viewing angle of an optical sensor as sensor parameters. Gaussian sputtering may also include utilizing Gaussian representations.
[0069] Alternatively or additionally, the renderer may utilize surface primitives (and / or line primitives), such as polygon meshes, and / or utilize isosurfaces for rendering.
[0070] Scene rendering can include a predetermined number of parameters, such as parameters for position, orientation, color (e.g., RGB), transparency (and / or opacity), scale, and / or spherical harmonics. For example, the predetermined number of 14 parameters could appear as three parameters for position, four parameters for orientation (e.g., for quaternions), three parameters for scaling, three parameters for RGB, and one parameter for opacity.
[0071] Rendering can include textures and / or lighting.
[0072] Any encoder (e.g., a first and / or a second encoder) can be a discrete encoder. A discrete encoder can include sensor-specific encoder components (also referred to as: sensor backbone) and sensor-agnostic encoder components. For a camera that serves as at least one optical sensor, the sensor-specific encoder component can be, for example, convolution-based (e.g., residual networks, ConvNet, and / or VovNet) and / or transformer-based (e.g., ViT and / or SwingTransformers). For a radar sensor that serves as at least one optical sensor, the sensor-specific encoder component can be, for example, PillarNetwork.
[0073] As a second encoder, the separable encoder has the advantage of reducing the amount of a second training dataset required for pre-training. The sensor-agnostic parts (and / or their parameters) can be the same for both the first and second encoders, and may only need to be pre-trained once on the first training dataset.
[0074] The first component of the decoupling encoder can be Flash3d (e.g., as described in https: / / arxiv.org / abs / 2406.04343, which is incorporated herein by reference). Flash3d maps the image to a Gaussian representation in the camera frame (e.g., the Gaussian position in the camera frame). In the second step, given or estimated (e.g., trained) camera parameters can be used to transform those Gaussian parameters from the camera to the surrogate and / or ego (e.g., vehicle) coordinate system. The output of this second step can be considered as the latent representation (also known as: latent scene state) z.
[0075] The pre-trained task head and / or pre-trained first encoder may be accompanied by one or more auxiliary self-supervised learning (SSL) tasks.
[0076] Pre-training can be improved by simultaneously performing several pre-training and / or SSL tasks on the first training dataset, and / or using the first encoder. For example, it can accelerate the optimization of corresponding (e.g., SSL and / or perception) task-specific losses and / or reconstruction losses.
[0077] One or more auxiliary SSL tasks may be performed without task-specific labels and / or based on optimized self-consistency loss.
[0078] Alternatively, sensor parameters (e.g., camera parameters) can be used as input for SSL tasks such as predictive masking, predictive impainting, predictive geometric transformations, and / or temporal prediction of the next frame (e.g., for video sequences that are optical sensor data).
[0079] There are several ways to incorporate temporal evolution into the pre-training of a perception task. In one embodiment, the latent representation can be extended temporally and / or spatially by using at least one previous latent representation in addition to the current latent representation.
[0080] Scene flow and / or temporal fusion models can be used, such as long short-term memory (LSTM) based models and / or temporal convolutions.
[0081] In an alternative embodiment, a generative model can be used.
[0082] The latent representation may include, or may be provided together with, a generative model that encodes the temporal evolution of optical sensor data.
[0083] The generative model G can operate on the encoder output (e.g., latent representation) z and generate sensor states x=G(z,p) given sensor parameterization p.
[0084] The latent representation (e.g., dense and / or sparse) can be the output of (e.g., first and / or second) encoders and inputs to both the renderer (and / or the generative model G) and the task head.
[0085] Regarding the second method, a method (particularly computer-implemented) for performing a perception task is provided. The method includes the step of receiving a dataset at the input layer of a (e.g., pre-trained) second encoder, the dataset comprising optical sensor data acquired by at least one optical sensor. The method further includes the step of encoding the received optical sensor data into a latent representation by one or more hidden layers of the second encoder. The method also includes the step of outputting the latent representation by the output layer of the second encoder. The method further includes the step of receiving the output latent representation at the input layer of a task head, which is pre-trained according to the method of the first method. The method also includes the step of transforming the received latent representation into a perception task result through one or more hidden layers of the (particularly pre-trained) task head. Finally, the method further includes the step of outputting the perception task result by the output layer of the (particularly pre-trained) task head.
[0086] The second encoder can be pre-trained to provide the same type of latent representation as (e.g., pre-trained) the first encoder.
[0087] The second encoder (e.g., pre-trained) can be independent of the first encoder (e.g., pre-trained). For example, the number of nodes and / or hidden layers and / or the pre-trained weights of the first and second encoders can be different.
[0088] In an alternative embodiment, the pre-trained second encoder may have the same architecture as the first encoder, but the weights are adjusted based on an unlabeled second training dataset in a second training database.
[0089] The pre-trained second encoder may include the step of receiving multiple second training datasets from a second training database at the input layer of the second encoder. Each second training dataset may include optical sensor data acquired by at least one optical sensor. The pre-trained second encoder may also include the step of encoding the optical sensor data of each received second training dataset into a latent representation by one or more hidden layers of the second encoder. The pre-trained second encoder may also include the step of outputting the latent representation at the output layer of the second encoder. The pre-trained second encoder may also include the step of receiving the output latent representation at the input layer of a renderer. The pre-trained second encoder may also include the step of reconstructing the optical sensor data based on the received latent representation by one or more hidden layers of the renderer. The pre-trained second encoder may further include pre-training the weights of the second encoder (e.g., at least one of the multiple hidden layers) by optimizing the reconstruction loss. Optimizing the reconstruction loss may include comparing the reconstructed optical sensor data with the received sensor data.
[0090] The second training dataset may be unlabeled, and / or the second training database may include unlabeled training data.
[0091] At least one optical sensor used to acquire optical sensor data may include one or more predefined sensor types. The predefined sensor type used for pre-training the second encoder may be the same as the predefined sensor type used for retraining the task head.
[0092] The optical sensors (and optional other sensors, such as motion-related sensors) used to acquire the first training dataset and the optical sensors (and optional other sensors, such as motion-related sensors) used to acquire the second training dataset may differ at least in part in terms of the number of sensors, sensor mounting location (also known as: mounting setup), sensor model (e.g., further development and / or newer variants), sensor-specific parameters, and / or sensor calibration.
[0093] Training any (e.g., a second) encoder may include training encoder parameters and / or encoder weights.
[0094] In some embodiments, the renderer can be the same for both the pre-trained first encoder and the pre-trained second encoder. In other embodiments, the renderers used for pre-training the first encoder and the second encoder can be selected independently.
[0095] In one embodiment, for example, if the renderer is a graybox renderer, the renderer can be pre-trained using a first training database in conjunction with a first encoder and / or with a task head. The renderer's pre-training can be based on optimizing the reconstruction loss, and / or can consider pre-training on a perceptual task using perceptual task labels, and / or can consider SSL for auxiliary tasks.
[0096] In another embodiment that can be combined with the previous embodiments, a second training database can be used to jointly pre-train the renderer with a second encoder. The pre-training of the renderer can be based on optimizing the reconstruction loss, and / or can take into account the SSL of auxiliary tasks.
[0097] In some embodiments, a pre-trained renderer may include providing sensor parameters from at least one sensor as input to the renderer.
[0098] In an alternative embodiment, the sensor parameters may be learned by (e.g., a first and / or second) encoder, renderer, and / or optional SSL task.
[0099] Perception tasks may include object detection and / or classification, particularly in the vehicle coordinate system and / or robot coordinate system. Alternatively or additionally, perception tasks may include occupancy estimation and / or visibility estimation. Alternatively or additionally, perception tasks may include map estimation, estimation of the drivable area of the vehicle, scene segmentation in non-sensor-specific representations, object tracking, object velocity estimation, depth estimation and / or distance estimation, out-of-domain instance detection, and / or time prediction.
[0100] Object detection and / or classification may include detecting and / or classifying traffic signs, road surfaces, pedestrians, vehicles, and / or (e.g., 2D or 3D) dense occupancy.
[0101] Object detection and / or classification can be based on low-level features, such as edge or pixel attributes of an image as optical sensor data.
[0102] Occupancy estimates and / or visibility estimates can be based on a voxel mesh (e.g., 2D or 3D) around the vehicle.
[0103] Map estimation can include estimations of road lane areas, for example, in a vehicle coordinate system.
[0104] Segmentation can be performed in the bird's-eye view grid representation.
[0105] Depth estimation and / or distance estimation may include estimations relative to the vehicle and / or relative to the robot in a predetermined direction.
[0106] Out-of-domain (OoD) can represent a scenario and / or situation that was not observed during training and is (e.g., in some sense) very different from the available training data (e.g., a person wearing certain clothing, an oddly shaped car, etc.). OoD instances can also be referred to as corner cases or long-tailed scenarios.
[0107] The first central question is: to what extent do the encoder and / or latent scene representation (hereinafter referred to as latent representation) accurately and / or reliably represent these scenes? This can now be checked by comparing the original input x with the reconstructed input G(E(x), p), where E is the encoder.
[0108] It is possible to check whether the encoding z is true for the OoD input. This is a strong property of the potential representation z, which is not true for most common encoders, such as CNNs and / or ViTs, which typically remove a large amount of information during the encoding x->z, making it impossible to regenerate x based on z.
[0109] The general idea behind techniques for pre-training task heads for performing perception tasks is that this strong latent representation z can be used for OoD detection. For example, by applying an outlier detection network to z.
[0110] Perception tasks can be applied to autonomous driving planning, to performing emergency braking, to planning robot motion (e.g., displacement and / or grasping), and / or generally to controlling vehicles or robots.
[0111] During the inference phase of the perception task, auxiliary SSL tasks and / or renderer-assisted reconstructions can be performed. Therefore, control mechanisms can be introduced to ensure that the perception task is executed correctly.
[0112] Pre-training techniques used for the task head, first and second encoders, and performing perception tasks can be compatible with network pruning, network quantization, and / or parallelization.
[0113] Pruning can include removing parameters from a neural network system or its components, such as a task head, any encoder, and / or renderer. Alternatively or additionally, pruning can be a compression method that includes removing weights from a pre-trained (and / or trained) neural network system or any of its components.
[0114] Quantization can correspond to reducing the precision of weights, biases, and / or activations, making them consume less memory.
[0115] Parallelizing neural network training across processors (e.g., CPU and / or GPU) can involve effectively allocating workloads and synchronizing tasks to maximize computational efficiency.
[0116] Regarding the first device, a task head for performing a perception task is provided. The task head includes an input layer configured to receive a plurality of latent representations. Each latent representation is associated with a first training dataset included in a first training database. Each first training dataset includes optical sensor data acquired by at least one optical sensor. Each first training dataset also includes a perception task label. The task head also includes one or more hidden layers configured to transform each of the plurality of received latent representations into a perception task result. The task head also includes an output layer configured to output the perception task result of each of the plurality of received latent representations. The task head can be pre-trained by optimizing a task-specific loss function. Optimizing the task-specific loss function may include comparing the output perception task result with the perception task label of each of the plurality of received latent representations. Alternatively or additionally, the task head can be pre-trained according to a method according to the first method aspect.
[0117] The task head can be configured to perform any of the steps described in the context of the first method, or to include any of the features. For example, the task head can include a pre-trained task-specific loss function.
[0118] The output layer (and / or output head) can be pre-trained on the source domain (and / or using a first training database), specifically leveraging available labels from the source domain. In a new target domain (e.g., for a different sensor setup), those pre-trained output layers (and / or task heads) can be reused without modification.
[0119] Generally, task-specific losses can also be used to train the complete architecture, including the encoder. However, this can lead to invariance in the encoder, resulting in higher reconstruction errors when using generative models. Therefore, task heads and / or losses are typically not used to train the encoder as well.
[0120] Regarding the second device aspect, a neural network system for performing a perception task is provided. The neural network system includes (e.g., pre-trained) a second encoder. The second encoder includes an input layer configured to receive a dataset comprising optical sensor data acquired by at least one optical sensor. The second encoder also includes one or more hidden layers configured to encode the received optical sensor data into a latent representation. The second encoder further includes an output layer configured to output the latent representation. The neural network system also includes a task head according to the first device aspect. The input layer of the task head is configured to receive the latent representation output by the second encoder. One or more hidden layers of the task head are configured to transform the received latent representation into a perception task result. The output layer of the task head is configured to output the perception task result.
[0121] The neural network system can be configured to perform any of the steps described in the context of the second method, or to include any of the features. For example, the neural network system may also include a renderer. Alternatively or additionally, the neural network system may also include at least one optical sensor.
[0122] The task head and / or neural network system can be implemented in a graphics processing unit (GPU) and / or a tensor processing unit (TPU), which is an application-specific integrated circuit (ASIC) for AI accelerators.
[0123] On another front, a computer program product is provided, comprising a program unit that, when loaded into the memory of a task head and a first encoder, causes the task head and the first encoder to perform steps of a method for pre-training a task head for a perception task according to a first method aspect. Alternatively or additionally, when the program unit is loaded into the memory of a computing device, the program unit causes a neural network system to perform steps of a method for performing a perception task according to a second method aspect.
[0124] In another aspect, a computer-readable medium is provided having program units stored thereon, which can be read and executed by a task head and a first encoder, such that when the program units are executed by the task head and the first encoder, the steps of the method for pre-training a task head for a perception task according to a first method aspect are performed. Alternatively or additionally, the program units can be read and executed by a neural network system, such that when the program units are executed by the neural network system, the steps of the method for performing a perception task according to a second method aspect are performed. Attached Figure Description
[0125] Figure 1 This is an example flowchart of a method for pre-training task heads for perception tasks.
[0126] Figure 2 This is an example flowchart of a method for performing perception tasks.
[0127] Figure 3 An example architecture for a task header is shown, which can be used... Figure 1 The method is pre-trained to perform perception tasks.
[0128] Figure 4 It shows how to use, for example, according to Figure 2 An example architecture of a neural network system that performs perception tasks using this method.
[0129] Figure 5 The diagram schematically illustrates a combination of a first encoder, a renderer, and a task head for pre-training, used, for example, according to... Figure 1 The method of pre-training.
[0130] Figure 6 The diagram schematically illustrates a combination of a second encoder, a renderer, and a pre-trained task head, used, for example, according to... Figure 2 The method performs perception tasks.
[0131] Figure 7 The pre-training of the task head is further illustrated, for example, based on... Figure 1 The method.
[0132] Figure 8 The pre-training of the second encoder is further illustrated, for example, based on... Figure 2 The method.
[0133] Figure 9 An example of the inference phase is illustrated, which uses a pre-trained task head and a pre-trained second encoder, for example, according to... Figure 1 and 2 The method.
[0134] Any reference numerals in the claims should not be construed as limiting the scope. Detailed Implementation
[0135] Figure 1 An example flowchart of method 100 for pre-training task heads for perception tasks (e.g., computer-implemented) is shown.
[0136] Method 100 includes a step S112 of receiving a plurality of latent representations at the input layer of a task head. Each latent representation is associated with a first training dataset included in a first training database. Each first training dataset includes optical sensor data acquired by at least one optical sensor and a perception task label. Method 100 also includes a step S114 of transforming each of the plurality of latent representations received in S112 into a perception task result. Step S114 is performed by one or more hidden layers of the task head. Method 100 also includes a step S116 of outputting a perception task result for each of the plurality of latent representations received in S112. Step S116 is performed by the output layer of the task head. Method 100 also includes a step S118 of pre-training the weights of at least one or more hidden layers of the task head by optimizing a task-specific loss function. Optimizing the task-specific loss function includes comparing the perception task result output in S116 with the perception task label for each of the plurality of latent representations received in S112.
[0137] Method 100 may further include a step S113 of receiving a sensing task label associated with each of the plurality of potential representations received in S112. The sensing task label may be received in S113 at the loss function of the task header.
[0138] Each potential representation of S112 can be received from the pre-trained first encoder.
[0139] The pre-trained first encoder may include step S102 of receiving optical sensor data included in at least one subset of each first training dataset. The optical sensor data of S102 may be received at the input layer of the first encoder. The pre-trained first encoder may also include step S104 of encoding the optical sensor data of each of the first training datasets received in the subset into a latent representation. Encoding may be performed by one or more hidden layers of the first encoder. The pre-trained first encoder may also include step S106-0 of outputting the latent representation. The latent representation may be output by the output layer of the first encoder. The pre-trained first encoder may also include step S106-I of receiving the latent representation of output S106-O at the input layer of a renderer. The pre-trained first encoder may also include step S108 of reconstructing the optical sensor data based on the latent representation received in S106-I through one or more hidden layers of the renderer. The pre-trained first encoder may also include pre-training the weights of at least one or more hidden layers of the first encoder by optimizing the reconstruction loss. Optimizing the reconstruction loss may include comparing the reconstructed optical sensor data of S108 with the sensor data received in S102.
[0140] Figure 2 An example flowchart of method 200 (e.g., computer-implemented) for performing a perception task is shown. Method 200 includes a step S212 of receiving a dataset, comprising optical sensor data acquired by at least one optical sensor, at the input layer of a (particularly pre-trained) second encoder. Method 200 also includes a step S214 of encoding the optical sensor data received in S212 into a latent representation. Step S214 is performed by one or more hidden layers of the second encoder. Method 200 also includes a step S216 of outputting the latent representation at the output layer of the second encoder. Method 200 also includes a step S218 of receiving the latent representation output in S214 at the input layer of a pre-trained task head. The task head can be pre-trained according to method 100. Method 200 also includes a step S220 of transforming the latent representation received in S218 into a perception task result. Step S220 is performed by one or more hidden layers of the pre-trained task head. Method 200 also includes a step S222 of outputting the perception task result at the output layer of the pre-trained task head.
[0141] The pre-training of the second encoder may include step S202, which involves receiving multiple second training datasets from a second training database at the input layer of the second encoder. Each second training dataset may include optical sensor data acquired by at least one optical sensor. The pre-training of the second encoder may also include step S204, which involves encoding the optical sensor data of each second training dataset received in S202 into a latent representation. Encoding the latent representation in S204 may be performed by one or more hidden layers of the second encoder. The pre-training of the second encoder may also include step S206-0, which involves outputting the latent representation at the output layer of the second encoder. The pre-training of the second encoder may also include step S206-1, which involves receiving the latent representation of output S206-0 at the input layer of a renderer. The pre-training of the second encoder may also include step S208, which involves reconstructing the optical sensor data based on the latent representation received in S206-1. Reconstruction S208 may be performed by one or more hidden layers of the renderer. The pre-training of the second encoder may also include pre-training the weights of at least one of the multiple hidden layers of the second encoder in S210 by optimizing the reconstruction loss. Optimizing the reconstruction loss may include comparing the optical sensor data reconstructed in S208 with the sensor data received in S202.
[0142] The renderer used for pre-training the first encoder (for pre-training the task head) and the renderer used for pre-training the second encoder (for performing the perception task) can be selected independently. In some embodiments, the renderer weights of the renderer can be fixed (e.g., for a white-box renderer), while in other embodiments, the renderer can be trained at least partially with the corresponding encoder.
[0143] Figure 3 An example architecture for a task head 300 used to perform a perception task is shown. The task head can be pre-trained according to method 100.
[0144] The task head 300 includes an input layer 312 configured to receive a plurality of latent representations. Each latent representation is associated with a first training dataset included in a first training database. Each first training dataset includes optical sensor data acquired by at least one optical sensor and a perception task label. The task head 300 also includes one or more hidden layers 314 configured to transform each of the plurality of received latent representations into a perception task result. The task head 300 also includes an output layer 316 configured to output the perception task result of each of the plurality of received latent representations.
[0145] The task head 300 can be pre-trained by optimizing a task-specific loss function. Optimizing the task-specific loss function may include comparing the output perceptual task result with the perceptual task label of each of a plurality of received latent representations. Alternatively or additionally, the task head 300 can be pre-trained by method 100.
[0146] The task head 300 may include a task-specific loss function 313, whereby a perceptual task label associated with each of the multiple received latent representations is received.
[0147] The task head 300 may include memory 319.
[0148] The task header 300 may include an input / output interface 321. The input layer 312 and / or the output layer 316 may be implemented by the input / output interface 321.
[0149] The task head 300 may include a processor 323. One or more hidden layers 314 and / or an optional task-specific loss function 313 may be implemented by the processor 323.
[0150] Figure 4 An example architecture of a neural network system 400 for performing a perception task is shown. The neural network system 400 includes a second encoder 401. The second encoder 401 can be pre-trained. The second encoder includes an input layer 402 configured to receive a dataset comprising optical sensor data acquired by at least one optical sensor. The second encoder 401 also includes one or more hidden layers 404 configured to encode the received optical sensor data into a latent representation. The second encoder 401 also includes an output layer 406 configured to output the latent representation.
[0151] The second encoder 401 may include an input / output interface 407. The input / output interface 407 may represent an input layer 402 and / or an output layer 406.
[0152] The second encoder 401 may include a processor. The processor may embody one or more hidden layers 404.
[0153] The second encoder 401 may include a memory 409.
[0154] The neural network system 400 also includes a task head 300. The input layer 312 of the task head 300 is configured to receive a latent representation output by the second encoder 401. One or more hidden layers 314 of the task head are configured to transform the received latent representation into a perception task result. The output layer 316 of the task head 300 is configured to output the perception task result.
[0155] The neural network system 400 can be configured to perform several perceptual tasks. Alternatively or additionally, the neural network system 400 may include multiple task heads 300.
[0156] The neural network system 400 may also include a renderer 411.
[0157] The renderer 411 may include an input layer 414 configured to receive a potential representation output by the second encoder 401.
[0158] The renderer 411 may also include one or more hidden layers 414 configured to reconstruct optical sensor data based on the received latent representations.
[0159] The renderer 411 may also include an output layer 416, which is configured to output reconstructed optical sensor data.
[0160] The renderer 411 may include an input / output interface 417. The input / output interface 417 may represent an input layer 412 and / or an output layer 416.
[0161] The renderer 411 may include a processor. The processor may include one or more hidden layers 414.
[0162] The renderer 411 may include memory 419.
[0163] The neural network system 400 also includes a reconstruction loss function 421.
[0164] The second encoder 401 (and / or the weights of at least one of its multiple hidden layers) can be pre-trained by optimizing the reconstruction loss. Optimizing the reconstruction loss may include comparing the optical sensor data reconstructed by the renderer 411 with the sensor data received by the second encoder 401.
[0165] The neural network system 400 may also include a set of sensors 422.
[0166] Techniques (e.g., including method 100, method 200, task head 300, and / or neural network system 400) can alternatively be represented as domain transfers with a sensor-agnostic world model.
[0167] The technology can use task-agnostic and sensor-agnostic latent representations (also known as world representations) to serve as a common interface between different sensor and downstream task settings.
[0168] like Figure 7 As illustrated in the example, the world model can include a latent representation z representing the current scene (e.g., some) plus a generative model. Given some sensor parameters (For example, outside and / or inside the camera), which can generate corresponding sensor inputs. : .
[0169] In other words, generative models Combine potential z and sensor parameters As input, and reconstruct sensor input .
[0170] like Figure 8 The key idea illustrated in the example is that, for a new target domain (For example, for a new sensor with the same modality as a previous sensor, such as a new camera with a higher resolution than the old camera) the converter and / or encoder are transformed via self-supervised loss. Learned in the latent (also known as: world model) representation z: .
[0171] This article, These are sensor parameters (e.g., in-camera and / or out-of-camera) in the target domain, which are known from the system settings or can be considered as additional, trainable encoder parameters. It is a reconstruction loss in the input sensor domain, and These are trainable encoder parameters.
[0172] According to this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400), a fairly general approach is applied to how to transfer different aspects from multiple potential source domains to a new target domain, including, for example, transferring a task head trained on labeled source domain data to a new target domain without requiring labels in the target domain at all. Because this technique is quite general, many solutions exist to fill certain gaps.
[0173] The source domain may include a set (e.g., at least one) of optical sensors with associated sensor parameters. A first training database may include a first training dataset, each first training dataset including optical sensor data acquired by at least one optical sensor in the set. Each first training dataset may also include at least one perception task label. In some embodiments, several task heads may be trained to perform different perception tasks, and the first training database may include a first training dataset including multiple perception task labels for different perception tasks. The size of the first training database may be limited by the availability of perception task labels.
[0174] The target domain may include another set (e.g., at least one) of optical sensors with relevant sensor parameters. The sensor mode (also referred to as: sensor type) in the target domain may be the same as the sensor mode (and / or sensor type) in the source domain. The optical sensors (also simply referred to as: sensors) may differ in the values of their sensor parameters (e.g., resolution) in the source and target domains.
[0175] Within the target domain, a second training dataset can be obtained, especially in the absence of any perceptual task labels. Therefore, it is possible to efficiently generate (e.g., extensive) second training databases that include the second training dataset.
[0176] The main advantage of this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) is that it can pre-train task heads. The target domain (and / or sensor settings) can be applied from the source domain S (and / or sensor settings) to a new target domain (and / or modified sensor settings) without requiring any labels or pseudo-labels in the target domain. Therefore, the target domain can be changed relative to the source domain sensor settings, particularly in the following ways: the number of sensors, the mounting location of one or more sensors, and / or sensor-specific parameters, such as those for a camera, such as in-camera, out-of-camera, and / or image resolution.
[0177] Alternatively or additionally, relative sensor calibration in the target domain can be learned in a data-driven manner, for example, sensor parameters. It can also be considered as trainable encoder parameters The part.
[0178] A key application of this technology (e.g., including method 100, method 200, task head 300, and / or neural network system 400) is to implement strongly pre-trained downstream perception tasks (e.g., object detection, occupancy estimation, segmentation, tracking, etc.) on a new target domain without any labels from the target domain (and / or through zero-trigger learning). This can be applied to reuse strongly pre-trained perception task heads in new sensor setups, for example, for other manufacturers (and / or customers) whose purpose is to perform perception tasks using optical sensors.
[0179] This technology (e.g., including method 100, method 200, task head 300, and / or neural network system 400) can be used to analyze data obtained from (particularly optical) sensors. The sensors can determine environmental measurements in the form of sensor signals, which can be provided by, for example, digital images such as video, radar, LiDAR, ultrasound, motion, and / or thermal images.
[0180] Generally, all non-sensor-specific downstream tasks in the context of autonomous driving and robotics can be considered. "Non-sensor-specific" may mean (e.g., perception) that the task output is not allowed in the sensor coordinate system. For example, semantic segmentation will not necessarily fall into the category of techniques (e.g., including method 100, method 200, task head 300, and / or neural network system 400) because its output is a pixel-by-pixel classification and therefore depends on specific camera settings.
[0181] Alternatively or additionally, the following downstream (e.g., perception) tasks may fall under this technology (e.g., including method 100, method 200, task head 300, and / or neural network system 400): object detection and / or classification in the vehicle coordinate system (e.g., 3D); occupancy and / or visibility estimation (e.g., in a 2D or 3D voxel grid around the vehicle); map estimation, such as estimation of road lane areas in the vehicle coordinate system; estimation of drivable areas relative to the vehicle; object tracking; segmentation in non-sensor-specific representations, such as bird's-eye view mesh representations; object velocity estimation; depth and / or distance estimation in a specific direction relative to the vehicle; and / or out-of-domain instance detection.
[0182] Virtual sensors can be used. Alternatively or additionally, information about elements encoded by sensor signals can be obtained based on (e.g., optical) sensor signals (e.g., indirect measurements can be performed based on sensor signals used as direct measurements).
[0183] Video and / or audio analysis (and / or any further analysis of sensor data) can be performed, for example, for classification and / or regression. This technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) can be used, for example, to classify sensor data, detect the presence of objects, such as traffic signs, road surfaces, pedestrians, vehicles, and / or 2D (and / or 3D) dense occupancy. Classification can be performed based on low-level features (e.g., edge or pixel attributes of an image). Alternatively or additionally, this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) can be used to determine continuous values or multiple continuous values, and / or to perform regression analysis, such as regarding distance, speed, and / or items in tracking data, such as objects. The determination and / or regression of continuous values can be performed based on low-level features (e.g., edge or pixel attributes of an image).
[0184] This technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) can be considered an upstream part of the machine learning (ML) toolchain. It does not necessarily directly improve the ML system that can be used for the aforementioned applications, but it is a method for training such an ML system, and / or a method for training a task head and / or encoder (and / or converter) for new input sensor settings.
[0185] After being trained in this way, the ML system can then be used downstream.
[0186] All standard techniques for network pruning, network quantization, or parallelization can be applied to this technique (e.g., including Method 100, Method 200, Task Header 300, and / or Neural Network System 400). This technique (e.g., including Method 100, Method 200, Task Header 300, and / or Neural Network System 400) itself does not inherently require permission for additional hardware optimizations.
[0187] The technology (e.g., including method 100, method 200, task head 300 and / or neural network system 400) can be a component of the perception and fusion architecture of an autonomous system, and / or can be used in various system settings and downstream applications.
[0188] The techniques (e.g., including method 100, method 200, task head 300 and / or neural network system 400) use latent representations and / or world models, such as those described in [7], [8], [9],
[10] .
[0189] For example, as described in
[11] , Gaussian sputtering and / or scene rendering can be utilized.
[0190] In one embodiment, the setup includes multiple tasks (e.g., object detection, time prediction, etc.), represented as follows: There exist sensor-specific (e.g., camera) parameters. At least one source domain (e.g., optical) sensor. Source domain data. Including optical sensor data and may have multiple source domains Prediction task tags It provides information from sensors with certain parameters. The renderer and / or generator that can potentially represent z as a mapping to the corresponding sensor signal (also known as: optical sensor data). Neural network (e.g., first and / or second) encoder Configured for receiving sensor signals Mapping to the latent representation z. Target domain data. Unmarked.
[0191] The goal is to get the task head From data with source domain source domain Apply and transfer data to new and unlabeled target domains .
[0192] This technique can be divided into (particularly temporary) three phases. In the first phase (also known as: the task head pre-training phase), such as Figure 5 As illustrated, one or more task heads 300 (or 300-1; 300-2; 300-3) are pre-trained (e.g., according to method 100) on a (e.g., sensor and / or task) agnostic (e.g., world model) latent representation z, which is in... Figure 5 As indicated by reference numeral 508 in the accompanying drawings, the data is obtained from source domain data 502-1, ..., 502-K, which is received from sources (and / or, in particular, optical sensors) 504-1, ..., 504-K and encoded by a first encoder 506. A renderer (also called a decoder) 510 provides reconstructed data 512-1. As indicated by the snowflake symbol, the renderer 510 can be a white-box renderer. As shown by the arrow from the optical sensor 504-1 to the renderer 510, sensor-specific parameters can be used for rendering.
[0193] Figure 7 An alternative schematic diagram for the first stage is provided. Figure 7 In the diagram, parallel to one or more task heads 300 trained via supervised loss (and / or perceptual task-specific loss) 313, an auxiliary SSL task head 702 can be trained without labels and using self-consistent loss 706, while reconstruction loss 708 is used to train the first encoder 506. Consecutive arrows illustrate the primary use of losses 313, 706, and 708. Dashed arrows indicate further potential impacts of losses 313, 706, and 708 on other components.
[0194] The first phase may resemble a standard multitasking setup. (See reference) Figure 5 and 7 In an embodiment, the first encoder (and / or the source domain encoder f) S )506 The sensor input at reference numbers 502-1, ..., 502-K in the attached diagram is... Its potential representation z mapped to reference numeral 508 in the attached figure: The multi-task head at reference numerals 300; 300-1, ... 300-3 in the attached diagram. (Also abbreviated as: The latent representation z 508 is mapped to the corresponding task output (and / or the execution of the corresponding downstream source task), which is labeled with the corresponding tag at reference numeral 313. Supervision. The auxiliary self-supervised (SSL) task at reference 702 in the attached figure. The latent representation z 508 is mapped to the corresponding output, which is supervised by self-supervision as indicated by reference numeral 706. SSL task 702 may or may not take sensor-specific (e.g., camera) parameters as input. Examples of SSL tasks include masking prediction, partial rendering, geometric transformation, and / or temporal prediction for the next frame. The renderer (also referred to as: decoder) is indicated by reference numeral 510. The reconstruction of the sensor-agnostic latent representation z 508 to a certain sensor input 512; 512-1 is supervised by the reconstruction loss 708 between the input 502; 502-1 and the output 512; 512-1: And can be trained accordingly. In an alternative embodiment, the decoder 510 may be a white-box renderer without additional and / or trainable parameters, i.e. It can be empty.
[0195] The distinguishing component of this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) (e.g., different from standard multitasking setups) may be a dedicated decoder 510 for signal reconstruction, which may be a white-box or gray-box renderer that takes a latent representation 508 and a set of sensor (and / or camera) parameters as input and maps them to the corresponding sensor signals 512.
[0196] The key to this technique (e.g., including method 100, method 200, task head 300 and / or neural network system 400) is that the resulting latent representation 508 can be a so-called “world model” that is both task-agnostic and domain and / or sensor-agnostic, for example, both due to multi-task 300-1; 300-2; 300-2 and additional self-supervised 702 training.
[0197] Note that in the context of this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400), domain and / or sensor agnosticity does not mean (or necessarily means) that the technique is expected to generalize from a purely RADAR source domain to entirely different physical sensor domains, such as vision (and / or camera sensors). Instead, in the context of this technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400), camera sensors in the source domain can be used to generalize to, for example, different video domains (e.g., in terms of the number of sensors, sensor mounting locations, and / or sensor-specific parameters). Alternatively or additionally, in cases where the source domain includes a mixture of different modalities (e.g., RADAR and vision / camera and Lidar) (e.g., by combining different datasets), the latent representation (and / or world model), and therefore the task head, can be transferred from that source domain to, for example, new mounting settings for those sensors or different camera models.
[0198] An example implementation of the white-box renderer 510 in the visual domain is Gaussian sputtering (GS), which renders camera images using the camera position and / or viewpoint as sensor parameters and a Gaussian representation. The GS renderer requires no training. Note that in the case of a white-box renderer only, during the first phase, only the downstream task head will be trained in a supervised manner on the white-box world representation.
[0199] In the second stage (also known as the target domain encoder and / or the second encoder, the pre-training stage), the transfer to the target domain is performed.
[0200] refer to Figure 6 and 8 In the example above, for a new target domain 604, the goal is to train a target domain-specific encoder (also known as: a second encoder). As shown by reference numeral 401 in the attached figure, it maps the target domain data 602 to a latent (also known as: world model) representation z 508. The two main losses for training the target domain-specific encoder (and / or the second encoder) 401 are a self-supervised loss 806 that can be applied to an additional SSL task 802 (which can be independently selected from the first-stage SSL task 702) for training the unlabeled target domain, and a reconstruction loss 421 given by the decoder (and / or renderer) 411 used to train the second encoder 401. .
[0201] The key in the second stage could be the renderer at reference 411 in the attached diagram. It can be fixed, i.e., the decoder parameters. It is possible to not train, which means that the meaning of the latent representation 508 remains the same across domains. Therefore, the task headers 300; 300-1; 300-2; 300-3 from the source domain are expected to generalize without further training.
[0202] The renderer 411 can be effective without modification and may be able to reconstruct the target domain sensor signal 602 from the latent representation (also known as: latent code) 508 by taking the target domain sensor parameter 604 as additional input, and obtain the reconstructed sensor signal 612.
[0203] The trainable part in the second stage can be (e.g., domain-specific and / or) a second encoder 401 that can correctly map the target domain to an agnostic latent (also known as: world model) representation 508.
[0204] In such Figure 9 The example illustrates the third stage (also known as: the inference stage and / or the inference target domain), and / or during inference time (particularly over the target domain), the model (and / or neural network system 400) includes a target domain encoder (also known as: the second encoder) at reference 401. It maps the target domain input 602 to an unknowable latent (also known as: world model) representation 508. The model (and / or neural network system 400) also includes one or more task heads that map from the latent representation 508 to the corresponding task outputs. .
[0205] During the inference phase, the given input at reference 602 in the target domain... First, the trained target domain (also known as: second) encoder 401 is mapped to the latent (also known as: world) representation z at the attached label 508, and then the downstream task head 300 can be applied to z.
[0206] exist Figure 9 The example auxiliary SSL task header 802' is shown in the figure. The SSL task header 802' in the inference phase can be the same as the SSL task header 802 in the training phase 401 of the second encoder.
[0207] exist Figure 9 The diagram illustrates an example application 902 of autonomous driving planning, wherein an autonomous vehicle 908 includes one or more sensors 604, a neural network system 400 incorporated in a computing unit 902 of the vehicle 908, and downstream tasks configured to control actuators 904. Further example applications 902 include emergency braking, distance estimation, out-of-domain activities, and / or planning and / or grasping for a robot.
[0208] Note that reference numerals 504 and 604 are used for both the sensor and its sensor-specific parameters for ease of labeling.
[0209] In one embodiment, the sensor parameters of the target domain sensor are defined. In the second phase, in addition to pre-training the target domain (also known as the second) encoder, the parameters can also be learned in the unknown context. These trainable parameters should be consistently used for the target domain (also known as the second) encoder, the decoder (also known as the renderer), and the potential SSL task that depends on the sensor parameters.
[0210] In another embodiment, the technique can be temporarily extended. In most real-world applications of autonomous agents (and / or autonomous driving), inputs are available in a streaming and / or time-varying manner, where new sensor inputs subsequently enter the system. And the world indicates z t It needs to be updated accordingly in a temporal manner. To capture this temporal aspect, the encoders used for the source and / or particularly the target domain (e.g., the first and / or second) can be extended to take into account previous world representations: In this embodiment, it is also beneficial to extend the latent representation (also known as the world representation) with, for example, a dynamic representation of the scene flow.
[0211] Note that mixing between the source and target domains is possible. For example, for some downstream tasks, only static (and / or non-temporal) datasets may be available. Therefore, during the first phase of pre-training the task head, only a static (e.g., first) encoder can be used in the source domain. Alternatively or additionally, a temporal (e.g., second) encoder can also be used in the target domain.
[0212] Simpler versions of this temporal fusion model include, but are not limited to, LSTM models and / or temporal convolutions.
[0213] In another embodiment, a world model predictor is applied. In the case of a time encoder, the time encoder is constructed as a world model predictor. It is beneficial, as it occurs over a certain continuous (and / or successive) time interval. Predicting the future from previous potential representations (also known as: world state): .
[0214] Note that the world model predictor is sensor-agnostic and / or independent of the sensor domain T (and / or S). Therefore, the predictor p can be trained only in the first phase of the pre-trained task head on the source domain (in this case, it needs to include temporal data, such as video data), and importantly, p can be fixed in the second phase of the pre-trained second encoder and / or can be trained outside the target domain.
[0215] Depending on the properties of the world representation z, different architectures of the predictor p may be beneficial. For example, in the case of dense components of z, a convolution-like dense architecture may be advantageous. In the case of sparse and / or token-based components of z, a transformer-based architecture may be advantageous.
[0216] In another embodiment, a domain-specific backbone can be applied. Since the world representation (also known as: latent representation) can be very abstract, it may be necessary to apply a domain-specific backbone to encoders with a large number of trainable parameters. This involves a fairly large architecture and / or model. To mitigate this and further reduce the amount of data required to train the target domain encoder, encoders are typically divided into sensor-specific encoders (often referred to as the sensor backbone) and sensor-agnostic encoder components. .This article, These are the trainable parameters of a sensor-agnostic world model. Sensor-agnostic model components. The task head only needs to be trained during the first phase of pre-training on the source domain, and / or can be kept fixed during the second phase of training on the target domain. In this way, f is sensor-agnostic.
[0217] In the context of this invention, the term "world model" can be used to indicate that a latent representation is sensor, domain, and / or task-agnostic. That is, it can model and / or represent the world (also called: scene) largely independently of the sensor setup used and / or the desired output task. Therefore, a "world model" can be a descriptive property of the latent representation being described, and does not need to be a concept in itself, e.g., it does not need to exist as any function, etc., that is a "world pattern." In this sense, the world model also does not need to be part of anything else. It is possible that the encoder and renderer can be considered (at least to some extent) inverses of each other, but this does not need to be directly related to the term "world model." It is possible (e.g., more precisely) through the flexibility of using encoders and renderers on different sensor setups that the latent representation z becomes a world model.
[0218] Otherwise, the encoder portion that depends on the sensor can also be formulated in a sensor-agnostic way, which to some extent makes it (or they) part of the “world model” representing z as (or representing).
[0219] Sensor-dependent encoders It can be trained on the source domain S or the target domain T, respectively.
[0220] Depending on the physical properties of the sensor, different domain-specific backbones may be beneficial. For example, convolution-based architectures (e.g., residual networks, ConvNext, VovNet, etc.) or transformer-based architectures (e.g., ViT, SwinTransformers) may be beneficial for camera inputs, while they may be beneficial for RADAR PillarNetwork or other simpler architectures.
[0221] According to further embodiments, different variations of the world representation can be used.
[0222] Modular world representations can be beneficial, for example, if the world and / or scene representation z has some internal structure with different complexities. This range can be from a single BEV (bird's-eye view) representation of the scene around the vehicle to several components representing different aspects of the scene, such as scene dynamics (e.g., scene flow, object flow, etc.) and / or scene appearance.
[0223] More modular (e.g., static versus time, and / or dense versus sparse) world representations (e.g., z = (z static , z dynamic (Especially compared to the unstructured "blockbox" representation) may be most beneficial because it allows for more SSL loss, is more transparent, and is more interpretable. Alternatively or additionally, a more structured and / or modular potential representation z may be beneficial, for example, consider z = (z... static , z dynamic ), where z is modularized into parts representing static scene features (e.g., where objects are, what they look like, etc.) and parts representing static scene features (e.g., in which direction and / or at what speed objects are moving).
[0224] According to another embodiment, multiple complementary source task domains for task head pre-training can be used in the first stage. For example, one source domain may include 3D object bounding box labels for certain categories, while another source domain may include labels for other and / or additional categories. This is beneficial in two main aspects. This embodiment can be more sensor-agnostic. The more and / or more diverse the available sensors during the first stage of pre-training the task head, the more general the world representation and decoder become, and the potential target domain is enabled through diverse source domain settings. Alternatively or additionally, this embodiment can be more task-agnostic. Downstream tasks from all used source domains (e.g., datasets) can also become available as downstream tasks in the target domain, thereby enhancing and extending the system's capabilities on the target domain.
[0225] Unless otherwise stated, all the embodiments given herein can be combined with each other. Combining several or all of the above embodiments may be particularly (or most) advantageous.
[0226] For example, specifically, combining a world model predictor and a domain-specific backbone can shift model complexity from sensor-specific encoders to sensor-agnostic world model components. This reduces the complexity of the sensor encoder. The complexity and / or size of the target domain may be advantageous, thus training the target domain encoder (and / or the second encoder) may require less data, especially in the target domain.
[0227] This technique (e.g., including method 100, method 200, task head 300, and / or neural network system 400) is specifically designed for transfer from a source domain to a target domain. This technique can be used, for example, in the context of autonomous driving.
[0228] refer to [1] Domain generalization: A survey (https: / / arxiv.org / abs / 2103.02503) [2] Real-Time Monocular Depth Estimation Using Synthetic Data WithDomain Adaptation via Image Style Transfer (https: / / openaccess.thecvf.com / content_cvpr_2018 / html / Atapour-Abarghouei_Real-Time_Monocular_Depth_CVPR_2018_paper.html) [3] Locality Preserving Joint Transfer for Domain Adaptation(https: / / ieeexplore.ieee.org / abstract / document / 874682) [4] https: / / www.cityscapes-dataset.com / [5] https: / / www.nuscenes.org / nuscenes [6] https: / / arxiv.org / abs / 2406.04343 [7] https: / / arxiv.org / abs / 1803.10122 [8] https: / / www.researchgate.net / publication / 354728319_World_model_learning_and_inference [9] https: / / openreview.net / pdf id=BZ5a1r-kVsf
[10] https: / / arxiv.org / pdf / 2403.00504
[11] https: / / openaccess.thecvf.com / content / CVPR2024 / html / Wu_4D_Gaussian_Splatting_for_Real-Time_Dynamic_Scene_Rendering_CVPR_2024_paper.html.
Claims
1. A method (100) for pre-training a task head for a perception task, comprising the following method steps: - Receive (S112) multiple latent representations at the input layer of the task head, wherein each latent representation is associated with a first training dataset included in a first training database, wherein each first training dataset includes optical sensor data acquired by at least one optical sensor, and wherein each first training dataset also includes a perception task label; - Each of the potential representations of multiple receptions (S112) is transformed (S114) into a perceptual task result by one or more hidden layers of the task header; - The perceptual task result of each of the multiple latent representations received (S112) is output (S116) at the output layer of the task header; and - Pre-train (S118) the weights of at least one or more hidden layers of the task head by optimizing a task-specific loss function, wherein optimizing the task-specific loss function includes comparing the output (S116) of the perceptual task result with the perceptual task label for each of the potential representations of a plurality of receivers (S112).
2. The method (100) according to claim 1 further includes the following step: - At the loss function of the task head, receive (S113) the perceptual task label associated with each of the potential representations of the multiple receptions (S112).
3. The method (100) according to any of the preceding claims, wherein At least one optical sensor includes at least one of the following: - camera; - LiDAR sensor; - Radar sensor; - Ultrasonic sensors; and - Thermal sensor.
4. The method (100) according to any one of the preceding claims, wherein, At least one subset of the first training dataset also includes motion sensor data, position sensor data, and / or navigation sensor data.
5. The method (100) according to the preceding claim, wherein, Motion sensor data is acquired through motion sensors, and / or position sensor data and / or navigation sensor data is acquired through a satellite-based navigation system.
6. The method (100) according to any one of the preceding claims, wherein, Multiple latent representations are obtained through a pre-trained first encoder.
7. The method (100) according to the preceding claim, wherein, The pre-trained first encoder includes the following steps: - The input layer of the first encoder receives (S102) optical sensor data including at least one subset of each first training dataset; - Optical sensor data of the first training dataset received in the subset (S102) is encoded (S104) into a latent representation by one or more hidden layers of the first encoder; - Output the latent representation (S106-O) at the output layer of the first encoder; - Receive the potential representation of the output (S106-O) at the input layer of the renderer; - Optical sensor data (S108) is reconstructed by one or more hidden layers of the renderer based on the latent representation of the received data (S106-I); and - Pre-train (S110) the weights of at least one or more hidden layers of the first encoder by optimizing the reconstruction loss, wherein optimizing the reconstruction loss includes comparing the reconstructed (S108) optical sensor data with the received (S102) sensor data.
8. The method (100) according to any one of the preceding claims, wherein, The pre-trained task head and / or the pre-trained first encoder are accompanied by one or more auxiliary self-supervised learning SSL tasks.
9. The method (100) according to any one of the preceding claims, wherein, the following At least one: - The potential representation includes or is provided together with a generative model that encodes the temporal evolution of optical sensor data; - The first encoder and / or any other encoder includes or is provided with a generative model that encodes the temporal evolution of optical sensor data; and - The task head and / or renderer includes or is provided with a generative model that encodes the temporal evolution of optical sensor data.
10. A computer-implemented method (200) for performing a perception task, comprising the following method steps: - A dataset (S212) is received at the input layer of a pre-trained second encoder, the dataset comprising optical sensor data acquired by at least one optical sensor; - The received optical sensor data (S212) is encoded (S214) into a latent representation by one or more hidden layers of the second encoder; - The latent representation output by the output layer of the second encoder (S216); - Receive (S218) a latent representation of the output (S216) at the input layer of the task head, the task head being pre-trained by the method according to any one of claims 1 to 10; - One or more hidden layers of the pre-trained task head will receive (S218) the latent representation transformation (S220) into a perceptual task result; and - The task result is perceived by the output layer of the pre-trained task head (S222).
11. The method (200) according to the preceding claim, wherein, The pre-trained second encoder includes the following steps: - The second encoder receives (S202) multiple second training datasets included in the second training database in the input layer, wherein each second training dataset includes optical sensor data acquired by at least one optical sensor; - The optical sensor data of each received (S202) second training dataset is encoded (S204) into a latent representation by one or more hidden layers of the second encoder; - The latent representation is output (S206-O) at the output layer of the second encoder; - Receive the potential representation of the output (S206-O) at the input layer of the renderer; - Optical sensor data (S208) is reconstructed by one or more hidden layers of the renderer based on the latent representation of the received data (S206-I); and - The weights of at least one of the multiple hidden layers of the second encoder are pre-trained (S210) by optimizing the reconstruction loss, wherein optimizing the reconstruction loss includes comparing the reconstructed (S208) optical sensor data with the received (S202) sensor data.
12. The method (200) according to claim 10 or 11, wherein, Perception tasks include at least one of the following: - Object detection and / or classification, especially in the vehicle coordinate system and / or robot coordinate system; - Occupancy estimates and / or visibility estimates; - Map estimation; - An estimate of the drivable area relative to the vehicle; - Scene segmentation in non-sensor-specific representations; - Object tracking; - Object velocity estimation; - Depth estimation and / or distance estimation; - Detection of instances outside the domain; and - Time prediction.
13. A task header (300) for performing a perception task, the task header (300) comprising: - An input layer (312) is configured to receive a plurality of latent representations, wherein each latent representation is associated with a first training dataset included in a first training database, wherein each first training dataset includes optical sensor data acquired by at least one optical sensor, and wherein each first training dataset also includes a perception task label; - One or more hidden layers (314) are configured to transform each of the multiple received latent representations into a perceptual task result; - Output layer (316), which is configured to output the perceptual task results of each of the multiple received latent representations; The task head (300) is pre-trained by optimizing a task-specific loss function (319), wherein optimizing the task-specific loss function (319) includes comparing the output perceptual task result with the perceptual task label of each of a plurality of received latent representations, and / or wherein the task head (300) is pre-trained by the method (100) according to any one of claims 1 to 9.
14. A neural network system (400) for performing a perception task, the neural network system (400) comprising: - In particular, the pre-trained second encoder (401) includes: An input layer is configured to receive a dataset comprising optical sensor data acquired by at least one optical sensor. One or more hidden layers are configured to encode received optical sensor data into a latent representation; The output layer, which is configured to output the latent representation; and - The task header (300) according to the preceding claim, wherein: The input layer (312) is configured to receive the latent representation output by the second encoder; One or more hidden layers (314) are configured to transform the received latent representation into a perceptual task result; and The output layer (316) is configured to output the results of the perception task; The second encoder (401) may be pre-trained according to the method of claim 11.