Computer-implemented method for sensor data processing in a vehicle

CN122528010APending Publication Date: 2026-08-07ZENSEACT AB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZENSEACT AB
Filing Date
2026-02-06
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

在资源有限的环境中运行大型模型可能导致时延问题、能耗增加和不切实际的硬件成本

Benefits of technology

[0006]The techniques disclosed herein aim to mitigate, alleviate, or eliminate one or more of the aforementioned defects and disadvantages in the prior art to address various problems related to sensor data processing. The disclosed techniques are inspired by the so-called Hybrid Expert (MoE) technique, which has been used in the field of large-scale large language models (LLM). A method is proposed that can be used to alleviate the tension between the limited computing budget of in-vehicle systems and the need for a large number of model parameters, enabling the model to understand all the scenarios it may encounter. Various aspects and implementations of the disclosed techniques are defined below and in the appended independent and dependent claims.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528010A_ABST
    Figure CN122528010A_ABST
Patent Text Reader

Abstract

The technology disclosed herein relates to a computer-implemented method for sensor data processing in a vehicle. The method comprises: obtaining a pre-processed representation of sensor data related to a surrounding of the vehicle at a current time instant, wherein the sensor data comprises sensor data associated with sensors of the vehicle; obtaining a plurality of fusion experts that differ from each other in one or more aspects; processing the pre-processed representation of sensor data and / or a fused world view representation of a previous time instant as input to a fusion gating module to generate a selected subset of fusion experts from the plurality of fusion experts based on an output of the fusion gating module; processing, by each fusion expert of the selected subset of fusion experts, the pre-processed representation of sensor data together with the fused world view representation of the previous time instant to generate a respective output; and providing, based on the outputs of the subset of fusion experts, an updated fused world view representation for the current time instant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to the field of autonomous driving systems. In particular, it relates to methods and apparatus for processing sensor data in a vehicle equipped with an autonomous driving system. Background Technology

[0002] Autonomous driving systems rely heavily on sensor data to accurately perceive and interpret their surroundings. Various sensors, such as LiDAR, radar, and cameras, generate vast amounts of data that must be processed in real time to ensure reliable and high-performance operation. A challenge in autonomous (or semi-autonomous) vehicle perception is balancing the need for highly complex models capable of handling diverse driving scenarios with the limited computing resources available in the vehicle.

[0003] An established trend in deep learning is that larger models trained on ample data outperform smaller models across a wide range of applications. This principle also applies to perception models in autonomous driving systems, where deep neural networks analyze sensor data for tasks such as object detection, lane recognition, and obstacle avoidance. The real world presents a near-infinite number of scenarios, requiring these models to generalize across diverse weather conditions, lighting variations, road geometry, and object types. For example, a LiDAR-based object detection model trained only in clear weather may struggle in rainy conditions due to variations in point cloud density and reflection patterns. Similarly, recognizing different object classes, such as moose versus trucks, requires diverse model parameters to account for different shapes, textures, and motion patterns.

[0004] While large models offer advantages, onboard computing power is inherently constrained by hardware limitations, power consumption considerations, and real-time processing requirements. Unlike cloud-based systems that can leverage massive computing resources, onboard inference must be efficient enough to operate within the available processing budget. Running large models in resource-constrained environments can lead to latency issues, increased energy consumption, and impractical hardware costs. As a result, autonomous vehicle developers face a fundamental trade-off: larger models improve performance but exceed available computing resources, while smaller models are computationally feasible but struggle with robustness and adaptability across diverse driving conditions.

[0005] Therefore, there is a need for improved techniques for processing sensor data in autonomous vehicles, which can make more efficient use of computing resources without compromising model accuracy and robustness, while balancing model size and computational efficiency. Summary of the Invention

[0006] The techniques disclosed herein aim to mitigate, alleviate, or eliminate one or more of the aforementioned defects and disadvantages in the prior art to address various problems related to sensor data processing. The disclosed techniques are inspired by the so-called Hybrid Expert (MoE) technique, which has been used in the field of large-scale large language models (LLM). A method is proposed that can be used to alleviate the tension between the limited computing budget of in-vehicle systems and the need for a large number of model parameters, enabling the model to understand all the scenarios it may encounter. Various aspects and implementations of the disclosed techniques are defined below and in the appended independent and dependent claims.

[0007] According to a first aspect, a computer-implemented method is provided for processing sensor data in a vehicle equipped with an autonomous driving system. The method includes: obtaining one or more preprocessed representations of sensor data relating to the vehicle's surrounding environment at a current moment. The sensor data includes sensor data associated with one or more sensors of the vehicle. The method further includes: obtaining a plurality of fusion experts. The fusion experts of the plurality of fusion experts differ from each other in one or more aspects. The method further includes: processing one or more preprocessed representations of the sensor data and / or a fused world view representation about a previous moment as input to a fusion gating module to generate a subset of selected fusion experts from the plurality of fusion experts based on the output of the fusion gating module. The method further includes: processing the preprocessed representations of the sensor data together with the fused world view representation about a previous moment through each fusion expert in the subset of selected fusion experts to generate a corresponding output. The method further includes: providing an updated fused world view representation about the current moment based on the output of the subset of fusion experts. Utilizing this aspect of the disclosed technology, there are similar advantages and preferred features as other aspects.

[0008] According to a second aspect, a computer program product is provided, comprising instructions that, when executed by a computing device, cause the computing device to perform a method according to any embodiment of the first aspect. According to an alternative embodiment of the second aspect, a (non-transient) computer-readable storage medium is provided. The non-transient computer-readable storage medium stores one or more programs configured to be executed by one or more processors of a processing system, the programs including instructions for performing a method according to any embodiment of the first aspect. This aspect of the technology disclosed has similar advantages and preferred features as other aspects.

[0009] As used herein, the term "non-transient" is intended to describe computer-readable storage media (or "memory") that exclude the propagation of electromagnetic signals, but is not intended to otherwise limit the types of physical computer-readable storage devices encompassed by the phrase computer-readable media or memory. For example, the terms "non-transient computer-readable media" or "tangible memory" are intended to encompass types of storage devices that include, for example, random access memory (RAM) that do not necessarily permanently store information. Program instructions and data stored in non-transient form on tangible computer-accessible storage media can be further transmitted via transmission media or signals such as electrical signals, electromagnetic signals, or digital signals, which can be transmitted via communication media such as networks and / or wireless links. Therefore, as used herein, the term "non-transient" is a limitation on the medium itself (i.e., tangible, not a signal), rather than a limitation on the persistence of data storage (e.g., RAM and ROM).

[0010] According to a third aspect, a computing device is provided for processing sensor data in a vehicle equipped with an autonomous driving system. The computing device includes control circuitry. The control circuitry is configured to obtain one or more preprocessed representations of sensor data relating to the vehicle's surrounding environment at a current moment. The sensor data includes sensor data associated with one or more sensors of the vehicle. The control circuitry is further configured to obtain a plurality of fusion experts. The fusion experts of the plurality of fusion experts differ from each other in one or more aspects. The control circuitry is further configured to process one or more preprocessed representations of the sensor data and / or a fused world view representation about a previous moment as input to a fusion gating module to generate a subset of selected fusion experts from the plurality of fusion experts based on the output of the fusion gating module. The control circuitry is further configured to process the preprocessed representations of the sensor data together with the fused world view representation about a previous moment through each fusion expert in the subset of selected fusion experts to generate a corresponding output. The control circuitry is further configured to provide an updated fused world view representation about the current moment based on the output of the subset of fusion experts. Utilizing this aspect of the technology disclosed herein offers advantages and preferred features similar to those of other aspects.

[0011] According to a fourth aspect, a vehicle is provided, including a computing device according to any embodiment of the third aspect. This aspect utilizing the disclosed technology possesses similar advantages and preferred features as the other aspects.

[0012] The disclosed aspects and preferred embodiments may be appropriately combined with each other in any manner that is obvious to those skilled in the art, such that one or more features or embodiments associated with one aspect may also be considered to be associated with embodiments of another aspect or another aspect.

[0013] Some implementations offer advantages such as improved overall performance of sensor data processing due to generalization across multiple experts. Furthermore, the efficiency of sensor data processing can be achieved by routing to a network of relevant experts. Additionally, computational costs are reduced by activating only the relevant experts among the multiple experts. Through collaboration among different experts, the overall performance of any downstream task can also be improved.

[0014] One advantage of some implementations is that the provided method of extending modern perception and planning models to an unknown number of parameters by introducing expert gating into sensor coding, fusion models and / or prediction models can overcome the bottleneck of limited computational budgets in vehicles.

[0015] One advantage of some implementations is that they can be trained end-to-end, allowing the learning process to make its own decisions about how to allocate responsibilities among experts. This makes the method more scalable and robust, as it allows for the addition of any number of experts and retraining of the model to extend it. More specifically, experts can be easily added or updated with new expertise over time to support the expansion of the vehicle's (or its ADS's) operational design domain and / or to meet new requirements in the operation of the ADS.

[0016] One advantage of some implementation methods is that the training process can be simplified, for example, in terms of training cost. More specifically, training smaller expert networks is generally associated with lower training costs compared to training a single, large model that requires incorporating a large amount of knowledge into a single model.

[0017] Further embodiments are defined in the dependent claims. It should be emphasized that when the term "comprising" and variations thereof are used in this specification, they are used to specify the presence of a described feature, integer, step, or component. They do not exclude the presence or addition of one or more other features, integers, steps, components, or groups thereof.

[0018] These and other features and advantages of the disclosed technology will be further illustrated below with reference to the embodiments described herein. Attached Figure Description

[0019] The above aspects, features, and advantages of the disclosed technology will be more fully understood when taken in conjunction with the accompanying drawings and by referring to the following illustrative and non-limiting detailed description of exemplary embodiments of the present disclosure, wherein:

[0020] Figures 1A to 1D It is a schematic flowchart representing a method according to some implementation methods;

[0021] Figure 2 This is a schematic illustration of a computing device according to some embodiments;

[0022] Figure 3 This is an illustrative illustration of a vehicle based on some implementation methods;

[0023] Figures 4A to 4C The sensor processing model according to some implementation methods is illustrated through the first set of examples;

[0024] Figure 5 A sensor processing model according to some implementation methods is illustrated by way of a second example;

[0025] Figure 6 A sensor processing model according to some implementation methods is illustrated by way of a third example;

[0026] Figure 7 It is a schematic flowchart representing a method according to some implementation methods;

[0027] Figure 8 This is a schematic illustration of a computing device according to some implementation methods. Detailed Implementation

[0028] This disclosure will now be described in detail with reference to the accompanying drawings, in which some exemplary embodiments of the disclosed technology are illustrated. However, the disclosed technology may be embodied in other forms and should not be construed as limited to the exemplary embodiments disclosed. The exemplary embodiments of the disclosure are provided to fully convey the scope of the disclosed technology to those skilled in the art. Those skilled in the art will understand that the steps, services, and functions explained herein can be implemented using separate hardware circuitry, using software that works in conjunction with a programmable microprocessor or general-purpose computer, using one or more application-specific integrated circuits (ASICs), using one or more field-programmable gate arrays (FPGAs), and / or using one or more digital signal processors (DSPs).

[0029] It will also be appreciated that when this disclosure is described in the form of a method, it can also be embodied in a device including one or more processors and one or more memories coupled to the one or more processors, wherein computer code is loaded to implement the method. For example, in some embodiments, the one or more memories may store one or more computer programs that, when executed by the one or more processors, cause the device to perform the steps, services, and functions disclosed herein.

[0030] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. It should be noted that the articles “a,” “an,” “the,” and “described,” as used in the specification and appended claims, are intended to mean the presence of one or more elements unless the context clearly indicates otherwise. Thus, for example, in some contexts, a reference to “unit” or “the unit” may refer to more than one unit, etc. Furthermore, the word “comprising” and its variations do not exclude other elements or steps. It should be emphasized that when the term “comprising” and its variations are used in this specification, they are used to specify the presence of a described feature, integer, step, or component. It does not exclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. The term “and / or” should be interpreted as also meaning “both” and each as an alternative.

[0031] It should also be understood that although the terms first, second, etc., may be used herein to describe various elements or features, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of the embodiments, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element. Both first elements and elements are elements, but unless otherwise stated, they are not the same element.

[0032] As used herein, the phrase "one or more of a set of elements" (such as "one or more of A, B, and C" or "at least one of A, B, and C") should be interpreted as conjunction or disjunction. In other words, it can refer to all elements in a set of elements, one element, or a combination of two or more elements. For example, the phrase "one or more of A, B, and C" can be interpreted as A or B or C, A and B and C, A and B, B and C, or A and C.

[0033] The term "obtain" is to be interpreted broadly herein and encompasses the direct and / or indirect receiving, retrieving, collecting, and acquiring of information between two entities configured to communicate with each other or further with other external entities. However, in some embodiments, the term "obtain" is to be interpreted as determining, deriving, forming, calculating, etc.

[0034] Overview

[0035] As explained earlier, the disclosed technology relates to sensor data processing in vehicles equipped with an Automated Driving System (ADS). This involves not only the processing itself but also downstream tasks such as perception and planning systems that rely on sensor data. In this context, an emerging trend in sensor data processing is the adoption of temporal fusion, potentially spanning several different sensors. As a high-level example, a processing pipeline can be described as follows. First, sensor data (from one or more sensors) is collected and encoded by one or more encoding networks from each sensor. Different encoding networks may be used, for example, for different sensor types (e.g., one network for LiDAR data, another for image data, etc.) or for different instances of the same sensor type (i.e., one network for camera 1, another for camera 2, etc.). However, there can also be a single encoding network capable of encoding sensor data of different sensor types. Second, the encoded sensor data is combined with a fusion state (or representation) from a previous time step via a sensor and temporal fusion network. The fusion state at a specific timestamp can be viewed as an abstract representation of relevant information about the world that the vehicle (or more precisely, its perception system) understands. The fusion network then takes the previous fusion state and combines it with the newly encoded sensor measurements to form an updated fusion state. Third, a task-specific prediction network is run on the latest fusion state. An example of such a prediction network could be a transformer decoder, which takes the fusion state as input to output, for example, a 3D bounding box for an object.

[0036] The disclosed technology is at least in part based on the concept of Hybrid Experts (MoE) in such sensor data processing. The idea behind MoE stems from the fact that LLMs must contain an extremely large number of parameters to reason about any subject. To avoid the rapidly increasing inference and training costs that come with the amount of knowledge in an LLM, it employs a gating mechanism within the transformer to route each lexical position to the appropriate expert. Individual experts have relatively low dimensionality and a small number of parameters, but because many experts are used, the overall model still has a very large number of parameters. Because only one expert runs at a time, increasing the number of model parameters by adding more experts does not increase the model's inference cost. The disclosed technology further takes this principle and adapts it to autonomous driving scenarios. As the inventors recognize, several different options exist for implementing this.

[0037] According to publicly available technology, one option is to use MoE as part of the encoding network. In this case, a relatively small and / or lightweight pre-encoder can be run on sensor data from one or more sensors. A set of encoder experts can then be provided. Each encoder expert can be similar to currently known sensor encoders. The output of the pre-encoder can then be fed into a gating module, which will select one or more encoder experts to run. Then, only these selected experts will run for this specific sensor data. The gating mechanism can be extremely simple and lightweight, for example, performing pixel pooling and then linear projection across the channel dimension followed by softmax. This gating would be very inexpensive to run compared to a full encoder. The following example will combine... Figure 5 and Figure 7 Let me describe this setting further.

[0038] Based on a similar principle, another option for the disclosed technique is to employ MoE as part of the fusion network. The fusion network typically constitutes a relatively large portion of the runtime of the sensor data processing pipeline, and its task is to generate the final representation of the world from which predictions are made. Therefore, the disclosed technique can be found to offer particular advantages in this application. In this case, the encoded sensor data and / or the previous fusion state can be fed into a gating mechanism. The gating mechanism can also be preceded by some simple sensor fusion, or depend solely on the previous fusion state. The gating mechanism then selects one or more fusion experts, each capable of handling sensor and temporal fusion. The selected fusion experts are run and the updated fusion state is output. The following will refer to Figure 1 and... Figures 4A to 4C Let me describe this setting further.

[0039] Another option is to use MoE as part of the prediction network, based on the same principles as above. This will be discussed below. Figure 6 Let me describe this setting further.

[0040] The settings described in the different options above can also be combined with each other. In other words, expert gating can be employed in encoding, fusion, and / or prediction. This can further enhance the advantages of the disclosed techniques.

[0041] The advantage of the described solution is that the encoder / fusion / prediction experts can be much smaller and faster during inference than a large encoder / fusion / prediction network that would have to handle all types of scenarios. The proposed processing pipeline (which can be implemented as part of the perception module of an ADS) can automatically learn the optimal distribution of experts and task partitioning during training, and different experts can learn to handle different weather conditions, geographic locations, times of day, or the number of traffic participants. The disclosed method utilizes end-to-end learning so that the learning process can make its own decisions about how to allocate responsibility among the experts. This makes the method more scalable and robust, as it allows adding any number of experts and retraining the pipeline network to expand it.

[0042] Furthermore, when the number of scenarios the perception model needs to handle needs to be expanded, or when it is desired to increase the number of parameters in the perception model, more experts can be simply added without affecting the in-vehicle assembly line's uptime. This contrasts with existing technologies, which become increasingly expensive to operate as more parameters are added.

[0043] In the following text, we will combine Figures 1A to 8 Further description of the disclosed technology.

[0044] definition

[0045] In this context, an Automated Driving System (ADS) refers to a complex combination of hardware and software components designed to control and operate a vehicle without direct human intervention. ADS technology aims to automate various aspects of driving, such as steering, acceleration, deceleration, and monitoring of the surrounding environment. The primary goal of ADS is to enhance safety, efficiency, and convenience in transportation. The range of ADS can vary from basic driver assistance systems to highly advanced autonomous driving systems, depending on their level of automation as classified by standards such as SAE J3016. These systems utilize various sensors, cameras, radar, lidar, and powerful computer algorithms to perceive the environment and make driving decisions. The specific capabilities and characteristics / functions of ADS can vary greatly, from systems providing limited assistance to systems capable of independently handling complex driving tasks under specific conditions.

[0046] Advanced Driver Assistance Systems (ADAS) are technologies that assist drivers during driving, although they do not necessarily provide complete autonomy. ADAS features are often used as building blocks for ADS. Examples include adaptive cruise control, lane keeping assist, automatic emergency braking, and parking assist. They enhance safety and convenience but typically require a certain level of human supervision and intervention. Autonomous Driving (AD), on the other hand, is a technology designed to control and navigate a vehicle without human supervision. Accordingly, the difference between ADAS and AD can be said to lie in the level of autonomy and control. ADAS systems are designed to assist and support the driver, while AD aims for complete control of the vehicle without constant human supervision. Accordingly, AD aims for a higher level of autonomy (such as Level 4 and Level 5 according to SAE International Standards), where the vehicle can operate independently in most or all driving scenarios without human intervention. As mentioned above, the term "ADS" used in this article is used as a collective term encompassing both ADAS and AD. In this context, ADS functions or ADS features can be understood as the specific functions or features of the entire ADS stack, such as highway navigation features, traffic congestion navigation features, and route planning features.

[0047] In the current context, "sensor" or "sensor device" refers to a specialized component or system designed to capture and collect information from the vehicle's surroundings. It can further refer to components used to collect information about the vehicle itself. These sensors play a crucial role in enabling Autonomous Vehicles (ADSs) to perceive and understand their environment, make informed decisions, and navigate safely. Sensors are typically integrated into the hardware and software systems of autonomous vehicles to provide real-time data for various tasks such as obstacle detection, localization, road model estimation, and object recognition. Commonly used sensor types in autonomous driving include LiDAR (Light Detection and Ranging), radar, cameras, ultrasonic sensors, inertial measurement units (IMUs), GPS, and wheel speed sensors. LiDAR sensors use laser beams to measure distances and create high-resolution 3D maps of the vehicle's surroundings based on point clouds. Radar sensors use radio waves to determine the distance and relative speed of objects around the vehicle. Camera sensors capture visual data, allowing the vehicle's computer system to identify traffic signs, lane markings, pedestrians, and other vehicles. Ultrasonic sensors use sound waves to measure proximity to objects. Various machine learning algorithms, such as artificial neural networks, can be used to process the output from sensors to understand the environment.

[0048] As used in this article, “network” (also known as machine learning algorithm, machine learning model, (artificial) neural network, deep learning network, etc.) refers to any computational system or algorithm that is trained on data to make predictions, decisions, or otherwise generate outputs, such as by learning patterns and relationships from training data and applying that knowledge to new input data.

[0049] The deployment of such networks typically involves a training phase, during which the network learns from labeled or unlabeled training data to achieve accurate predictions in the subsequent inference phase. Training data (and input data during inference) can be, for example, images or image sequences, LiDAR data (i.e., point clouds), radar data, or any other form of data. Furthermore, training / input data can include combinations or fusions of one or more different data types. Additionally, or in combination, it can include combinations or fusions of two or more instances of the same data type, such as two or more images from different cameras.

[0050] In some implementations, the network or machine learning model may be implemented using publicly available, appropriate software development machine learning code elements in any manner deemed appropriate by a person skilled in the art, such as code elements available in PyTorch, TensorFlow, and Keras, or in any other appropriate software development platform.

[0051] Implementation

[0052] Figures 1A to 1D A schematic flowchart illustrates a method 100 implemented by a computer for processing sensor data in a vehicle equipped with an automated driving system (ADS). Figure 1B Illustration Figure 1A Several optional sub-steps of step S104. Similarly, Figure 1C Illustration Figure 1A Several optional sub-steps of the step marked S110. Similarly, Figure 1D Illustration Figure 1A Several optional sub-steps of the step marked S116.

[0053] In some implementations, method 100 can be viewed as a method 100 for sensor data fusion. Method 100 can be executed in a vehicle (i.e., using the vehicle's computing resources). More specifically, method 100 can be executed as part of an Adaptive Data System (ADS). The following is in conjunction with... Figure 3 This vehicle 300 is described by way of example. However, it should be noted that the principles described herein can also be implemented in offline environments, such as by a server. In this case, the techniques described herein can be used as part of a pseudo-annotation system, i.e., for automatically (or semi-automatically) generating annotated data for training data, which can then be used to develop other machine learning models for use in the context of autonomous driving systems.

[0054] The different steps of method 100 are described below in more detail. Even though illustrated in a specific order, the steps of method 100 can be performed in any suitable order and multiple times. Therefore, although the accompanying drawings may show a specific order of the method steps, the order of the steps may differ from what is depicted. Furthermore, two or more steps can be performed simultaneously or partially simultaneously. For example, steps labeled S106 and S108 can be performed independently of each other. This variation will depend on the chosen software and hardware system and the designer's choice. All such variations are within the scope of the disclosed technology. Similarly, software implementation can be accomplished using standard programming techniques based on rule-based logic and other logics to complete the various steps. Further variations of method 100 will become apparent from this disclosure. The embodiments mentioned and described herein are given by way of example only and should not be limited to the technology of this disclosure. Other solutions, uses, purposes, and functions within the scope of the disclosed technology claimed in the patent claims described below will be apparent to those skilled in the art.

[0055] It should be further recognized that, Figures 1A to 1D Method 100 includes steps illustrated with solid lines and steps illustrated with dashed lines. The steps illustrated with solid lines are those included in the most extensive exemplary embodiment of method 100. The steps included with dashed lines are examples of several optional steps that may form part of several alternative embodiments. It should be understood that the optional steps do not need to be performed sequentially. Furthermore, it should be understood that not all optional steps need to be performed. Optional steps can be performed in any order and in any combination.

[0056] Method 100 includes obtaining one or more preprocessed representations of the sensor data in S106. The sensor data includes sensor data associated with one or more sensors of the vehicle. In other words, the sensor data may be captured by one or more sensors of the vehicle and includes data from each of the sensors. Therefore, the sensor data may include one or more sensor data types. As an example, the sensor data may include image data captured by a camera of the vehicle, LiDAR data captured by a LiDAR sensor, radar data captured by a radar, etc. More specifically, the sensor data may include one or more of image data, LiDAR data, radar data, and ultrasonic data.

[0057] Sensor data pertains to the vehicle's surroundings at the current moment. In other words, sensor data can depict or represent the surrounding environment in any other way. The vehicle's surroundings can be understood as the general area around the vehicle, where objects (such as traffic signs, other vehicles, landmarks, obstacles, etc.) can be detected and identified by the vehicle's sensors (radar, LiDAR, cameras, etc.), i.e., within the vehicle's sensor range.

[0058] Each of one or more preprocessed representations of sensor data can be associated with sensor data from a corresponding sensor among one or more sensors. In other words, each preprocessed representation can correspond to sensor data of a sensor data type. One preprocessed representation can be associated with image data (i.e., generated based on image data), while another preprocessed representation can be associated with LiDAR data, and so on. However, in some implementations, sensor data from more than one sensor can be associated with a single preprocessed representation. In other words, the preprocessed representation can be formed based on both image data and LiDAR data.

[0059] In the current context, the term "preprocessed representation" can be understood as a transformed or processed version of sensor data. A preprocessed representation can be obtained by applying certain calculations or modifications to the sensor data. Such calculations can be used, for example, to enhance usability, reduce noise, integrate multiple data sources, etc. Depending on the specific processing technique employed, this representation can take various forms.

[0060] In some implementations, the preprocessed representation of sensor data is an encoded representation of the sensor data. In other words, preprocessing sensor data may include encoding the sensor data. Examples of this are given below. Figure 4A The following is illustrated and further explained. The term "encoded representation" can be understood as a compressed form of sensor data. Sensor data can be transformed into different representations, such as a vector representation in latent space or any other numerical representation. In some cases, sensor data is transformed into a low-dimensional representation. The encoded representation can be generated by processing the sensor data via an encoding network. The encoding network can be any suitable network commonly used for encoding sensor data. As an example, the encoding network can include a convolutional neural network (CNN). In another example, the encoding network can include a vision transformer for processing image data. In yet another example, the encoding network can include PointNet or PointPillars for processing LiDAR or radar data.

[0061] In some implementations, the preprocessed representation of the sensor data is a pre-fused representation of the sensor data. In other words, preprocessing the sensor data may include performing some form of sensor fusion. The pre-fused representation can be generated by processing the sensor data via a pre-fusion network. Examples of this are given below. Figure 4B Show and further explain.

[0062] In some implementations, the preprocessed representation of the sensor data is the internal state of the fusion network. More specifically, the concepts of the disclosed techniques can be used as part of the fusion network. For example, the fusion network can be a transformer. The transformer includes several transformer layers. Each transformer layer further includes an attention layer and a feedforward neural network (FFN), where the concept of MoE can be employed. The attention layer can perform cross-attention and self-attention. Examples of this are combined below. Figure 4C Show and further explain.

[0063] Method 100 further includes: obtaining S108 multiple fusion experts. The multiple fusion experts differ from each other in one or more aspects. The multiple fusion experts can therefore be (or form) a mixture of experts. Each fusion expert can be a separate neural network. In other words, a fusion expert can be viewed as a specialized neural network for the task of fusing sensor data. One aspect in which fusion experts can differ from each other is that they can have different sets of weights. Another aspect is that the fusion experts may have been trained on different datasets. Yet another aspect is that the fusion experts have different architectures. In a further example, the fusion experts can differ in that they focus on different regions in the input, or focus on recognizing different patterns in the input. It should be noted that these aspects should be considered as non-limiting examples, as fusion experts can also differ in other ways. In general, the fusion experts differ from each other in that, as part of the task of fusing sensor data, they are specifically designed for different things.

[0064] Method 100 further includes: processing one or more preprocessed representations of the S110 sensor data and / or a fused world view representation about a previous time as input to a fusion gating module to generate a subset of fusion experts from a plurality of fusion experts based on the output of the fusion gating module.

[0065] A world-view representation can be understood as an internal model (or state) of the vehicle (or more precisely, the Adaptive Digital Sensor) of its surroundings, i.e., a representation of how the vehicle perceives its environment. The world-view representation can therefore be considered a representation of the surrounding environment. This can be used to allow the vehicle to interpret and navigate its surroundings, and serves as a basis for decision-making, path planning, and control. A world-view representation is "fused" in the sense that it is constructed by integrating data from multiple sensors and multiple moments in time. A fused world-view representation at a given moment can therefore be a representation of the surrounding environment based on sensor data relating to the vehicle's environment at that moment, and sensor data relating to the vehicle's environment at at least one previous moment. Different moments associated with the captured sensor data, as referred to herein, can correspond to the points in time when the sensor data was captured. Therefore, two successive moments can constitute two successive capture frames.

[0066] As described above, the fused world view representation is a representation relative to a previous moment. The terms "current" and "previous," used in phrases like "current moment" and "previous moment," are intended only to indicate the relative relationship between moments. Therefore, the term "current" does not necessarily refer to the actual time of performing the method (or its steps). "Previous" also does not necessarily refer to the immediately preceding moment. Rather, "previous" simply indicates a moment earlier than the "current" moment. In this specific case, the current moment refers to the moment when the most recently acquired sensor data depicts the surrounding environment (i.e., the time when the sensor data was captured). This sensor data has not yet been fused into the fused world view representation. The previous moment refers to the moment up to which the fused world view representation has been generated. In other words, the fused world view representation is a representation of the world up to that moment. As will be explained further below, when the sensor data at the current moment has been integrated into the world view representation, an "updated fused world view representation for the current moment" can be provided.

[0067] The fusion gating module is configured to generate an output indicative of a subset of the fusion expert's selection, based on one or more preprocessed representations of the received sensor data, a fused world view representation about a previous time step, or both as input. The output may include any information from which the subset of the fusion expert's selection can be obtained. In some embodiments, the output includes a subset of the fusion expert's selection. In other words, the step of processing one or more preprocessed representations of the S110 sensor data and / or a fused world view representation about a previous time step as input to the fusion gating module can be understood as using the fusion gating module to determine a subset of the fusion expert's selection from a plurality of fusion experts, wherein the fusion gating module is configured to process the input and output the corresponding subset of the fusion expert's selection.

[0068] The selection of a subset of fusion experts can be viewed as one or more (but not all) fusion experts selected from a pool of fusion experts to process the preprocessed representation of sensor data. Consistent with the MoE concept, the gating module thus selects the fusion expert best suited for processing the input by applying some gating function or mechanism. The gating module can therefore be viewed as routing the input to the selected experts. The gating function can be configured to map the input (i.e., the preprocessed representation of sensor data and / or the previous world view representation) to a ranking score for the corresponding fusion expert. From another perspective, the gating module can be configured to dynamically select which fusion experts should contribute to the final output for a given input. The gating module can thus determine the weight of each fusion expert's contribution, allowing the fusion process to adaptively allocate computational resources based on the complexity and characteristics of the input, in the best possible way and to achieve the best possible result performance.

[0069] More specifically, the fusion gating module can be configured to generate a subset of fusion experts by first determining the ranking score of each of the multiple fusion experts in S110a based on the input. Then, a subset of the top k fusion experts in S110b is selected based on the determined ranking scores. Here, k is a positive integer greater than 0 and less than the number of fusion experts. Finally, the subset of the top k fusion experts selected in S110c is provided as output. The ranking score can be considered as a metric indicating the "relevance" or "performance" of these (i.e., the fusion experts) in processing the input. A higher ranking score for a fusion expert can therefore indicate that the fusion expert is more suitable to be used than a fusion expert with a lower ranking score. The ranking score can thus be considered as a weight for different fusion experts.

[0070] The gating module can further consider additional data when generating the output. More specifically, the step of selecting a subset of the top k fusion experts in S110b can be further based on metadata associated with one or more environmental conditions related to the surrounding environment. Environmental conditions may, for example, involve the ambient lighting level, weather, time of day, etc. This data may be obtained, for example, from the vehicle's onboard system or from an external data source. By incorporating metadata associated with one or more environmental conditions, the weights of the fusion experts can be improved because this information can provide additional signals to the decision. For example, if the gating module knows that its lighting level is low (e.g., due to nighttime conditions), it can assign more weight to fusion experts who perform well in such scenarios.

[0071] Furthermore, the gating module can further consider the expert selections made in previous time steps. More specifically, the selection of a subset of the top k fusion experts in S110b can be further based on a subset of fusion experts used to provide the world view representation of fusion in previous time steps. Thus, fusion experts can be alternated relative to subsequent time steps. As an example, the gating module can assign lower weights to fusion experts used in previous iterations. Information about which fusion experts were used and when can be explicitly fed to the gating module, or implicitly fed in by incorporating it into the world view representation of fusion in previous time steps fed to the gating module.

[0072] Furthermore, some of the multiple fusion experts can be selected for each iteration of the method (i.e., each time new sensor data is fused into the world view representation). The gating module can then select these fusion experts as part of the generated set of selected fusion experts each time the method is executed.

[0073] In some implementations, the fusion gating module is configured to apply a learned gating function to the input to generate an output. The gating module may, for example, employ a neural network trained to output a subset of the fusion experts' choices based on the input. The gating function may include a Softmax activation function or a Sigmoid activation function, each configured to determine a ranking score associated with each fusion expert based on the input. The learned gating function may be jointly trained with the rest of the network / model (i.e., the encoding network, the fusion network including the fusion experts, and the prediction network). This can be done in an end-to-end manner.

[0074] In some implementations, the fusion gating module is configured to apply a deterministic gating function to the input to generate the output. Unlike learned gating functions, the deterministic gating function is not learned but rather applies a rule-based approach to determine the ranking score. The deterministic gating function can exist as part of the training of the rest of the network (i.e., the fusion expert). In this way, the fusion expert can be optimized based on the deterministic gating function. The state in which the deterministic gating function operates can also be optimized, as it can be trained end-to-end with the rest of the network / model.

[0075] Method 100 further includes: S112 processing the preprocessed representation of the sensor data together with the fused world view representation from the previous time step by each fusion expert in a subset of selected fusion experts to generate a corresponding output. In other words, the preprocessed representation of the sensor data and the fused world view representation from the previous time step can be fed to each of the subset of selected fusion experts. The output generated by each fusion expert can therefore be a fused representation of the preprocessed representation of the sensor data and the fused world view representation from the previous time step.

[0076] In some implementations, the preprocessed representation of the sensor data is an attention representation of the sensor data, which is generated by applying cross-attention and / or self-attention together with a partially updated fused world view representation, one or more encoded representations of the sensor data, and a fused world view representation from a previous time step. The partially updated fused world view representation is obtained from the pre-fusion network. The pre-fusion network in this case may be a previous iteration performing steps labeled S110 and S112, or received from the processed data via a previous fusion step of a sequential fusion process. More specifically, the principles of the disclosed technique can be applied within a fusion network comprising several transformer layers (or fusion layers). Sensor data fusion can therefore be performed by processing data via several transformer layers. The several transformer layers can thus form a sequence of fusion layers. The steps of method 100 can be considered as steps performed in one of the transformer layers. The partially updated fused world view representation can then be the output of a previous transformer layer. For further details, refer to the following... Figure 4C .

[0077] Method 100 further includes: providing an updated fused world view representation of the current moment based on the output of a subset of fusion experts, S114.

[0078] In some implementations, the updated fused world view representation for S114 is provided by processing the output of each fusion expert in a subset of the selected fusion experts via a post-fusion network configured to generate an updated fused world view representation. The post-fusion network can therefore be a learned neural network trained to generate the final updated fused world view representation from the outputs of the fusion experts.

[0079] In some implementations, the updated fusion world view representation of S114 is provided by combining the outputs of each fusion expert in a subset of the selected fusion experts into the updated fusion world view representation. The outputs can be combined by weighted summation of S114b. The outputs can also be combined by averaging or weighted averaging of S114b. The weights for the weighted average can be ranking scores (or weights) determined as part of a gating function. In other words, the outputs can be combined based on how the corresponding fusion experts are weighted in the gating module. In another example, the outputs of each fusion expert in a subset of the selected fusion experts in S114b can be combined by summing the outputs or by applying max pooling to the outputs.

[0080] As explained above, some of the fusion experts can be selected for each iteration of the method. In other words, a set of experts can always be run, and the remaining experts can be selected dynamically according to the principles disclosed herein. This can be alternatively described as method 100 further comprising: processing a preprocessed representation of the sensor data together with a fused world view representation from a previous time step using a preselected set of fusion experts to generate a corresponding output. The updated fused world view representation for S114 can then be provided further based on the output of the preselected set of fusion experts.

[0081] Method 100 may further include generating a task-specific output S116 by processing the updated fused world view representation via a task-specific prediction network configured to generate task-specific output. The task-specific output can then be used for vehicle operation, for example, by sending it to a subsequent module of the ADS responsible for, for example, path planning or vehicle decision-making and control.

[0082] Task-specific prediction networks can be configured to perform tasks associated with the operation of an autonomous driving system. These tasks can be, for example, any of the following: object detection, object classification, object tracking, lane estimation, free space estimation, trajectory prediction, obstacle avoidance, path planning, scene classification, traffic sign classification, 3D scene flow, and occupancy prediction.

[0083] To date, publicly available techniques have been described as employing the concept of MoE (i.e., acting as a fusion network) in the process of sensor data fusion. However, it should be noted that the same techniques can be applied to the process of encoding sensor data (in...). Figure 5 (illustrated in the example), and the process of generating task-specific predictions (in...) Figure 6 (See the example illustrations below). This will be described further below.

[0084] In some embodiments, method 100 further includes: obtaining sensor data related to the vehicle's surrounding environment at the current moment, S102. Method 100 may then further include: encoding sensor data associated with each of one or more sensors, S104, by processing the sensor data via an encoding network to generate an encoded representation of the sensor data associated with each of the one or more sensors. The encoding network may include a pre-encoder and multiple encoder experts. Similar to multiple fusion experts, the multiple encoder experts may be different from each other and form a mixture of experts. Processing the sensor data associated with each of the one or more sensors via the encoding network may then include: for each of the one or more sensors: processing the sensor data associated with said sensor via a pre-encoder, S104a, to generate a pre-encoded representation of the sensor data. Then, using an encoder gating module, determining a subset of encoder experts from the multiple encoder experts, S104b, is a selection of encoder experts. This can be done by processing the pre-encoded representation of the sensor data via the encoder gating module. Then, processing the pre-encoded representation of the sensor data with each encoder expert in the subset of encoder experts, S104c, generates a corresponding output. Finally, based on the output of a subset of encoder experts, an encoded representation of the S104d sensor data is provided. This encoded representation of the sensor data (potentially after some further processing) can then be processed by the fusion gating module as a preprocessed representation of the sensor data, as explained above. The principles associated with the fusion gating module and multiple fusion experts also apply to the encoder gating module and multiple encoder experts. To avoid unnecessary repetition, refer to the above.

[0085] In some implementations, method 100 further includes generating a task-specific output S116 by processing an updated fused world view representation via a task-specific prediction network configured to generate task-specific outputs. The task-specific prediction network may include multiple prediction experts. Generating the task-specific output S116 may include determining a subset of prediction experts selected from the multiple prediction experts, S116a, using a prediction gating module. The prediction gating module may be configured to process at least the updated fused world view representation and generate outputs indicative of the subset of prediction experts selected. Then, the updated fused world view representation is processed by each prediction expert in the subset of prediction experts to generate a corresponding output. Finally, a task-specific output S116c is provided based on the outputs of the subset of prediction experts. The task-specific output S116c can be provided by processing the outputs of the subset of prediction experts via a post-prediction network configured to generate a final prediction output (such as a bounding box, object classification, planned trajectory, etc.). The principles associated with the fusion gating module and multiple fusion experts also apply to the prediction gating module and multiple prediction experts. For the avoidance of unnecessary repetition, refer to the above.

[0086] In some implementations, method 100 further includes: acquiring map data associated with the vehicle's surrounding environment. More specifically, the map data may be acquired based on the vehicle's location at the time the sensor data was acquired. The step labeled S110 may further include: processing the map data as input to the fusion gating module. Alternatively or in combination, the step labeled S112 may further include: processing the map data through each fusion expert in a selected subset of fusion experts. Thus, the map data can be incorporated to provide the fused world view representation updated in S114.

[0087] Executable instructions for performing these functions are optionally included in a non-transient computer-readable storage medium or other computer program product configured for execution by one or more processors.

[0088] Generally, computer-accessible media can include any tangible or non-transient storage medium or storage media such as electrical, magnetic, or optical media—for example, a hard disk or CD / DVD-ROM bus-connected to a computer system. As used herein, the terms “tangible” and “non-transient” are intended to describe computer-readable storage media (or “memory”) excluding those that propagate electromagnetic signals, but are not intended to otherwise limit the types of physical computer-readable storage devices encompassed by the phrases “computer-readable medium” or “memory.” For example, the terms “non-transient computer-readable medium” or “tangible memory” are intended to encompass types of storage devices that include, for example, random access memory (RAM) that do not necessarily permanently store information. Program instructions and data stored in non-transient form on tangible computer-accessible storage media can be further transmitted via transmission media or signals such as electrical, electromagnetic, or digital signals, which can be transmitted via communication media such as networks and / or wireless links.

[0089] Figure 2 This is a schematic illustration of a computing device 200 according to some embodiments of the disclosed technology. The computing device 200 can be configured to perform, as in combination with... Figures 1A to 1D Method 100 is described. Therefore, computing device 200 can be a computing device 200 for processing sensor data in a vehicle equipped with an autonomous driving system.

[0090] As described herein, computing device 200 refers to a computer system, or any device or general-purpose computing system configured to perform various functions. Even though computing device 200 is illustrated herein as a single device, it can also be a distributed computing system comprised of several different devices. Computing device 200 may be provided as part of vehicle 300.

[0091] The computing device 200 includes a control circuit 202. The control circuit 202 may physically comprise a single circuit device. Alternatively, the control circuit 202 may be distributed across several circuit devices.

[0092] like Figure 2 As shown in the example, computing device 200 may further include transceiver 206 and memory 208. Control circuitry 202 is communicatively connected to transceiver 206 and memory 208. Control circuitry 202 may include a data bus, and control circuitry 202 may communicate with transceiver 206 and / or memory 208 via the data bus.

[0093] Control circuitry 202 can be configured to perform overall control of the functions and operations of computing device 200. Control circuitry 202 may include processor 204, such as a central processing unit (CPU), microcontroller, or microprocessor. Processor 204 can be configured to execute program code stored in memory 208 to perform the functions and operations of computing device 200. Control circuitry 202 is configured to perform the above-described combination... Figures 1A to 1D The steps of method 100 are described. The steps can be implemented in one or more functions stored in memory 208.

[0094] Transceiver 206 is configured to enable computing device 200 to communicate with other entities such as other devices. Transceiver 206 can both transmit and receive data from computing device 200. Computing device 200 may, for example, be part of a vehicle. Transceiver 206 can then allow computing device 200 to communicate with other systems within the vehicle, or with external entities (such as other vehicles) or remote servers.

[0095] Memory 208 may be a non-transient computer-readable storage medium. Memory 208 may be one or more of the following: buffer, flash memory, hard disk drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or other suitable devices. In a typical arrangement, memory 208 may include non-volatile memory for long-term data storage and volatile memory used as system memory for computing device 200. Memory 208 may exchange data with circuit 202 via a data bus. Accompanying control lines and address buses may also exist between memory 208 and circuit 202.

[0096] The functions and operations of computing device 200 can be implemented in the form of executable logic routines (e.g., lines of code, software programs, etc.) stored on a non-transient computer-readable recording medium (e.g., memory 208) in computing device 200 and executed by circuit 202 (e.g., using processor 204). In other words, when circuit 202 is described as being configured to perform a specific function, processor 204 of circuit 202 can be configured to execute a portion of program code stored in memory 208, wherein the stored portion of program code corresponds to a specific function. Furthermore, the functions and operations of circuit 202 can be independent software applications or part of a software application that performs additional tasks related to circuit 202. The described functions and operations can be considered as methods by which the corresponding device is configured to perform, such as those described above. Figures 1A to 1DMethod 100 is discussed. Furthermore, although the described functions and operations can be implemented in software, the functions can also be implemented via dedicated hardware or firmware, or a combination of one or more of hardware, firmware, and software. The functions and operations of computing device 200 are described below.

[0097] Control circuit 202 is configured to acquire one or more preprocessed representations of sensor data relating to the vehicle's surrounding environment at the current moment. The sensor data includes sensor data associated with one or more sensors of the vehicle. This can be performed, for example, by executing a first acquisition function 210.

[0098] The control circuit 202 is further configured to acquire multiple fusion experts. The fusion experts differ from each other in one or more aspects. This can be performed, for example, by executing a second acquisition function 212.

[0099] The control circuit 202 is further configured to process one or more preprocessed representations of sensor data and / or a fused world view representation about a previous time step as input to a fusion gating module to generate a subset of fusion experts from a plurality of fusion experts based on the output of the fusion gating module. This can be performed, for example, by executing a first processing function 214.

[0100] The control circuit 202 is further configured to process the preprocessed representation of the sensor data together with the fused world view representation from the previous time step by each fusion expert in a subset of selected fusion experts to generate a corresponding output. This can be performed, for example, by executing a second processing function 216.

[0101] The control circuit 202 is further configured to provide an updated fused world view representation based on a subset of the fusion experts' output. This can be performed, for example, by executing the provision function 218.

[0102] The control circuit 202 can be further configured to generate a task-specific output by processing the updated fused world view representation via a task-specific prediction network configured to generate a task-specific output. This can be performed, for example, by executing prediction function 220.

[0103] It should be noted that the first acquisition function 210 and the second acquisition function 212 can be implemented as two separate functions or as a common acquisition function. Similarly, the first processing function 214 and the second processing function 216 can be implemented as two separate functions or as a common processing function.

[0104] Further attention should be paid to the following, as in the above combination Figures 1A to 1DThe principles, features, aspects, and advantages of method 100 described herein also apply to computing device 200 as described herein. To avoid unnecessary repetition, reference is made to the foregoing. Therefore, control circuitry can be configured to perform any of the steps described as part of method 100.

[0105] Figure 3 This is an illustrative illustration of a vehicle 300 according to some embodiments. The vehicle 300 may be equipped with an automated driving system (ADS) 310. As used herein, "vehicle" means any form of motorized transport. For example, the vehicle 300 may be any road vehicle such as a car (as illustrated herein), a motorcycle, a (freight) truck, a bus, a smart bicycle, etc. The vehicle 300 may be equipped with a computing device 200 as described above. Therefore, the vehicle 300 is capable of performing the disclosed technologies.

[0106] Vehicle 300 includes several components typically found in autonomous or semi-autonomous vehicles. It should be understood that vehicle 300 may have... Figure 3 Any combination of the various elements shown herein. Furthermore, vehicle 300 may include more than Figure 3 The elements shown herein are further elements. Although the various elements are shown herein as being located inside the vehicle 300, one or more of the elements may be located outside the vehicle 300. Furthermore, even though the various elements are depicted herein in certain arrangements, as will be readily understood by those skilled in the art, the various elements may also be implemented in different arrangements. It should be further noted that the various elements may be communicatively connected to each other in any suitable manner. Figure 3 Vehicle 300 should be considered only as an illustrative example, as the components of vehicle 300 can be implemented in several different ways.

[0107] Vehicle 300 includes a control system 302. The control system 302 is configured to perform overall control of the functions and operations of vehicle 300. The control system 302 includes control circuitry 304 and memory 306. Control circuitry 302 may physically comprise a single circuit device. Alternatively, control circuitry 302 may be distributed across several circuit devices. As an example, control system 302 may share its control circuitry 304 with other parts of the vehicle. Control circuitry 302 may include one or more processors such as a central processing unit (CPU), microcontroller, or microprocessor. One or more processors may be configured to execute program code stored in memory 306 to perform the functions and operations of vehicle 300. The processor may be or may include any number of hardware components for performing data or signal processing or for executing computer code stored in memory 306. In some embodiments, control circuitry 304 or some of its functions may be implemented on one or more so-called system-on-a-chip (SoC). As an example, ADS 310 may be implemented on an SoC. Memory 306 optionally includes high-speed random access memory such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices; and optionally includes non-volatile memory such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state memory devices. Memory 306 may include database components, object code components, script components, or any other type of information structure used to support the various activities of this specification.

[0108] In the illustrated example, memory 306 further stores map data 308. Map data 308 can be used, for example, by the ADS 310 of vehicle 300 to perform autonomous functions of vehicle 300. Map data 308 may include high-definition (HD) map data and / or standard-definition (SD) map data. Even though memory 308 is illustrated as a separate element from ADS 310, it is conceivable that memory 12 could be provided as an integrated element of ADS 310. In other words, according to some embodiments, any distributed or local memory device can be used to implement the inventive concept. Similarly, control circuitry 304 can be distributed, for example, such that one or more processors of control circuitry 304 are provided as integrated elements of ADS 310 or any other system of vehicle 300. In other words, according to exemplary embodiments, any distributed or local control circuitry device can be used to implement the technology of this disclosure.

[0109] Vehicle 300 further includes a sensor system 320. Sensor system 320 is configured to acquire sensing data about the vehicle itself or its surroundings. Sensor system 320 may, for example, include a Global Navigation Satellite System (GNSS) module 322 (such as GPS) configured to collect geographic location data of vehicle 300. Sensor system 320 may further include one or more sensors 324. Sensors 324 may be any type of onboard sensor such as a camera, LiDAR and RADAR, ultrasonic sensors, gyroscope, accelerometer, odometer, etc. It should be appreciated that sensor system 320 may also provide the possibility of acquiring sensing data directly or via dedicated sensor control circuitry within vehicle 300.

[0110] Vehicle 300 further includes a communication system 326. Communication system 326 is configured to communicate with external units such as other vehicles (i.e., via vehicle-to-vehicle (V2V) communication protocols), remote servers (e.g., cloud servers), databases, or other external devices, i.e., vehicle-to-infrastructure (V2I) or vehicle-to-everything (V2X) communication protocols. Communication system 326 can communicate using one or more communication technologies. Communication system 326 may include one or more antennas. Cellular communication technologies can be used for remote communication, such as to remote servers or cloud computing systems. Additionally, if the cellular communication technology used has low latency, it can also be used for V2V, V2I, or V2X communication. Examples of cellular radio technologies include GSM, GPRS, EDGE, LTE, 5G, 5G NR, and so on, including future cellular solutions. However, in some solutions, short-to-medium range communication technologies such as wireless local area networks (LANs), for example, solutions based on IEEE 802.11, can be used for communication with other vehicles near vehicle 300 or with local infrastructure components. ETSI is developing cellular standards for vehicle communications, and 5G is considered a suitable solution, for example, due to its high bandwidth, low latency, and efficient processing of communication channels.

[0111] The communication system 326 can further provide the possibility of transmitting output to a remote location (e.g., a remote server, operator, or control center) via one or more antennas. Moreover, the communication system 326 can be further configured to allow various components of the vehicle 300 to communicate with each other. As an example, the communication system can provide a local network setup such as CAN bus, I2C, Ethernet, fiber optics, etc. Local communication within the vehicle can also be a wireless type with protocols such as WiFi, LoRa, Zigbee, Bluetooth, or similar medium / short-range technologies.

[0112] Vehicle 300 further includes a control system 320. The control system 328 is configured to control the handling of vehicle 300. The control system 328 includes a steering module 330 configured to control the heading of vehicle 300. The control system 328 further includes a throttle module 332 configured to control the actuation of the throttle valve of vehicle 300. The control system 328 further includes a brake module 334 configured to control the actuation of the brakes of vehicle 300. The various modules of the steering system 328 can receive manual input from the driver of vehicle 300 (i.e., from the steering wheel, accelerator pedal, and brake pedal, respectively). However, the control system 328 can be communicatively connected to the vehicle's ADS 310 to receive instructions on how the various modules should operate. Therefore, ADS 310 can control the handling of vehicle 300.

[0113] As described above, vehicle 300 includes ADS 310. ADS 310 may be part of the vehicle's control system 302. ADS 310 is configured to perform autonomous functions and operations of vehicle 300. ADS 310 may include several modules, each responsible for a different function of ADS 310.

[0114] ADS 310 may include a positioning module 312 or a positioning block / system. The positioning module 312 is configured to determine and / or monitor the geographic location and heading of the vehicle 300, and may utilize data from sensor systems 320, such as data from GNSS module 322. Alternatively or in combination, the positioning module 312 may utilize data from one or more sensors 324. The positioning system may alternatively be implemented as a real-time dynamic (RTK) GPS.

[0115] ADS 310 may further include a perception module 314 or a perception block / system. Perception module 314 may refer to any known module and / or function, for example, included in one or more electronic control modules and / or nodes of vehicle 300, adapted and / or configured to interpret sensing data related to driving of vehicle 300 to identify, for example, obstacles, lanes, relevant signs, appropriate navigation paths, etc. Perception module 314 may therefore be adapted to rely on and receive input from multiple data sources such as automotive imaging, image processing, computer vision, and / or in-vehicle networks, and combine with sensing data, for example, from sensor system 320. The device 200 described above may be provided, for example, as part of perception module 314.

[0116] The positioning module 312 and / or the sensing module 314 can be communicatively connected to the sensor system 320 to receive sensor data from the sensor system 320. The positioning module 312 and / or the sensing module 314 can further transmit control commands to the sensor system 320.

[0117] The ADS may further include a path planning module 316. The path planning module 316 is configured to determine a planned path for the vehicle 300 based on the vehicle's perception and position, as determined by the perception module 314 and the positioning module 312, respectively. The planned path determined by the path planning module 316 can be sent to the control system 328 for execution. As an example, the determined current position of the vehicle on the navigation map can be sent to the path planning module 316.

[0118] The ADS may further include a decision and control module 318. The decision and control module 318 is configured to perform control of the ADS 310 and make decisions. For example, the decision and control module 318 may decide whether the planned path determined by the path planning module 316 should be executed. The decision and control module 318 may be further configured to detect any deviation behavior of the vehicle, such as deviating from the planned path or the expected trajectory of the path planning module 316. This includes evasive maneuvers performed by the ADS 310 and by the vehicle's driver.

[0119] It should be understood that portions of the described solution can be implemented in vehicle 300, in a system located outside the vehicle, or in a combination of inside and outside the vehicle; for example, in a server communicating with the vehicle, i.e., a so-called cloud solution. Different features and principles of the implementation methods can be combined in combinations other than those described. Furthermore, the elements (i.e., systems and modules) of vehicle 300 can be implemented in combinations different from those described herein.

[0120] Figure 4A , Figure 4B and Figure 4C The diagrams illustrate three distinct examples of how the disclosed techniques can be implemented as part of a sensor fusion process. More specifically, these diagrams schematically depict a processing pipeline from new sensor data input to a task-specific predicted output. The illustrated processing pipeline may also be referred to as Sensor Processing Model 400 (or simply the “Model”). Figure 4A The diagram illustrates the entire sensor processing pipeline, although the disclosed technology specifically addresses the fusion portion of the pipeline in its broadest form.

[0121] First go to Figure 4A This demonstrates an example of using the concept of MoE as a fully fused network. In other words, Figure 4A This illustrates the situation where each of the fusion experts constitutes a complete (or independent) fusion network.

[0122] The input to Model 400 is sensor data from one or more sensors, referred to in this paper as "Sensor 1" to "Sensor N". N in this paper refers to any positive integer. The sensor data from the sensors is then fed into an encoding network. In the illustrated example, one encoding network is provided for each sensor. Thus, the encoding network can be specialized for encoding sensor data of a specific sensor data type. However, as explained previously, it is also possible to have encoding networks capable of handling combinations of sensor data from more than one sensor type or sensor instance.

[0123] The encoded representation of the sensor data (as the output from the encoding network) is then fed together with the previous fusion state (referred to above as the "fusion world view representation of the previous moment") into a gate (i.e., the "gating module" above). Therefore, the encoded representation of the sensor data in this paper constitutes the preprocessed representation of the sensor data fed into the gating module.

[0124] The gating module can route the preprocessed representation of sensor data and its previous fusion state to multiple fusion experts (referred to herein as "fusion expert 1", "fusion expert 2" through "fusion expert N"). Arrows (both solid and dashed lines) indicate possible ways the gate can route the data. However, in the illustrative example, dashed arrows are used to indicate routes not selected by the gate, and solid arrows indicate routes selected by the gate.

[0125] It should be noted that gating does not require the use of both the preprocessed representation of the sensor data and the previous fusion state in its decision-making. For example, it can be gating based on either the preprocessed representation of the sensor data or the previous fusion state. However, the gate can still then feed both the preprocessed representation of the sensor data and the previous fusion state to the selected fusion expert.

[0126] The outputs of the selected fusion experts (fusion expert 2 and fusion expert N in this example) are then used to provide the new fusion state (i.e., the "updated fusion world view representation" as mentioned above).

[0127] The new fusion state can be fed into a task-specific prediction network, which generates a task-specific output, i.e., a prediction.

[0128] Figure 4B Another example of sensor processing model 400' is shown. Figure 4B Model 400' and Figure 4A The difference in Model 400 is that the preprocessed representation of the sensor data of the selected subset of the feed gate and fusion expert is the pre-fused representation of the sensor data.

[0129] More specifically, as the output of the encoding network, the encoded representation of the sensor data is processed by a pre-fusion network configured to generate a pre-fused representation of the sensor data. Pre-fusion may include several optical fusion steps upon which gating functions can be applied. Figure 4A Compared to the example shown, fusion experts can then construct a simpler fusion network.

[0130] As previously explained, in some implementations of the disclosed technology, the concept of MoE can be implemented as part of a fusion network. In one such example, the fusion network can be formed by a transformer. The transformer can include several transformer layers. Each transformer layer can further include an attention module and a feedforward neural network (FFN). The attention module can be configured to perform both cross-attention and self-attention. The concept of MoE can then be used as part of the FFN. This example is shown in Figure 4C In this context, one interpretation is that, as described above, the pre-fusion network can be an instance of the fusion network as previously described. Another interpretation is that the preprocessed representation of the sensor data, as mentioned above, can be the output of the previous transformer layer, and combined with... Figures 1A to 1D The described steps occur in the next transformer layer. Therefore, the technique described herein can be repeated (i.e., iteratively executed) to generate an updated fused world view representation. This is in Figure 4C The example illustration is shown in the text.

[0131] Figure 4C Another example of a sensor processing model 400'' is shown. In this example, the fusion network includes several converter layers, or fusion layers 1 to N as shown herein. The fusion layers can be arranged sequentially such that the output of one fusion layer is fed as input to the next. Figure 4C An example of such a fusion layer architecture is shown. It should be noted that fusion layers can have the same or different architectures. However, at least one of the fusion layers employs the concept of MoE as described herein. The fusion layer illustrated is such a fusion layer.

[0132] The first fusion layer in the fusion layer sequence receives encoded sensor data and the previous fusion state as input from the encoding network. This fusion layer then outputs a partially updated fusion state (i.e., a partially updated fused world view representation), which can be fed into the next fusion layer in the sequence. In subsequent fusion layers, the input will be the partially updated fusion state, along with the encoded sensor data and the previous fusion state. The final fusion layer in the sequence then outputs a new fusion state (i.e., an updated fused world view representation).

[0133] Specifically Figure 4CThe diagram illustrates a fusion layer employing multiple fusion experts. In this case, the preprocessed representation of the sensor data fed to the gating module is the output of the attention module from the fusion layer, referred to below as the attention representation of the sensor data. The attention representation of the sensor data can be generated by applying cross-attention and / or self-attention along with a partially updated fused world view representation (i.e., the output of the previous fusion layer, also referred to in this context as the pre-fusion network), one or more encoded representations of the sensor data, and the previous fusion state. Figure 4A and Figure 4B Similar to the previous example, the gate then routes the preprocessed representation of the sensor data to a subset selected by the fusion expert. The output of the fusion layer can either be fed into the next fusion layer in the sequence or provided as a new fusion state.

[0134] It should be noted that the architecture of the fusion layer can also be implemented differently. For example, the attention module can be divided into two separate modules. One module performs cross-attention, and the other module performs self-attention. Furthermore, more than one FFN can be used. One or more of the FFNs can employ fusion experts.

[0135] In short, Figure 4C The example illustrates how a fusion expert can be employed within a transformer-based fusion network. It should be noted that the principles for employing a fusion expert within a transformer-based fusion network can also be applied to transformer-based prediction networks, such as in the following combination... Figure 6 Example of the description.

[0136] Figure 5 The illustration shows another example of how the disclosed techniques can be implemented as part of the coding process of a sensor processing model 500. More specifically, it shows an example of employing expert gating within the coding network.

[0137] Similar to the example above, the input to Model 500 is sensor data from one or more sensors. In this example, the sensor data from each sensor is fed into the corresponding encoding network. However, as explained above, the same encoding network can be used, for example, for several types of sensor data.

[0138] For example, view Figure 5 At the top, sensor data from sensor 1 is fed to the pre-encoder of the encoding network. The pre-encoder is configured to perform some form of precoding to generate a pre-encoded representation of the sensor data.

[0139] The pre-encoded representation of the sensor data can then be fed into the door (i.e., the encoder-gated door control module as mentioned above). Figures 4A to 4CSimilar to the example in [the example], the gate can route its input to a subset of selected experts, in this case, one or more of the encoder experts labeled as encoder expert 1 to encoder expert N.

[0140] The outputs of the encoder experts can then be fed into the fusion network along with the previously fused state. Although the outputs of the encoder experts illustrated are fed into the fusion network individually, they can also be combined in some ways before being fed into the fusion network. For example, the output of an encoder expert in one of the encoding networks (e.g., the network associated with sensor 1) can be formed into an encoded representation of the sensor data before being fed into the fusion network. Similarly, the output of an encoder expert in another encoding network (e.g., the network associated with sensor N) can be formed into another encoded representation of the sensor data before being fed into the fusion network.

[0141] The fusion network then generates a new fusion state, which can then be fed into a task-specific prediction network configured to generate prediction outputs.

[0142] Undoubtedly, this combination Figure 5 The principles of the encoder expert presented can be found in Figures 4A to 4C Implemented in any of the encoding networks. Similarly, in the current example, as combined with Figures 4A to 4C The principles of the fusion expert presented can be found in Figure 5 This is implemented in the fusion network of this example. Therefore, expert gating can be used in both the encoding and fusion processes of the sensor processing pipeline.

[0143] at last, Figure 6 The illustration shows an example of the principles upon which the disclosed techniques are based during the prediction process of sensor processing model 600. More specifically, it shows an example of employing expert gating within the prediction network.

[0144] Similar to the example above, sensor data from one or more sensors (sensor 1 to sensor N) are processed through an encoding network. The encoded representation can then be processed together with the previous fused state through a fusion network to generate a new fused state.

[0145] The new fused state can then be fed into the prediction gating module (labeled Gate in the diagram). The gating module then routes the new fused state to a subset of prediction experts (prediction expert 1 to prediction expert N) who form the prediction network.

[0146] The output of a subset of the prediction experts' selections can then be used to generate a prediction output. This can be done by feeding the output of the subset of prediction experts' selections into a post-prediction network configured to generate the final prediction output. The post-prediction network can, for example, output bounding boxes, planned trajectories, or object classifications. Thus, prediction experts can be used as part of the prediction network. The post-prediction network can be configured differently depending on the downstream task to which its output is targeted. For example, in image classification, the post-prediction network can be configured to determine the average class probability from each prediction expert's output. In trajectory planning / prediction, where each prediction expert outputs a corresponding trajectory with a relevant score, the post-prediction network can be configured to apply softmax to the scores to output the selected trajectory.

[0147] Similar to the above, such as Figure 6 The principles of predictive experts shown in the diagram can be applied to... Figures 4A to 4C and Figure 5 The prediction network in the example. Correspondingly, Figure 5 coding experts and Figures 4A to 4C The principles of fusion experts can be applied to Figure 6 Model 600.

[0148] As those skilled in the art will recognize, the above examples are merely considered as a few non-limiting examples. Modifications may be made to these examples without departing from the disclosed technology.

[0149] Figure 7 The illustration shows a schematic flowchart of a computer-implemented method 700 for processing sensor data in a vehicle equipped with an Automated Driving System (ADS). In some embodiments, method 700 can be viewed as a method 700 for encoding sensor data. Method 700 can be executed in the vehicle (i.e., using the vehicle's computing resources), as in the above combination. Figure 3 Vehicle 300 is described. More specifically, method 700 can be performed as part of ADS.

[0150] The different steps of method 700 are described below in more detail. Even though illustrated in a specific order, the steps of method 700 can be performed in any suitable order and multiple times. Therefore, although the accompanying drawings may show a specific order of method steps, the order of steps may differ from what is depicted. Furthermore, two or more steps may be performed simultaneously or partially simultaneously. This variation will depend on the chosen software and hardware system and the designer's choice. All such variations are within the scope of the disclosed technology. Similarly, software implementation can be accomplished using standard programming techniques based on rule-based logic and other logics to complete the various steps. Further variations of method 700 will become apparent from this disclosure. The embodiments mentioned and described herein are given by way of example only and should not be limited to the technology of this disclosure. Other solutions, uses, purposes, and functions within the scope of the disclosed technology claimed in the patent claims described below will be apparent to those skilled in the art.

[0151] It should be further recognized that, Figures 1A to 1D Method 700 includes steps illustrated with solid lines and steps illustrated with dashed lines. The steps shown with solid lines are those included in the broadest example embodiment of method 700. The steps included with dashed lines are examples of several optional steps that may form part of several alternative embodiments. It should be understood that the optional steps do not need to be performed sequentially. Furthermore, it should be understood that not all optional steps need to be performed. Optional steps can be performed in any order and in any combination.

[0152] Some example implementations of a method 700 for processing sensor data in a vehicle equipped with an autonomous driving system are listed in the following items.

[0153] Item 1. A computer-implemented method 700 for processing sensor data in a vehicle equipped with an autonomous driving system, the method 700 comprising:

[0154] S702 obtains sensor data related to the vehicle's surrounding environment at the current moment, wherein the sensor data includes sensor data associated with one or more sensors of the vehicle;

[0155] The S704 is obtained by including a pre-encoder and a coding network of multiple encoder experts, wherein the encoder experts of the multiple encoders are different from each other in one or more aspects;

[0156] For each of one or more sensors:

[0157] The sensor data associated with the sensor in S706 is processed by a pre-encoder to generate a pre-encoded representation of the sensor data.

[0158] The pre-encoded representation of the S708 sensor data is processed as input to the encoder gating module to generate a subset of encoder experts from multiple encoder experts based on the output of the expert gating module;

[0159] Each encoder expert in a selected subset processes a pre-encoded representation of the S710 sensor data to generate a corresponding output; and

[0160] Based on the output of a subset of encoder experts, an encoded representation of S712 sensor data is provided.

[0161] Item 2, according to the method 700 of Item 1, further includes: processing one or more encoded representations of sensor data associated with each of one or more sensors together with a fused world view representation from a previous time step via a fusion network S714 to generate an updated fused world view representation.

[0162] Item 3, according to method 700 of item 2, wherein the fusion network includes multiple fusion experts; and

[0163] Specifically, processing the encoded representation of sensor data associated with each of one or more sensors together with the fused world view representation from the previous moment in S714 includes:

[0164] The encoded representation of the processed sensor data and / or the fused world view representation about the previous time step are used as input to the fusion gating module to generate a subset of fusion experts from multiple fusion experts based on the output of the fusion gating module;

[0165] Each fusion expert, selected from a subset of fusion experts, processes the encoded representation of the sensor data together with the fused world view representation from the previous time step to generate the corresponding output; and

[0166] Based on the output of a subset of fusion experts, an updated fusion world view representation is provided for the current moment.

[0167] Item 4, the method 700 according to item 2 or 3, further includes: generating the S716 task-specific output by processing the updated fused world view representation via a task-specific prediction network configured to generate task-specific output.

[0168] Item 5, according to method 700 of item 4, wherein the task-specific prediction network includes multiple prediction experts, and

[0169] The outputs specific to the S716 task include:

[0170] The prediction gating module is used to determine a subset of prediction experts from a pool of prediction experts.

[0171] The updated fused world view representation is processed by each predictor in a subset of the predictors to generate the corresponding output; and

[0172] The output of a subset of prediction experts provides task-specific output.

[0173] Item 6, Method 700 according to any one of Items 1 to 5, wherein the multiple encoder experts are a mixture of experts.

[0174] Item 7, Method 700 according to any one of Items 1 to 6, wherein the encoder gating module is configured to apply the learned gating function to the input to generate the output.

[0175] Item 8, Method 700 according to any one of Items 1 to 6, wherein the encoder gating module is configured to apply a deterministic gating function to the input to generate the output.

[0176] Item 9, Method 700 according to any one of Items 1 to 8, wherein the encoder gating module is configured to generate a subset of the encoder expert's selection through the following steps:

[0177] Based on the input, determine the ranking score of each of the multiple encoder experts.

[0178] A subset of the top k encoder experts is selected based on a given ranking score, where k is a positive integer greater than 0 and less than the number of encoder experts among the plurality of encoder experts; and

[0179] Provide a subset of the top k encoder experts as the output.

[0180] Item 10, according to the method 700 of Item 9, wherein the selection of a subset of the top k encoder experts is further based on metadata associated with one or more environmental conditions, which are related to the surrounding environment.

[0181] Item 11, according to the method 700 of Item 9 or 10, wherein the selection of a subset of the top k encoder experts is further based on a subset of encoder experts used to provide a fused world view representation of the previous moments.

[0182] It should be noted that, in conjunction with the above Figures 1A to 1D Any features, principles, or advantages associated with the described method 100 also apply to the combination Figure 7 Method 700 is described, and vice versa.

[0183] Figure 8This is a schematic illustration of a computing device 800 according to some embodiments of the disclosed technology. The computing device 800 can be configured to perform actions such as those described in conjunction with… Figure 7 The method 700 is described. Therefore, the computing device 800 can be a computing device 800 for processing sensor data in a vehicle equipped with an autonomous driving system.

[0184] As described herein, computing device 800 refers to a computer system, or any device or general-purpose computing system configured to perform various functions. Even though computing device 800 is illustrated herein as a single device, it can also be a distributed computing system comprised of several different devices. Computing device 800 may be provided as part of vehicle 300, as described above. Figure 3 Vehicle 300 is described.

[0185] The computing device 800 includes a control circuit 802. The control circuit 802 may physically comprise a single circuit device. Alternatively, the control circuit 802 may be distributed across several circuit devices.

[0186] like Figure 8 As shown in the example, computing device 800 may further include transceiver 806 and memory 808. Control circuitry 802 is communicatively connected to transceiver 806 and memory 808. Control circuitry 802 may include a data bus, and control circuitry 802 may communicate with transceiver 806 and / or memory 808 via the data bus.

[0187] Control circuit 802 can be configured to perform overall control of the functions and operations of computing device 800. Control circuit 802 may include processor 804, such as a central processing unit (CPU), microcontroller, or microprocessor. Processor 804 can be configured to execute program code stored in memory 808 to perform the functions and operations of computing device 800. Control circuit 802 is configured to perform the above-described combination. Figure 7 The steps of method 700 are described. The steps can be implemented in one or more functions stored in memory 808.

[0188] Transceiver 806 is configured to enable computing device 800 to communicate with other entities such as other devices. Transceiver 806 can both transmit data to and receive data from computing device 800. Computing device 800 may, for example, be part of a vehicle. Transceiver 806 can then allow computing device 800 to communicate with other systems within the vehicle, or with external entities (such as other vehicles) or remote servers.

[0189] Memory 808 may be a non-transient computer-readable storage medium. Memory 808 may be one or more of the following: buffer, flash memory, hard disk drive, removable media, volatile memory, non-volatile memory, random access memory (RAM), or other suitable devices. In a typical arrangement, memory 808 may include non-volatile memory for long-term data storage and volatile memory used as system memory for computing device 800. Memory 808 may exchange data with circuit 802 via a data bus. Accompanying control lines and address buses may also exist between memory 808 and circuit 802.

[0190] The functions and operations of the computing device 800 can be implemented in the form of executable logic routines (e.g., lines of code, software programs, etc.), which are stored on a non-transient computer-readable recording medium (e.g., memory 808) in the computing device 800 and executed by the circuit 802 (e.g., using processor 804). In other words, when the circuit 802 is described as being configured to perform a specific function, the processor 804 of the circuit 802 can be configured to execute a portion of program code stored in the memory 808, wherein the stored portion of program code corresponds to the specific function. Furthermore, the functions and operations of the circuit 802 can be independent software applications or part of a software application that performs additional tasks related to the circuit 802. The described functions and operations can be considered as methods by which the corresponding device is configured to perform, such as those described above. Figure 7 Method 700 is discussed. Furthermore, although the described functions and operations can be implemented in software, the functions can also be implemented via dedicated hardware or firmware, or a combination of one or more of hardware, firmware, and software. The functions and operations of computing device 800 are described below.

[0191] Some example implementations of a computing device 800 for processing sensor data in a vehicle equipped with an autonomous driving system are listed in the following items.

[0192] Item 12. A computing device 800 for processing sensor data in a vehicle equipped with an autonomous driving system, the computing device 800 including a control circuit 802 configured to:

[0193] Acquire sensor data relating to the vehicle’s surrounding environment at the current moment, wherein the sensor data includes sensor data associated with one or more sensors of the vehicle;

[0194] Obtain an encoding network that includes a pre-encoder and multiple encoder experts, wherein the encoder experts of the multiple encoders differ from each other in one or more aspects;

[0195] For each of one or more sensors:

[0196] Sensor data associated with the sensor is processed by a pre-encoder to generate a pre-encoded representation of the sensor data;

[0197] The pre-encoded representation of the processed sensor data is used as input to the encoder gating module to generate a subset of encoder experts from multiple encoder experts based on the output of the expert gating module;

[0198] Each encoder expert in a selected subset processes a pre-encoded representation of the sensor data to generate a corresponding output; and

[0199] Based on the output of a subset of encoder experts, an encoded representation of the sensor data is provided.

[0200] Item 13. The computing device 800 according to Item 12, wherein the control circuit 802 is further configured to process one or more encoded representations of sensor data associated with each of one or more sensors together with a fused world view representation from a previous moment via a fusion network to generate an updated fused world view representation.

[0201] Item 14. The computing device 800 according to item 13, wherein the control circuit 802 is further configured to generate task-specific output by processing the updated fused world view representation via a task-specific prediction network configured to generate task-specific output.

[0202] Further attention should be paid to the following, as in the above combination Figure 7 The principles, features, aspects, and advantages of method 700 are also applicable to computing devices 800, as described herein. To avoid unnecessary repetition, refer to the above. Furthermore, in conjunction with the above... Figure 2 Any features, principles, or advantages associated with the described computing device 200 also apply to the combination Figure 8 The computing device 800 is described, and vice versa.

[0203] The disclosed technology has been presented above with reference to specific embodiments. However, other embodiments besides those described above are possible and within the scope of the disclosed technology. Within the scope of the disclosed technology, method steps that perform the method by hardware or software, different from those described above, can be provided. Thus, according to an exemplary embodiment, a non-transient computer-readable storage medium is provided storing one or more programs configured to be executed by one or more processors of a vehicle control system, the one or more programs including instructions for performing the method according to any of the embodiments described above. Alternatively, according to another exemplary embodiment, a cloud computing system can be configured to perform any of the methods presented herein. The cloud computing system may include distributed cloud computing resources that jointly perform the methods presented herein under the control of one or more computer program products.

[0204] It should be noted that any reference numerals in the drawings do not limit the scope of the claims, the disclosed technology can be implemented at least in part by both hardware and software means, and the same hardware item can represent several “apparatus” or “units”.

Claims

1. A computer-implemented method (100) for processing sensor data in a vehicle equipped with an autonomous driving system, the method (100) comprising: (S106) Obtain (S106) one or more preprocessed representations of sensor data relating to the vehicle’s surrounding environment at the current moment, wherein the sensor data includes sensor data associated with one or more sensors of the vehicle; (S108) Obtain multiple fusion experts, wherein the fusion experts among the multiple fusion experts are different from each other in one or more aspects; Processing (S110) the one or more preprocessed representations of the sensor data and / or the fused world view representations about previous moments as input to the fusion gating module to generate a subset of fusion experts from the plurality of fusion experts based on the output of the fusion gating module; Each fusion expert in a subset of the selected fusion experts processes the preprocessed representation of the sensor data together with the fused world view representation from the previous time step (S112) to generate a corresponding output; and Based on the output of a subset of the fusion experts, an updated fusion world view representation is provided (S114) with respect to the current moment.

2. The method according to claim 1, wherein, The preprocessed representation of the sensor data is an encoded representation of the sensor data, which is generated by processing the sensor data via an encoding network.

3. The method according to claim 1, wherein, The preprocessed representation of the sensor data is a pre-fused representation of the sensor data, which is generated by processing the sensor data via a pre-fusion network.

4. The method according to any one of claims 1 to 3, wherein, The preprocessed representation of the sensor data is an attention representation of the sensor data, which is generated by applying cross-attention and / or self-attention together with a partially updated fused world view representation, one or more encoded representations of the sensor data, and the fused world view representation from the previous time step. The partially updated fused world view is represented as obtained from the pre-fused network.

5. The method according to claim 1, wherein, Providing the updated fused world view representation as described in (S114) includes: The output of each fusion expert in a subset of the selected fusion experts is processed by the post-fusion network process (S114a) configured to generate the updated fused world view representation, or The outputs of each fusion expert in the selected subset of fusion experts are combined (S114b) into the updated fusion world view representation.

6. The method (100) according to claim 1, wherein, The multiple fusion experts are a mixture of experts.

7. The method (100) according to claim 1, wherein, The fusion gating module is configured to apply the learned gating function to the input to generate the output.

8. The method (100) according to claim 1, wherein, The fusion gating module is configured to apply a deterministic gating function to the input to generate the output.

9. The method (100) according to claim 1, wherein, The fusion gating module is configured to generate a subset of the fusion experts' selections through the following steps: Based on the input, determine (S110a) the ranking score of each of the plurality of fusion experts. Based on the determined ranking scores, a subset of the top k fusion experts is selected (S110b), where k is a positive integer greater than 0 and less than the number of fusion experts; and The subset of the top k fusion experts selected in (S110c) is provided as the output.

10. The method (100) according to claim 9, wherein, The selection (S110b) of a subset of the top k fusion experts is further based on metadata associated with one or more environmental conditions related to the surrounding environment.

11. The method (100) according to claim 9 or 10, wherein, The selection (S110b) of the subset of the first k fusion experts is further based on the subset of fusion experts used to provide the world view representation of the fusion at the previous moment.

12. The method (100) according to claim 1, further comprising: (S102) Obtain (S102) the sensor data relating to the vehicle’s surrounding environment at the current moment; as well as By processing the sensor data via an encoding network, the sensor data associated with each of the one or more sensors is encoded (S104), thereby generating an encoded representation of the sensor data associated with each of the one or more sensors; The encoding network includes a pre-encoder and multiple encoder experts; and The processing of the sensor data associated with each of the one or more sensors via the encoding network includes, for each of the one or more sensors: The sensor data associated with the sensor is processed by the pre-encoder (S104a) to generate a pre-encoded representation of the sensor data; The encoder gating module is used to determine (S104b) a subset of the encoder experts from the plurality of encoder experts; The pre-encoded representation of the sensor data is processed (S104c) by each encoder expert in the subset of encoder experts to generate a corresponding output; and The output of the subset based on the encoder expert provides (S104d) the encoded representation of the sensor data.

13. The method (100) according to claim 1, further comprising: The task-specific output is generated (S116) by processing the updated fused world view representation via a task-specific prediction network configured to generate task-specific output; The task-specific prediction network includes multiple prediction experts, and Generating the task-specific output described in (S116) includes: The prediction gating module is used to determine (S116a) a subset of the prediction experts from the plurality of prediction experts; The updated fused world view representation is processed (S116b) by each prediction expert in the subset of prediction experts to generate a corresponding output; and The output of the subset based on the prediction expert provides (S116c) the task-specific output.

14. A computer program product comprising instructions that, when executed by a computing device, cause the computing device to perform the method (100) according to claim 1.

15. A computing device (200) for processing sensor data in a vehicle equipped with an autonomous driving system, the computing device (200) including a control circuit (202) configured to: Obtain one or more preprocessed representations of sensor data relating to the vehicle's surrounding environment at the current moment, wherein, The sensor data includes sensor data associated with one or more sensors of the vehicle; A plurality of fusion experts are obtained, wherein the fusion experts among the plurality of fusion experts are different from each other in one or more aspects; The one or more preprocessed representations of the sensor data and / or the fused world view representations about previous moments are used as input to the fusion gating module to generate a subset of fusion experts from the plurality of fusion experts based on the output of the fusion gating module; Each fusion expert in a subset of the selected fusion experts processes the preprocessed representation of the sensor data together with the fused world view representation from the previous time step to generate a corresponding output; and Based on the output of a subset of the fusion experts, an updated fusion world view representation is provided for the current moment.