Change detection criteria for updating a sensor-based reference map

By using a self-supervised learning machine learning model and employing a common sense engine to evaluate the differences between sensor data and reference maps, the problem of inaccurate map updates in existing technologies is solved. This enables fast and accurate map updates under extreme conditions, thereby improving the safety and navigation accuracy of autonomous driving systems.

CN114627440BActive Publication Date: 2026-05-29APTIV TECHNOLOGIES AG

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
APTIV TECHNOLOGIES AG
Filing Date
2021-12-10
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately identify environmental changes under extreme conditions, leading to a decline in the safety and performance of autonomous driving systems.

Method used

A self-supervised learning machine learning model is used to evaluate the differences between sensor data and reference maps through a common sense engine, and the map is updated using change detection criteria to ensure the accuracy and timeliness of the map.

Benefits of technology

It enables rapid and accurate map updates under extreme conditions, improving the safety and navigation accuracy of autonomous driving systems, reducing unnecessary map updates, and maintaining the real-time nature and detail of maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114627440B_ABST
    Figure CN114627440B_ABST
Patent Text Reader

Abstract

This document describes change detection criteria for updating a sensor-based map. Based on detecting an indication of a registered object in the vicinity of a vehicle, a processor determines a difference between a characteristic of the registered object and a characteristic of a sensor-based reference map. A machine learning model is trained using self-supervised learning to identify change detection from inputs. The model is executed to determine whether the difference satisfies change detection criteria for updating the sensor-based reference map. If the change detection criteria are satisfied, the processor causes the sensor-based reference map to be updated to reduce the difference, which enables the vehicle to safely operate in an autonomous mode using the updated reference map for navigating the vehicle in a vicinity of a coordinate location of the registered object. The map can be updated concurrently as changes occur in the environment and without impacting performance, enabling real-time perception to support control and improve driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 124,512, filed December 11, 2020, pursuant to 35U.SC119(e), the disclosure of which is incorporated herein by reference in its entirety. Background Technology

[0003] Some automotive systems rely on reference maps for autonomous or semi-autonomous driving. For example, when operating in extreme conditions, such as at night on a poorly lit road, radar can be a useful sensor for conveying features, vegetation, embankments, bridge extensions, access hatches, or other obstacles as depicted in a reference map. Reliance on these reference maps derived from sensor data can lead systems operating vehicles or convoys to make safe driving decisions. Responsible feedback (e.g., from humans or machines) and quality assurance can be used to ensure the maps remain up-to-date. Automated updates to simultaneously capture changes occurring in the real world contribute to a higher level of driving safety. The difficulty of this automation stems from attempting to quickly and accurately identify so-called “change detections” in the environment. Change detections are markers or indicators in sensor data that correspond to identifiable or missing features in a reference map of that environment. Some systems analyze camera images (e.g., airborne, infrastructure) or other sensor data to help identify change detections and automatically trigger reference map updates. However, these automated attempts often fail or become too cumbersome to rely on, especially when attempting to update for any possible change detections that may occur; relying on it not only degrades performance but can also impede driving safety. Summary of the Invention

[0004] This document describes change detection criteria for updating a sensor-based map. In one example, the method includes: receiving from a sensor device of a vehicle an indication that a registered object has been detected in a vicinity of the vehicle; and having a processor of the vehicle determine, based on the indication, a difference between features of the registered object and features of a sensor-based reference map, the features of which include a map location corresponding to the coordinate location of the registered object. The method further includes: having the processor execute a machine learning model trained using self-supervised learning to identify change detections from inputs to the model, whether the difference satisfies the change detection criteria for updating the sensor-based reference map; and, in response to determining that the difference satisfies the change detection criteria, having the processor update the sensor-based reference map to reduce the difference. The method additionally includes: having the processor operate the vehicle in an autonomous mode that relies on the sensor-based reference map for navigating the vehicle in a vicinity of the coordinate location of the registered object.

[0005] This document also describes a system comprising a processor configured to perform the methods described herein and other methods, and a computer-readable storage medium including instructions that, when executed, cause the processor to perform the methods described herein and other methods. Furthermore, this document describes other systems configured to perform the methods summarized above and other methods described herein.

[0006] This invention presents a simplified concept for a change detection criterion for updating sensor-based maps, which will be further described below in the detailed description and accompanying drawings. This invention is not intended to identify essential features of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter. That is, one advantage provided by the described change detection criterion is the ability to quickly and accurately identify change detections from sensor data, relying on which map updates are triggered. Although primarily described in the context of radar-based maps and language-based self-supervised learning methods, the change detection criterion for updating sensor-based maps described herein can be applied to other sensor-based reference maps (e.g., LiDAR-based, image-based), where improved navigation and control accuracy is desired while preserving processing resources and keeping the map up-to-date, and other self-supervised learning methods besides language-based methods can be used. Attached Figure Description

[0007] This document describes in detail one or more aspects of the change detection criteria used to update sensor-based maps, with reference to the following diagram:

[0008] Figure 1-1 An example environment is shown in which change detection criteria are used to update a sensor-based map according to the technology of this disclosure;

[0009] Figure 1-2 An example process for updating a sensor-based map using change detection criteria, according to the techniques of this disclosure, is shown.

[0010] Figure 1-3 An example vehicle is shown for updating a sensor-based map according to the use change detection criteria of the technology disclosed herein;

[0011] Figure 2-1 and Figure 2-2 An example scenario is shown illustrating the use of the techniques according to this disclosure for updating change detection criteria for sensor-based maps;

[0012] Figure 3This is a conceptual diagram illustrating adversarial matching as part of a change standard used to update sensor-based maps, based on the technology disclosed herein.

[0013] Figure 4-1 and Figure 4-2 An example of a common sense engine for updating sensor-based maps using change detection criteria, according to the techniques of this disclosure, is shown.

[0014] Figure 5 Another example process for updating a sensor-based map using change detection criteria, according to the techniques of this disclosure, is shown;

[0015] Figure 6-1 , Figure 6-2 , Figure 7-1 , Figure 7-2 , Figure 8-1 , Figure 8-2 , Figure 9-1 and Figure 9-2 Additional example scenarios are shown, illustrating the use of the techniques according to this disclosure for updating change detection criteria for sensor-based maps.

[0016] The same numbers are often used throughout the accompanying drawings to refer to similar features and parts. Detailed Implementation

[0017] Overview

[0018] Automatically identifying change detections for updating sensor-based maps can be challenging. This document describes a method for updating sensor-based maps using change detection criteria, compared to other methods of updating reference maps. Based on indications of a registered object detected near a vehicle, the processor determines the differences between the features of the registered object and the features of the sensor-based reference map. A machine learning model is trained using self-supervised learning to identify change detections from the input. The model is executed to determine if the differences meet the change detection criteria used to update the sensor-based reference map. If the change detection criteria are met, the processor updates the sensor-based reference map to reduce the differences, enabling the vehicle to operate safely in autonomous mode using an updated reference map for navigating the area near the coordinates of the registered object. The map can be updated synchronously as changes occur in the environment without being over-updated for changes that shouldn't be reflected, resulting in better real-time perception to aid control and improve driving safety.

[0019] Therefore, the technology disclosed herein enables a self-supervised learning approach to create criteria to be applied when determining whether a change detection is sufficient to warrant a map update. The commonsense engine is a machine learning model that evaluates each change detection to identify differences in features or attributes necessary for map updates. Through self-supervised learning techniques, the commonsense engine learns and constructs change detection criteria. These criteria represent a knowledge base that enables the commonsense engine to answer questions based on natural language and point clouds related to phenomena observed in a pretextual task. Unlike other techniques used to identify change detections, the commonsense engine can process point cloud data quickly and accurately, which has rich associative features not only at the geographic layer but also at the semantic layer (e.g., for security). Thus, when changes in road geometry or traffic objects are detected in sensor data relative to a sensor-based reference map, the commonsense engine uses real-time criterion operations to detect roundabout types, construction closures, erosion, and other features that may be missing from the reference map because they were not visible or present at the time of map creation.

[0020] Example Environment

[0021] Figure 1-1 An example environment 100 is illustrated in which change detection criteria are used to update a sensor-based map according to the techniques of this disclosure. Environment 100 includes a vehicle 102, a network 104 (e.g., the Internet), a remote system 106 (e.g., a server), and several other vehicles 116. The change detection criteria can be used by one or more of the entities shown in environment 100 to update a sensor-based map for autonomous or semi-autonomous vehicle navigation and vehicle control. The techniques described below can be performed at the remote system 106 by communicating with vehicle 102 on network 104. The several other vehicles 116 can perform similar techniques, such as executing a machine learning model trained to identify the change detection criteria. Similarly, vehicle 102 can execute a model that performs the same operation to update a sensor-based reference map.

[0022] The vehicle includes a processor 108 (or other similar control circuitry system) operatively coupled to sensor device 110. As some examples, sensor device 110 is shown to include multiple cameras 110-1 (e.g., optical, infrared), multiple position sensors 110-2 (e.g., positioning systems, accelerometers, barometers), and multiple distance / distance change rate sensors 110-3, such as radar, lidar, and ultrasonic sensors. Sensor device 110 generates sensor data 112, and processor 108 analyzes sensor data 112 to obtain change detection.

[0023] Sensor device 110 is configured to identify indications of registered objects 118 identifiable in the field of view and report these indications to processor 108. Indications of registered objects 118 can be stored as sensor data 112. Sensor data 112 is compared with a sensor-based reference map 114 to enable vehicle 102 to navigate safely, and in some cases, to navigate safely in the immediate vicinity of registered objects 118. The sensor-based reference map 114 can be stored locally by vehicle 102 (as shown) or at least accessible by vehicle 102 via network 104 (e.g., stored at a remote system 106 and accessible as a map service).

[0024] For ease of description, the following examples are described primarily in the context of execution on the processor 108 of vehicle 102. Remote system 106 or multiple other vehicles 116 may perform similar techniques to update sensor-based maps in response to criteria used for change detection. In other words, the described techniques can be distributed and executed across components of environment 100, or executed individually on remote system 106 or solely on the processor 108 of vehicle 102.

[0025] The locations represented by sensor data 112 and map 114 can be so accurate that comparisons and matches of changes in road geometry (e.g., roundabout type, lane width, number of lanes) or road infrastructure (e.g., removal or addition of traffic cones, removal or addition of signs, removal or addition of traffic barriers) can be based on their overlap. This is the basis of change detection theory.

[0026] Change detection theory enables vehicle 102 to handle significant changes in environment 100, such as closed ramps or lanes, changed routes, and new roundabouts. However, autonomous vehicles, such as vehicle 102, require higher-resolution or more detailed reference maps to achieve safe and accurate autonomous driving. The detailed features of reference map 114 are affected by changes in the real world due to weather, time, or various other factors. These changes in detailed features can be categorized into two types based on their impact on use cases.

[0027] A so-called "minor change" can render map 114 invalid because it is an incorrect representation of environment 100 or registered object 118. Minor changes do not impede the ultimate goal of safe navigation and operation of vehicle 102. Conversely, "major changes," compared to minor changes, limit or hinder the use of map 114, thus limiting or completely prohibiting autonomous driving functions. For example, minor changes are primarily caused by traffic accidents or weather and are therefore usually unintentional. Examples are vehicles parked on the side of the road, dents, scratches, or turf in guardrails, missing or worn lane markings, damaged or displaced signs and poles. By definition, these types of minor changes can alter the orientation of vehicle 102 or impede the usability of map 114 in supporting autonomous driving modes. Positioning systems that may use such map features as landmarks typically rely on a large number of these landmarks, allowing the assumption that most landmarks remain unchanged and the positioning system still functions. Therefore, while map 114 cannot be fully verified for such minor changes, minor changes do not render map 114 invalid for autonomous driving. Major changes are primarily caused by severe weather, road works, or maintenance projects, and unlike minor changes, they are intentional or drastic. Examples include repaving or resurfacing roads, road erosion, landslides across roads, adding one or more lanes, or rebuilding roads to update the road layout. Almost everything on Map 114 should be reliable if used as a reference for positioning and navigation. Therefore, combined changes to landmarks, such as replacing guardrails, can also constitute major changes.

[0028] Figure 1-2 An example system 102-1 is shown that updates a sensor-based map using a change detection criterion according to the technology of this disclosure. System 102-1 is an example of vehicle 102 and similar to vehicle 102, including a processor 108, sensor devices 110, sensor data 112, and a map 114. Components of system 102-1 communicate via a bus 160, which may be a wired or wireless communication bus. Communication device 120, computer-readable storage medium 122, and driving system 130 are coupled to other components of system 102-1 via bus 160. The computer-readable storage medium stores a common sense engine 124, which includes a machine learning model 126 that relies on change detection criterion 128 to determine whether changes in the environment necessitate updating the map 114.

[0029] The common sense engine 124 may be implemented at least partially in hardware, for example, when causing the software associated with the common sense engine 124 to execute on processor 108. Therefore, the common sense engine 124 may include hardware and software, such as instructions stored on computer-readable storage medium 122 and executed by processor 108. The common sense engine 124 constructs a machine learning model 126, which relies on change detection criteria 128 to identify major changes that need to be made to map 114. The machine learning model 126 constructed by the common sense engine 124 is configured to identify change criteria 128 for updating map 114.

[0030] Figure 1-3 An example process 140 for updating a sensor-based map using change detection criteria according to the technology of this disclosure is shown. Process 140 is shown as a set of operations 142 to 154, which may be referred to as actions or steps, and are performed in, but not limited to, the order or combination of operations shown or described. Furthermore, any of the operations 142 to 154 may be repeated, combined, or rearranged to provide other methods. Reference may be made in the sections discussed below. Figure 1-1 Environment 100 and Figure 1-1 and Figure 1-2 The entities detailed herein are for illustrative purposes only. This technique is not limited to being performed by one or more entities.

[0031] Figure 1-3 This illustrates how change detection can generally work. At 142, a reference map is acquired from sensor device 110 of vehicle 102. At 144, objects detected by sensor device 110 (such as radar) are registered on the reference map. At 146, sensor data 112 is received and common sense engine 124 executes at processor 108 to answer natural language-based and point cloud-based questions in reasoning about common sense phenomena observed from the sensor data while vehicle 102 is in motion. At 148, common sense engine 124 determines the differences between reference map 114 and the registered objects in the region of interest (ROI) associated with vehicle 102. In other words, common sense engine 124 can limit its evaluation of features of map 114 for updates based on a portion of sensor data 112 (specifically, a portion indicating one or more objects in the ROI).

[0032] At point 150, the self-supervised learning of the common sense engine 124 enables it to create its own criteria 128 through nominal tasks in reasoning based on natural language and point clouds, for checking whether change detection is sufficient. Self-supervised learning is a version of unsupervised learning (where data typically provides supervision), and the neural network's task is to predict the remaining missing data. Self-supervised learning enables the common sense engine 124 to fill in details indicating features in environment 100 that differ from expectations or are expected but missing from sensor data 112. These details are predicted and, depending on the quality of sensor data 112, can yield acceptable semantic features without applying actual labels. If sensor data 112 includes change detection indicating features in map 114 that are sufficiently different from the attributes of the corresponding location in environment 100, the common sense engine 124 updates map 114.

[0033] At point 152, the difference between sensor data 112 and map 114 is quantified. And at point 154, map 114 is modified to eliminate or at least reduce the difference between sensor data 112 and map 114.

[0034] Figure 2-1 and Figure 2-2 An example scenario is shown illustrating the use of the techniques according to this disclosure for updating change detection criteria for sensor-based maps. Figure 2-1 and Figure 2-2 Scenes 200-1 and 200-2 respectively illustrate the geographical and semantic driving scenarios in two-dimensional bird's-eye view.

[0035] exist Figure 2-1 In scenario 200-1, as vehicle 102 travels along the road, tow truck 202 is parked in the rightmost lane. The vehicle 102's sensor device 110 detects one or more traffic cones 204 and / or flares 206, positioned on the road to warn other drivers of the parked tow truck 202. The vehicle 102's common sense engine 124 determines that, under normal circumstances, such a parked vehicle would occupy at most a single lane. The common sense engine 124 can determine that the traffic cones 204 are positioned outside a single lane and therefore could constitute a change detection for updating map 114. However, because they are traffic cones 204 and not obstacles or some other structure, the common sense engine 124 can avoid updating map 114, as scenario 200-1 is unlikely to be a construction zone associated with a major change and is likely only a temporary, minor change, not worth updating map 114.

[0036] In contrast, Figure 2-2In scenario 200-2, traffic cones 210 for construction zone 208 with construction sign 212 typically occupy at least one lane. The sensor device 110 of vehicle 102 can report construction sign 212 and / or traffic cones 210 as a point cloud portion of sensor data 112 identifying one or more registered objects, where construction sign 212 and / or traffic cones 210 would likely be interpreted by common sense engine 124 as a major variation. The ROIs of these objects 210 and 212 can be analyzed by common sense engine 124 along with scene context, such as scene labels and detection labels. This analysis aims to question and resolve challenging visual questions indicating differences between objects 210 and 212 and their representation or absence in map 114, while also providing root causes to support these answers.

[0037] Conceptually, the common sense engine 124 can reason in multiple different geographic or semantic spaces. The common sense engine is a machine learning model that processes sensor data 112 and can therefore “reason” in two different scenarios: a geographic scenario and a semantic scenario. Given point cloud sensor data 112, a list of registered objects from the sensor data 112, and a question (e.g., “What is this?”), the common sense engine 124 is trained through self-supervision to answer the question and provide root causes explaining why the answer is correct. Self-supervision enables the common sense engine 124 to perform seemingly infinite loop learning by answering increasingly challenging questions that go beyond simple visual or recognition-level understanding, moving towards a higher-order cognitive and common-sense understanding of the world depicted by the point cloud from the sensor data 112.

[0038] The Common Sense Engine 124's task can be broken down into two multiple-choice subtasks, corresponding to answering question q with a response r and providing a valid reason or root cause. An example of a subtask might include:

[0039] 1. Point cloud P,

[0040] 2. The sequence of object detections, o. Each object detection o i Composed of bounding box b, segmentation mask m, and category label l i ∈L constitutes.

[0041] 3. The query q is presented using a mixture of natural language and pointers. Each word q in the query... i It is a word in the vocabulary V, or a label pointing to an object in o.

[0042] 4. A set of N responses, where each response r (i)The syntax for this is the same as for queries: it uses natural language and directives. Only one response is correct.

[0043] 5. The model selects a single optimal response.

[0044] In a question-and-answer sequence with a response r and a valid reason or root cause q, the query is question q, and the response r can be the answer option. In a valid reason or root cause answer, the query is the concatenated question and the correct answer, while the response is the root cause option.

[0045] Common Sense Engine 124 can execute two models: one for calculating the correlation P between the question q and the response. rel And another is used to calculate the similarity P between two response options. sim It uses Bidirectional Encoder Representations from Transformers (BERT) for natural language inference. BERT can be based on convBERT (see https: / / arxiv.org / pdf / 2008.02496.pdf). In a given dataset example (q... i ,r i ) 1≤i≤N In this case, the weight matrix W∈R can be adjusted. N×N (by W) i,j =log(P rel (q i ,r j ))+μlog(1-P sim (r i ,r j Given that maximum weighted bipartite matching is performed for each q, i Obtain counterfacts. μ>0 controls the trade-off between similarity and relevance. To obtain multiple counterfacts, several binary matching operations can be performed. To ensure that the negations are diverse, candidate responses r can be used during each iteration. j With the current assignment to q i The maximum similarity among all responses is used to replace the similarity item.

[0046] BERT and convBERT are just two examples of transformers. Other types of transformers are unified cross-modal transformers, which model both image and text representations. Other examples include ViLBERT and LXMERT, which are based on two-stream cross-modal transformers that bring more specific representations to images and language.

[0047] Although described mainly in the context of radar-based maps and language-based self-supervised learning methods, the change detection criteria for updating sensor-based maps described herein can be applied to other sensor-based reference maps (e.g., lidar-based, image-based), in which case it is desirable to improve the accuracy of navigation and control while still conserving processing resources and keeping the map up-to-date, and other self-supervised learning methods besides language-based methods can be used. For example, an alternative to LSTM is the Neural Circuit Policy (NCP), which is much more efficient than LSTM and uses far fewer neurons than LSTM (see https: / / www.nature.com / articles / s42256-020-00237-3).

[0048] In some examples, the commonsense engine 124 employs adversarial matching techniques to create a robust multi-option dataset at scale. An example of such a dataset is conceptualized in Figure 3 .

[0049] Figure 3 is a conceptual diagram showing adversarial matching as part of the technique according to the present disclosure for using change criteria to update a sensor-based map. On the left, the circles represent queries q1 to q4, and on the right, the circles are used to show potential responses r1 to r4. Incorrect options are obtained via a maximum weight bipartite matching between queries q1 to q4 and responses r1 to r4. The weights associated with the pairing of responses and questions are shown as line segments connecting the circles; where thick line segments represent higher weights and, by contrast, thin line segments represent lower weights. The weights are scores from a natural language inference model.

[0050] To narrow the gap between recognition (e.g., detecting objects and their attributes) and the level of cognition (e.g., inferring the possible intentions, goals, and social dynamics of moving objects), the commonsense engine 124 performs adversarial matching to achieve grounding of the meaning of natural language passages in sensor data 112, understanding of responses in the context of questions, and reasoning about the underlying understanding of questions, shared understanding of other questions and answers, to identify meaning from the differences between expected point cloud data and measured point cloud data when using map 114 as a relative baseline for change.

[0051] As in Figure 4-1 and Figure 4-2As explained in the description, the common sense engine 124 executes a machine learning model that performs three reasoning steps. First, the model grounds the meaning of natural language paragraphs about directly referenced objects from sensor data 112. The model then contextualizes the meaning of the response to the queried question, as well as global objects from sensor data 112 that are not mentioned in the question. Finally, the model reasones on this shared representation to arrive at the correct response. The common sense engine 124 is configured to collect questions, correct answers, and correct root causes through adversarial matching. One approach to collecting questions and answers to various common sense reasoning questions on a large scale is to carefully select situations of interest that involve many different registered objects or scenarios where many things may change.

[0052] Adversarial matching involves reusing or repeating exactly three times the negative answers for three other questions for each correct answer to a question. Therefore, each answer has the same probability of being correct (25%): this addresses the problem of only answering biases and prevents the machine from always choosing the most general answer, which doesn't offer much benefit when there's a better understanding. Commonsense Engine 124 can formulate the answer reuse problem as a constrained optimization based on the relevance and entailment scores between each candidate negative answer and the best answer as measured by a natural language inference model. This adversarial matching technique allows any language-generated dataset to be transformed into a multiple-choice test with minimal human intervention.

[0053] One challenge encountered was obtaining counterfactuals (i.e., incorrect responses to a question). This could be addressed by performing two separate subtasks: ensuring that counterfactuals were as relevant to the context of the environment as possible, thus attracting the machine; however, counterfactuals could not be too similar to correct responses to prevent them from accidentally becoming correct. These two objectives were balanced to create a training dataset that was challenging for the machine but easy for humans to verify accuracy. Adversarial matching is characterized by the ability to use variables to set a tradeoff between difficulty for humans and difficulty for the machine. In most examples, the question should be difficult for the machine but easy for humans. For example, adjusting a variable in one direction could make the question more difficult for the common sense engine 124 to respond to, but easier for an operator to know through experience and intuition whether the response was correct. This visual understanding of the sensor data 112 could correctly answer the question; however, the confidence of the common sense engine 124 came from its understanding of the underlying reasons it provided for reasoning.

[0054] The common sense engine 124 is configured to provide root causes explaining why an answer is correct. The question, answer, and root cause can be preserved as a mixture of rich natural language and other indications of cloud data density and feature shape (e.g., detection labels). Maintaining the question, answer, and root cause together in a single model allows the common sense engine 124 to provide explicit links between the textual description of the registered object (e.g., “cone traffic sign 5”) and the corresponding point cloud region in 3D space. To simplify the evaluation, the common sense engine 124 divides the final task into answer and proof phases in a multi-option setting. For example, given a question q1 and four answer options r1 to r4, the common sense engine 124 model first selects the correct answer. If its answer is correct, four root cause options (not shown) are provided, which can claim to prove the answer is correct, and the common sense engine 124 selects the correct root cause. Whether a prediction made by the common sense engine 124 is correct can depend on the correctness of both the selected answer and the subsequently selected root cause.

[0055] Figure 4-1 and Figure 4-2 An example of a common sense engine 144-1 for updating a sensor-based map using a change detection criterion, according to the technology of this disclosure, is shown. The common sense engine 144-1 is divided into four parts, including an initialization component 402, a basicization component 404, a contextualization component 406, and a reasoning component 408.

[0056] Initialization component 402 may include a convolutional neural network (CNN) and BERT to learn a joint point cloud language representation of each token in the sequence passed to contextualization component 404. Because both queries and responses can contain a mixture of labels and natural language words, the same initialization component 402 is applied to each (allowing them to share parameters). At the core of initialization component 402 is a bidirectional LSTM, which is passed as input at each location, w i Word representation and The features are then used. CNNs are used for optional object-level features: the visual representation of each region o is aligned with the ROI from its boundary regions. Additionally, the object's category label l is encoded. o The relevant information will be l o The embeddings (along with the visual features of the object) are projected into a shared hidden representation. The output of the LSTM at all locations is r for the response and q for the query.

[0057] Alternatives to CNNs can be used; for example, Faster R-CNN can extract visual features (e.g., pooled ROI features for each region), which can encode the localization features of each region via a normalized multidimensional array including elements of coordinates (e.g., top, left, bottom, right), dimensions (e.g., width, height, area), and other features. Thus, the array could include: [x1, y1, x2, y2, w, h, w*h]. Both the visual and location features from this array are then fed through fully connected (FC) layers to be projected into the same embedding space. The final visual embedding for each region is obtained by adding the two outputs from the FC and then passing this sum through a layer normalization (LN) layer.

[0058] Given initial representations of the query and response, the basic component 402 uses an attention mechanism to contextualize these sentences relative to each other and the point cloud context. For each position i in the response, the query representation of interest is defined using the following equation:

[0059] α i,j =softmax(r i Wq j )and To contextualize the answer, including implicitly relevant objects not yet extracted from the base component 402, another bilinear attention is performed at the contextualization component 406 between the response r and the features of each object o. The result of this object attention is...

[0060] Finally, the reasoning component 408 of the machine learning model 126 of the commonsense engine 124 infers the response, the query of interest, and the object to output an answer. The reasoning component 408 uses a bidirectional long short-term memory (LSTM) to achieve this, where the LSTM is given context for each position i. r i ,as well as To better facilitate gradient flow through the network, the output of the inference LSTM is concatenated with the question and answer representations at each time step: the resulting sequence is max-pooled and passed through a multilayer perceptron, which predicts the query-response compatibility logic.

[0061] In some examples, the neural network of machine learning model 126 can be based on a previous model, such as ResNet50 for image features. To obtain a strong representation of the language, BERT representation can be used. BERT is applied to the entire question and answer options, and extracts a feature vector for each word from the penultimate layer. Machine learning model 126 minimizes the feature vector for each response r.i The model is trained using the multi-class cross-entropy between the predictions and the golden label. The aim is to provide a fair comparison between Machine Learning Model 126 and BERT, therefore using BERT-Base for each model is also a possibility.

[0062] The goal of machine learning model 126 is to use BERT as simply as possible and treat it as a baseline. Given a query q and response options r... (i) Both are merged into a single sequence to be provided to BERT. Each token is a sequence corresponding to a different transformer unit in BERT. Later layers can then be used in BERT to extract contextualized representations for each token in the query and response.

[0063] This provides a different representation for each response option i. The frozen BERT representation can be extracted from the penultimate layer of its transformer. Intuitively, this makes sense because these layers are used for two pre-training tasks of BERT: next-sentence prediction (the unit corresponding to the token at the last layer L focuses on all units at layer L-1, and is also used to focus on all other units). The trade-off is that pre-compiling the BERT representation significantly reduces runtime, and the machine learning model 126 focuses on learning a more robust representation.

[0064] In some cases, it is desirable to include simple settings in the machine learning model 126, which would allow for adjustments for certain scenarios and, where possible, to use a similar configuration for the baseline, especially in terms of learning rate and hidden state size.

[0065] Based on the performance of the described technique, it has been found in some examples that the projection of point cloud features maps a 2176-dimensional hidden size (2048 and 128-dimensional class embeddings from ResNet50) to a 512-dimensional vector. The basification component 404 can include an LSTM as a single-layer bidirectional LSTM with a 1280-dimensional input size (768 from BERT and 512 from point cloud features) and using 256-dimensional hidden states. The inference component 408 can rely on an LSTM that is a two-layer bidirectional LSTM with a 1536-dimensional input size (512 from point cloud features and 256 for each direction in the basification query and basification answer of interest). This LSTM can also use 256-dimensional hidden states.

[0066] In some examples, the representations from the LSTM of the inference unit 408, the basified answer, and the question of interest are max-pooled and projected onto a 1024-dimensional vector. This vector can be used to predict the i-th multivariate logit. The hidden-hidden weights of all LSTMs in the commonsense engine 124 can be set using orthogonal initialization and pdrop =0.3 applies cyclic dropout to the LSTM input. This model can be optimized to have 2*10 -4 learning rate and 10 -4 Weight decay. When a plateau occurs (validation accuracy does not increase for two consecutive epochs), pruning the gradients to have the total L2 norm can reduce the learning rate by half. In some examples, each model can be trained for up to 20 epochs.

[0067] Figure 5 Another example process for updating a sensor-based map using change detection criteria according to the technology of this disclosure is shown. Method 500 is shown as a set of operations 502 to 510 performed in the order or combination of the operations shown or described. Furthermore, any of operations 502 to 510 may be repeated, combined, or rearranged to provide other methods, such as process 120. In the various sections of the following discussion, reference may be made to environment 100 and the entities detailed above, which are referred to by way of example only. This technology is not limited to being performed by one or more entities.

[0068] At point 502, an indication of a registered object is detected in the vicinity of the vehicle. For example, sensor device 110 generates sensor data 112, which includes point cloud data of the environment 100 and the object 118. Processor 108 acquires sensor data 112 via bus 160.

[0069] At point 504, based on this indication, the differences between the features of the registered object and the features of the sensor-based reference map are determined. The features of the sensor-based reference map include map locations corresponding to the coordinate locations of the registered object. For example, portions of sensor data 112 and portions of map 114 may overlap at the same coordinate locations; differences between features at the same coordinate locations indicate reasonable possible changes detected in updating map 114.

[0070] At point 506, a machine learning model is executed, trained using self-supervised learning to identify change detections from the inputs given to the model. For example, processor 108 executes a commonsense engine 124, which compares differences to change detection criteria. The commonsense engine 124 can be designed to update a map 114, which can be a radar-based reference map or any sensor-based reference map. Map 114 can include multiple layers, each for a different sensor. For example, a first layer for recording radar-based features can be aligned and matched with a second layer (such as a LiDAR layer or a camera layer) that records features aligned with the features of the first layer.

[0071] At point 508, in response to determining that the difference meets the change detection criteria, the sensor-based reference map is updated to reduce the difference. For example, common sense engine 124 relies on change detection criteria 128 to determine whether map 114 should be updated in response to a specific change detection. Differences between features of sensor data 112 and features of map 114 can be observed at common coordinate locations. These differences can be identified as inconsistencies between sensor data 112 and map 114 regarding things such as:

[0072] • The expected distance from vehicle 102 to the center island and obstacles, which can be used to infer the roundabout type;

[0073] • The expected distance between vehicle 102 and adjacent obstacles, which can be used to determine lane width and the number of lanes;

[0074] • Expected obstacle curvature of ramp curvature;

[0075] • The expected distance between the cone-shaped traffic sign and its overall shape;

[0076] • Expected distance to other traffic obstacles and distance to the guardrails of the traffic obstacles;

[0077] • Expected traffic signs;

[0078] When sensor data 112 includes radar data, differences in radar-specific features can be identified. Utilizing these differences allows for more accurate identification of changes detected in the radar layer of map 114. These radar features may include:

[0079] Expected signal strength

[0080] ·Expected peak sidelobe ratio

[0081] ·Expected signal-to-noise ratio

[0082] • Expected radar cross-section

[0083] • Expected constant false alarm rate

[0084] • Expected transmit / receive antenna gain

[0085] • Expected static object detection

[0086] ○Vegetation

[0087] ○ Embankment

[0088] ○ Bridge expansion

[0089] ○ Speed ​​bumps, access holes, or drain pipes

[0090] ○ Figure 2-1 and Figure 2-2 Scenes 200-1 and 200-2 respectively illustrate the geographical and semantic driving scenarios in two-dimensional bird's-eye view.

[0091] At point 510, the vehicle operates in an autonomous mode that relies on a sensor-based reference map for navigation in the vicinity of the registered object's coordinates. For example, vehicle 102 avoids construction zone 208, traffic cone 210, and sign 212 in response to the identification of a construction zone, which is achieved by updating map 114, with the common sense engine 124 ensuring that features of construction zone 208 appear in the sensor-based reference map 114. In this way, the techniques of this disclosure enable the use of point clouds with detailed features at both the geometric and semantic (e.g., safety) layers. Self-supervised learning enables the common sense engine to create its own supervision through questions and responses to nominal tasks.

[0092] Figure 6-1 , Figure 6-2 , Figure 7-1 , Figure 7-2 , Figure 8-1 , Figure 8-2 , Figure 9-1 and Figure 9-2 Additional example scenarios are shown, illustrating the use of the techniques according to this disclosure for updating change detection criteria for sensor-based maps. Figure 6-1 , Figure 6-2 , Figure 7-1 , Figure 7-2 , Figure 8-1 , Figure 8-2 , Figure 9-1 and Figure 9-2 Scenes 600-1, 600-2, 700-1, 700-2, 800-1, 800-2, 900-1, and 900-2 are displayed in a 3D perspective view of the driving scene.

[0093] exist Figure 6-1 In scenario 600-1, traffic sign 602-1 is located on the right-hand side of the road on which vehicle 102 is traveling. Vehicle 102's common sense engine 124 determines that under normal circumstances, this sign is a yield sign. Next, the vehicle turns... Figure 6-2At a later point in time, when vehicle 102 travels along the same road for the second time, the common sense engine 124 of vehicle 102 expects to see the yield sign 602-1, but instead detects a different sign 602-2, such as a stop sign. The ROI of traffic sign 602-2 can be analyzed by the common sense engine 124 along with the scene context, for example, in the form of scene labels and detection labels. This analysis resolves the visual differences between traffic signs 602-1 and 602-2, and their representation or absence in map 114, while also providing the root cause. The common sense engine 124 can determine that a road change has occurred in response to identifying the change from sign 602-1 to 602-2 (yield to stop); the sign change constitutes a road change detection for updating map 114, and the common sense engine 124 can cause map 114 to change.

[0094] exist Figure 7-1 In scenario 700-1, vehicle 102 is traveling on road 702-1. The common sense engine 124 of vehicle 102 determines that under normal circumstances, road 702-1 has no sidewalk or shoulder. Next, we move to... Figure 7-2 At a later point in time, when vehicle 102 travels along the same road a second time, the common sense engine 124 of vehicle 102 expects to see features of road 702-1 that are absent (e.g., no shoulder, no sidewalk), but instead detects road 702-2 that includes a shoulder and a sidewalk. The ROI of road 702-2 can be analyzed by the common sense engine 124 along with scene context, such as scene labels and detection labels. This analysis resolves the visual differences between road 702-1 and road 702-2. The common sense engine 124 can determine that adding a shoulder and a sidewalk to one side of road 702-1 to create road 702-2 constitutes another road change detection for updating map 114, and the common sense engine 124 can change map 114.

[0095] exist Figure 8-1 In scenario 800-1, vehicle 102 is traveling on road 802-1. The common sense engine 124 of vehicle 102 determines that under normal circumstances, road 802-1 has no intersections. Next, [the text continues...] Figure 8-2At a later point in time, when vehicle 102 travels along the same road for the second time, the common sense engine 124 of vehicle 102 expects to see no intersection, but instead detects road 802-2, which includes an intersection leading to another street. The ROI of road 802-2 can be analyzed by the common sense engine 124 along with scene context, such as scene labels and detection labels. This analysis resolves the visual differences between road 802-1 and road 802-2. The common sense engine 124 can determine that the intersection in road 802-2 constitutes a third road change detection for updating map 114, and the common sense engine 124 can change map 114.

[0096] Now, unlike scenarios 600-1, 600-2, 700-1, 700-2, 800-1, and 800-2, Figure 9-1 and Figure 9-2 In scenarios 900-1 and 900-2, map 114 does not need to be updated because scenarios 900-1 and 900-2 display vegetation change detection instead of road change detection. Vehicle 102 travels on a road arranged with vegetation 902-1. The common sense engine 124 of vehicle 102 determines that under normal circumstances, the road has vegetation 902-1 consisting of fir trees on either side of the road, especially on either side closest to the road. Next, we move to... Figure 9-2 At a later point in time, when vehicle 102 travels along the same road a second time, the common sense engine 124 of vehicle 102 expects to see vegetation 902-1, but instead detects a different vegetation 902-2 arranged on that side of the road, which has fewer trees than vegetation 902-1. The ROI of vegetation 902-2 can be analyzed by the common sense engine 124 along with scene context, such as scene labels and detection labels. This analysis resolves the visual differences between vegetation 902-1 and vegetation 902-2. The common sense engine 124 can determine that vegetation 902-2 is not a change used to update map 114, but rather merely constitutes a vegetation change detection and avoids updating the vegetation change to map 114. In other words, Figure 9-1 and Figure 9-2 This is an example scenario used to identify areas of interest (e.g., everything except vegetation 902-1 and 902-2) and filter them out. In other scenarios, map 114 may include vegetation (e.g., for off-road navigation in a national park or uninhabited area), and in such cases, similar to road changes, vegetation changes can trigger updates to map 114.

[0097] Additional examples

[0098] The next section provides further examples of change detection criteria for updating sensor-based maps.

[0099] Example 1. A method comprising: receiving from a sensor device of a vehicle an indication that a registered object has been detected in a vicinity of the vehicle; having a processor of the vehicle determine, based on the indication, a difference between features of the registered object and features of a sensor-based reference map, the features of the sensor-based reference map including a map location corresponding to the coordinate location of the registered object; having the processor execute a machine learning model trained using self-supervised learning to identify change detections from inputs to the model, whether the difference satisfies a change detection criterion for updating the sensor-based reference map; in response to determining that the difference satisfies the change detection criterion, having the processor update the sensor-based reference map to reduce the difference; and having the processor operate the vehicle in an autonomous mode, the autonomous mode relying on the sensor-based reference map for navigating the vehicle in a vicinity of the coordinate location of the registered object.

[0100] Example 2. The method of Example 1, wherein the sensor device includes a radar device, and the sensor-based reference map includes a reference map derived at least in part from radar data.

[0101] Example 3. The method of Example 1 or 2, wherein the sensor device includes a lidar device, and the sensor-based reference map includes a reference map derived at least in part from point cloud data.

[0102] Example 4. A method in any of the preceding examples further includes: enabling a machine learning model to be trained using self-supervised learning by the processor generating multiple change detection criteria for determining whether to update the sensor-based reference map.

[0103] Example 5. The method of Example 4, wherein generating multiple change detection criteria for determining whether to update a sensor-based reference map includes: performing self-supervised learning based on training data, said training data comprising a nominal task expressed in natural language.

[0104] Example 6. The method of Example 4 or 5, wherein generating multiple change detection criteria for determining whether to update a sensor-based reference map includes: performing self-supervised learning based on training data, said training data further including sensor-based questions and answers.

[0105] Example 7. The method of Example 6, wherein the sensor-based questions and answers include questions and answers related to point cloud data, which indicates the 3D features of registered objects located at different map locations in the environment.

[0106] Example 8. The method of Example 1, where the map location includes a three-dimensional region of space, and the coordinate location of the registered object includes a three-dimensional coordinate location in space.

[0107] Example 9. A computer-readable storage medium comprising instructions, which, when executed, cause a processor of a vehicle to: receive from a sensor device of the vehicle an indication that a registered object has been detected in a vicinity of the vehicle; determine, based on the indication, a difference between features of the registered object and features of a sensor-based reference map, the features of which include a map position corresponding to the coordinate position of the registered object; execute a machine learning model trained using self-supervised learning to identify change detections from inputs to the model, whether the difference satisfies a change detection criterion for updating the sensor-based reference map; in response to determining that the difference satisfies the change detection criterion, cause the sensor-based reference map to be updated to reduce the difference; and cause the vehicle to operate in an autonomous mode, the autonomous mode relying on the sensor-based reference map for navigating the vehicle in a vicinity of the coordinate position of the registered object.

[0108] Example 10. A computer-readable storage medium of Example 9, wherein the sensor device includes a radar device, and the sensor-based reference map includes a reference map derived at least in part from radar data.

[0109] Example 11. The computer-readable storage medium of Example 9, wherein the sensor device includes a lidar device, and the sensor-based reference map includes a reference map derived at least in part from point cloud data.

[0110] Example 12. The computer-readable storage medium of Example 9, wherein the instructions, when executed, further enable the processor of a vehicle system to: enable a machine learning model to be trained using self-supervised learning by generating multiple change detection criteria for determining whether to update a sensor-based reference map.

[0111] Example 13. A computer-readable storage medium of Example 12, wherein instructions, when executed, cause a processor to: use self-supervised learning based on training data to generate multiple change detection criteria for determining whether to update a sensor-based reference map, the training data comprising a nominal task expressed in natural language.

[0112] Example 14. A computer-readable storage medium of Example 13, wherein instructions, when executed, cause a processor to: use self-supervised learning based on additional training data to generate multiple change detection criteria for determining whether to update a sensor-based reference map, the additional training data including sensor-based questions and answers.

[0113] Example 15. The computer-readable storage medium of Example 14, wherein sensor-based questions and answers include questions and answers related to point cloud data, which indicates the three-dimensional features of registered objects located at different map locations in the environment.

[0114] Example 16. The computer-readable storage medium of Example 9, wherein the map location includes a three-dimensional region of space, and the coordinate location of the registered object includes a three-dimensional coordinate location in space.

[0115] Example 17. A system comprising: a processor configured to: receive from a sensor device of a vehicle an indication that a registered object has been detected in a vicinity of the vehicle; determine, based on the indication, a difference between features of the registered object and features of a sensor-based reference map, the features of the sensor-based reference map including a map position corresponding to the coordinate position of the registered object; execute a machine learning model trained using self-supervised learning to identify change detections from inputs to the model, whether the difference satisfies a change detection criterion for updating the sensor-based reference map; in response to determining that the difference satisfies the change detection criterion, update the sensor-based reference map to reduce the difference; and enable the vehicle to operate in an autonomous mode, the autonomous mode relying on the sensor-based reference map for navigating the vehicle in a vicinity of the coordinate position of the registered object.

[0116] Example 18. The system of Example 17, wherein the sensor device includes a radar device, and the sensor-based reference map includes a reference map derived at least in part from radar data.

[0117] Example 19. The system of Example 17, wherein the sensor device includes a lidar device, and the sensor-based reference map includes a reference map derived at least in part from point cloud data.

[0118] Example 20. The system of Example 17, wherein the processor is further configured to enable a machine learning model to be trained using self-supervised learning by generating multiple change detection criteria for determining whether to update the sensor-based reference map.

[0119] Example 21. A system comprising means for performing any of the methods in the preceding examples.

[0120] in conclusion

[0121] While various embodiments of the present disclosure have been described in the foregoing description and illustrated in the accompanying drawings, it should be understood that the present disclosure is not limited thereto, but can be practiced in various ways within the scope of the following claims. It will be apparent from the foregoing description that various modifications can be made without departing from the spirit and scope of the present disclosure as defined by the following claims. The complexities and delays associated with updating reference maps, particularly when considering the detection of all possible changes, can be overcome by relying on the described change detection criteria, which, in addition to improving performance, also promotes driving safety.

[0122] Unless the context explicitly states otherwise, the use of "or" and grammatically related terms indicates an unrestricted, non-exclusive alternative. As used herein, the phrase referring to "at least one" of a list of items means any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, ab, ac, bc, and abc, as well as any combination with multiple identical elements (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other ordering of a, b, and c).

Claims

1. A method for use in a means of transport, comprising: Obtain a sensor-based reference map including map locations corresponding to the coordinate locations of at least one registered object, wherein the at least one registered object is an object detected by the sensor device of the vehicle; Sensor data is received from the sensor device of the vehicle, the sensor data including an indication that the registered object was detected in the vicinity of the vehicle at the coordinate location; The processor of the vehicle determines the difference between the characteristics of the registered object defined by the sensor data and the characteristics of the registered object defined by the sensor-based reference map; The processor executes a machine learning model to determine whether the difference meets multiple change detection criteria for determining whether to update the sensor-based reference map. The machine learning model is trained to create the multiple change detection criteria itself using self-supervised learning based on a natural language nominal task and point cloud-based inference. The machine learning model executes a first model to compute the correlation between the question and the response, and executes a second model to compute the similarity between two response options. In response to determining that the difference meets the plurality of change detection criteria, the processor updates the features of the registered object defined by the sensor-based reference map based on the features of the registered object defined by the sensor data; as well as The processor enables the vehicle to operate in an autonomous mode, which relies on the sensor-based reference map for navigating the vehicle in the vicinity of the registered object's coordinate location.

2. The method as described in claim 1, characterized in that, The sensor device includes a radar device, and the sensor-based reference map includes a reference map that is at least partially derived from radar data.

3. The method as described in claim 1, characterized in that, The sensor device includes a lidar device, and the sensor-based reference map includes a reference map that is at least partially derived from point cloud data.

4. The method as described in claim 1, characterized in that, The machine learning model was further trained to use sensor-based questions and answers.

5. The method as described in claim 4, characterized in that, The sensor-based questions and answers include questions and answers related to point cloud data, which indicates the three-dimensional features of registered objects located at different map locations in the environment.

6. The method as described in claim 1, characterized in that, The map location includes a three-dimensional area in space, and the coordinate location of the registered object includes a three-dimensional coordinate location in space.

7. A computer-readable storage medium storing instructions, which, when executed, cause a processor of a vehicle to: Obtain a sensor-based reference map including map locations corresponding to the coordinate locations of at least one registered object, wherein the at least one registered object is an object detected by the sensor device of the vehicle; Sensor data is received from the sensor device of the vehicle, the sensor data including an indication that the registered object was detected in the vicinity of the vehicle at the coordinate location; Determine the differences between the characteristics of the registered object defined by the sensor data and the characteristics of the registered object defined by the sensor-based reference map; A machine learning model is executed to determine whether the difference meets multiple change detection criteria for determining whether to update the sensor-based reference map. The machine learning model is trained to create the multiple change detection criteria itself using self-supervised learning based on a natural language nominal task and point cloud-based inference. The machine learning model executes a first model to compute the correlation between the question and the response, and executes a second model to compute the similarity between two response options. In response to determining that the difference meets the plurality of change detection criteria, the features of the registered object defined by the sensor-based reference map are updated based on the features of the registered object defined by the sensor data; and The vehicle is then operated in an autonomous mode, which relies on the sensor-based reference map for navigating the vehicle in the vicinity of the registered object's coordinate location.

8. The computer-readable storage medium as claimed in claim 7, characterized in that, The sensor device includes a radar device, and the sensor-based reference map includes a reference map that is at least partially derived from radar data.

9. The computer-readable storage medium as claimed in claim 7, characterized in that, The sensor device includes a lidar device, and the sensor-based reference map includes a reference map that is at least partially derived from point cloud data.

10. The computer-readable storage medium as claimed in claim 7, characterized in that, The machine learning model was further trained to use sensor-based questions and answers.

11. The computer-readable storage medium as claimed in claim 10, characterized in that, The sensor-based questions and answers include questions and answers related to point cloud data, which indicates the three-dimensional features of registered objects located at different map locations in the environment.

12. The computer-readable storage medium as claimed in claim 7, characterized in that, The map location includes a three-dimensional area in space, and the coordinate location of the registered object includes a three-dimensional coordinate location in space.

13. A system for a vehicle, the system comprising: Processor, the processor being configured to: Obtain a sensor-based reference map including map locations corresponding to the coordinate locations of at least one registered object, wherein the at least one registered object is an object detected by the sensor device of the vehicle; Sensor data is received from the sensor device of the vehicle, the sensor data including an indication that the registered object was detected in the vicinity of the vehicle at the coordinate location; Determine the differences between the characteristics of the registered object defined by the sensor data and the characteristics of the registered object defined by the sensor-based reference map; A machine learning model is executed to determine whether the difference meets multiple change detection criteria for determining whether to update the sensor-based reference map. The machine learning model is trained to create the multiple change detection criteria itself using self-supervised learning based on a natural language nominal task and point cloud-based inference. The machine learning model executes a first model to compute the correlation between the question and the response, and executes a second model to compute the similarity between two response options. In response to determining that the difference meets the plurality of change detection criteria, the features of the registered object defined by the sensor-based reference map are updated based on the features of the registered object defined by the sensor data; and The vehicle is then operated in an autonomous mode, which relies on the sensor-based reference map for navigating the vehicle in the vicinity of the registered object's coordinate location.

14. The system as described in claim 13, characterized in that, The sensor device includes a radar device, and the sensor-based reference map includes a reference map that is at least partially derived from radar data.

15. The system as described in claim 13, characterized in that, The sensor device includes a lidar device, and the sensor-based reference map includes a reference map that is at least partially derived from point cloud data.