Display enhancement

By combining video frame data and environmental models in autonomous vehicles and using user input to transmit information between different views, the problem of inconsistent environmental perception in autonomous vehicles is solved, the accuracy of operator decision-making and control is improved, and a more comprehensive understanding of the environment and vehicle navigation is achieved.

CN121753088APending Publication Date: 2026-03-27ZOOX INC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-13
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

During the environmental perception and navigation process of autonomous vehicles, the inconsistency between video frame data and perception output data makes it difficult for operators to fully understand the environment around the vehicle, affecting the accuracy of decision-making and control.

Method used

By combining video frame data with an environmental model, and utilizing user input to indicate features and transfer data between different views, the operator's understanding of the environment is enhanced. This includes displaying indications of missing features in the video frame data or highlighting features in the model, thereby achieving comprehensive information display and vehicle action commands.

Benefits of technology

It improves the operator's understanding of the autonomous vehicle environment, enhances the clarity of decision-making and the accuracy of vehicle control, and provides a more comprehensive environmental overview by comprehensively utilizing the advantages of video frame data and environmental models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753088A_ABST
    Figure CN121753088A_ABST
Patent Text Reader

Abstract

A method is provided that includes: (i) receiving, at a system, from an autonomous vehicle: (a) sensor data captured by sensors of the vehicle, and (b) output data generated based on environmental data captured by one or more sensors of the vehicle and used by a planning component to navigate in an environment, (ii) causing one or more displays to display at least one of: a representation of the sensor data or a model of the environment based on the output data, (iii) determining a location of a feature within the sensor data or the model, (iv) causing the one or more displays to display an indication of the feature at an orientation corresponding to the location, (v) receiving a user input at the system, and (vi) sending, by the system, data to the vehicle based on the user input to cause the vehicle to take an action.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] Autonomous and partially autonomous vehicles are being tested and used more and more frequently, not only for convenience but also to improve road safety. Autonomous vehicles can use combinations of different sensors that can be used to detect nearby objects to help the vehicle navigate its environment. Attached Figure Description

[0002] The accompanying drawings are used for detailed description. The same reference numerals are used in different drawings to indicate similar or identical components or features.

[0003] Figure 1 The system is illustrated in the example diagram, which includes a vehicle and a remote system for monitoring the vehicle. Figures 2 to 15 This is an illustration of an example of the output on one or more monitors at a remote system. Figure 16 A flowchart illustrating the method based on the example is provided; and Figure 17 This is a block diagram of an example vehicle system. Detailed Implementation

[0004] This application relates to systems, methods, and computer-readable media for improving the ability of a human operator (also known as a teleoperator) to understand the environment in which an autonomous vehicle is situated by combining one or more elements from a view presented to the operator (such as video captured by the vehicle) with one or more elements from another view presented to the operator (such as a model of the environment). This allows the operator to have a clearer overview of the environment, which in turn allows the operator to make more informed decisions, such as providing more informed instructions to remotely control the vehicle.

[0005] A human operator can monitor the autonomous vehicle and / or provide instructions, such as driving instructions, to the autonomous vehicle at certain times. In some scenarios, such as when the vehicle is unsure how to proceed, it may request instructions from the operator. The operator can be located remotely from the vehicle and can therefore monitor and / or control the vehicle from a remote system. To enable the operator to monitor / control the vehicle, data from the vehicle can be presented to the operator. For example, the vehicle may have one or more sensors (such as one or more video cameras) that capture sensor data (such as video frame data) and transmit the sensor data to a remote system. A representation of the sensor data (such as video frame data) can be displayed to the operator on one or more displays, allowing the operator to see the environment around the vehicle. In some examples, the representation of the video frame data can be shown to the operator as a view.

[0006] Video frame data can be human-interpretable. For example, video frame data can correspond to the view a human would see if they were in the same location as the video camera recording the video frame data. Other sensors on the vehicle can also provide human-interpretable data for operator viewing. For example, an infrared camera can also provide human-interpretable sensor data, even if this may differ from how a human would perceive the environment. Sensor data other than video frame data may be useful under specific weather conditions. For example, in foggy conditions, other sensor data may be more useful than video frame data. It should be understood that throughout this disclosure, any reference to “video data” or “video frame data” can be replaced by “sensor data,” where sensor data is recorded / captured by one or more sensors (such as one or more cameras) on the vehicle. Sensor data can include, for example, video frame data, or can include other sensor data that can be visually displayed to an operator. In the example, video frame data can be raw sensor data received from the vehicle, which may include sensor data processed by an ASIC (Application-Specific Integrated Circuit) or other processor of the sensor (e.g., an image signal processor) rather than further processed by the sensing components.

[0007] The vehicle can also generate output data for its own navigation in the environment (although in some cases, video frame data can also be used by the vehicle for navigation). For example, environmental data can be captured by one or more sensors (such as lidar devices, radar devices, etc.) and processed to provide output data, which is used by the vehicle's planning components for navigation in the environment. Therefore, the output data can indicate how the vehicle perceives its environment and can be based on the environmental data. The output data can also be sent to the system.

[0008] The output data (at least initially) may not be particularly human-interpretable; this could be the case, for example, with video frame data. However, the output data can be used by the vehicle and / or system to generate a model of the environment, such as a 2D or 3D rendered graphical model. Such a model can be human-interpretable, while the data used to generate the model may not be. The model can be displayed to the operator. In some examples, the display of the model can be shown to the operator as a view.

[0009] As will be explained in more detail below, the output data may include perception output data, which is generated by the vehicle's perception components. The perception components can detect features / objects in the environment surrounding the vehicle by combining environmental data from one or more sensors. The perception components can classify objects and determine other characteristics associated with the objects, such as position, velocity, etc. The output data may also include prediction output data, which is generated by the vehicle's prediction components. The prediction components can generate one or more probability maps representing the predicted probabilities of the possible locations of one or more objects in the environment. Therefore, the output data can be used to generate models and thereby represent objects in the environment.

[0010] When displaying a model and video frame data for operator viewing, the operator can more easily perceive information in the video frame data representation compared to the model. Video cameras can have different fields of view compared to the sensors used to generate the output data (which is used to generate the model). Therefore, features / objects may be visible in the video frame data representation but not in the model. Conversely, considering that the output data can be based on environmental data collected by several different types of sensors (as opposed to video frame data which can be captured by one or more video cameras), the model can provide a more complete overview of the environment around the vehicle. Therefore, in some cases, the operator can more easily perceive information in the model compared to the video frame data representation. For example, in the video frame data, features / objects may be occluded by other features / objects (and therefore may not be visible to the operator), but those features may be visible in the model.

[0011] Given that different information can be gathered from two different "views" presented to the operator, it can be useful to combine information from video frame data and the model into a single view. For example, if a feature is visible in the model but not in the video frame data, an indication of that feature can be provided / displayed in the representation of the video frame data (or vice versa). Thus, a rich view of the environment can be provided to the operator within a single view without having to refer to two views, and / or the differences between the two views can be visualized.

[0012] As will become apparent, this “enhancement” of features from different views can be done automatically (such as when the system detects that a feature is missing from one view or is more difficult to see), or it can be initiated by the operator based on user input. For example, the operator can provide user input at a specific location in one view (such as in the model view presented to the operator), which causes the system to indicate the corresponding location associated with the user input in another view.

[0013] As an example, an operator can provide user input (such as "click," "press," or "drag") on a specific feature in the video view, which causes an indication of the feature to appear in the model view. The user can then see where to find the same feature in the model view, which can help the operator plan the vehicle's navigation path. Indicators of features in another view can be markers, such as dots, stars, crosses, etc. For example, if the operator clicks on a pedestrian in the video view, a marker can appear on top of the same pedestrian depicted in the model view. If a pedestrian is not depicted in the model view (perhaps due to missing data), a marker can appear in the corresponding location where the pedestrian should be. In this example, if the feature already exists in the model view, it can be indicated in a way that makes it more visually apparent to the operator. For example, if the operator clicks on a vehicle in the video view, the model of the same vehicle in the model view can be highlighted, such as by changing or emphasizing its color, or by adding a boundary, such as a circle, around the vehicle. In another example, an indication of a feature in another view can be an image of the feature. For example, if an operator clicks on a vehicle in the video view, an image of the vehicle (or a general vehicle or a graphical model of the vehicle) can be generated in the model view. Images of features can be added when a feature is not fully or partially present in another view.

[0014] In an example where a feature / object is visible to the operator in the video view but is missing from the output data (and therefore the same feature does not exist in the model view), additional output data (such as perception data) associated with the feature can be generated and sent to the vehicle for use by the vehicle's planning components to navigate in the environment. The additional output data may include data associated with the feature (such as the classification of the feature) and the feature's position within the environment, based on the feature's orientation within the view. The position can be within the coordinate system of the output data. Therefore, additional output data can be provided to the vehicle, which can improve the vehicle's perception of the environment. Thus, the vehicle is informed of additional features it should be aware of in its environment.

[0015] As another example, an operator can provide user input in one view to modify or generate a vehicle's navigation path. Therefore, features (i.e., the modified or generated navigation path) can be modified or generated in one view, and indications of those features can be displayed in another view. User input can include one or more clicks, touches, drags, etc., to modify or generate the navigation path. For example, an operator can draw the vehicle's navigation path in a model view and overlay it in a video view (or vice versa). This shows the operator where the vehicle should navigate in the video view.

[0016] As another example, an operator can provide user input (such as clicking on a visible object in one view) that causes an indication of a predicted path associated with the object to be displayed in another view. For instance, an operator can click on a pedestrian in a model view, which causes the predicted path associated with the pedestrian to be overlaid in a video view (or vice versa).

[0017] In another example, the operator can provide user input to provide additional information associated with the features. This additional information provided by the operator can be sent to the vehicle, and the vehicle's planning components can use this additional information to navigate the environment.

[0018] As an example, a vehicle may incorrectly or fail to classify objects / features, or may classify objects with low confidence. The classification of objects in the environment surrounding the vehicle can be performed by the vehicle's perception components, and the classification of objects can be sent as part of the output data to a remote system. Therefore, in this scenario, an operator can provide user input to classify or reclassify a specific feature / object. The result of the classification or reclassification is then sent to the vehicle, allowing it to navigate appropriately. For example, a feature / object may be misclassified in the model, and the user can provide a first user input (such as a click) on a feature in the video view, causing an indication to be displayed for the feature in the model view (or vice versa). A second user input from the user can provide a classification for the feature / object, and the classification is then sent to the vehicle.

[0019] As another example, an operator can provide user input to deliver driving instructions for the vehicle. For instance, a first user input can select a specific feature / object, and a second user input can provide driving instructions associated with that feature / object, such as "avoid," "stop," "wait," "ignore," "continue," etc. As an example, there might be litter on the road. In some cases, the vehicle might be unsure how to classify the object, or might classify it incorrectly. Therefore, the operator can click on the litter and instruct the vehicle to ignore or avoid it. As mentioned earlier, clicking on an object in one view (such as a video view) can cause an indication of the object to appear in another view (such as a model view), which informs the operator that they are clicking on the correct object, thus associating the driving instruction with the same object perceived by the vehicle.

[0020] Therefore, in the examples described herein, a system is provided comprising one or more processors and one or more non-transitory computer-readable media having instructions stored thereon, which, when executed by the one or more processors (or the system), cause the one or more processors (or the system) to perform operations including: (i) receiving from an autonomous vehicle at the system: (a) video frame data captured by a video camera of the vehicle, and (b) perception output data generated by the vehicle, wherein the perception output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by a planning component of the vehicle for navigation in the environment, wherein the video frame data and the perception output data indicate the environment; (ii) causing a first area of ​​one or more displays to display a representation of the video frame data; (iii) The system causes a second area of ​​the one or more displays to display a model of the environment based on the perceived output data; (iv) receives user input at the system, the user input being associated with a feature within one of the first or second areas; (v) determines a first orientation associated with the feature within the first or second area; (vi) determines a second orientation of the feature within the other of the first or second area based at least in part on the first orientation; (vii) causes the one or more displays to display an indication of the feature at the second orientation within the other of the first or second area, based at least in part on whether the first orientation is within the first or second area of ​​the one or more displays; and (viii) sends data based at least in part on the user input to the vehicle to cause the vehicle to take action.

[0021] The system can be a remote system, such as a remote system used by an operator to monitor and / or control the vehicle. The system can be communicatively coupled to the vehicle, for example, via one or more wired and / or wireless networks.

[0022] Therefore, this disclosure relates to receiving user input associated with a feature in one “view” (i.e., a first region or a second region) and responsively causing an indication of the feature to be displayed in another view (i.e., in the other of the first or second regions). This can be achieved by determining where the user input is received in the view (i.e., a first orientation), determining the location in the environment corresponding to that orientation (such as physical coordinates), and then determining where the feature should be indicated in the other view (i.e., a second orientation). Thus, the first orientation is an orientation on one or more displays (within the first or second region), and the second orientation is another orientation on one or more displays (within the other of the first or second region). For example, the first and second orientations may correspond to x, y pixel coordinates on one or more displays. Both the first and second orientations correspond to the same physical location in the environment (because they are associated with the same feature / object). In some cases, video frame data, as well as model and output data (such as perceptual output data), may be associated with a coordinate system. In the example, the first and second orientations (in 2D space) can be converted into 3D locations in the environment, as explained in more detail below.

[0023] In the example, the feature could be an environmental characteristic, such as an agent, vehicle, pedestrian, or bicycle. Therefore, the feature may already be displayed when user input is received and associated with it. In other examples, the feature could be a feature generated based on user input, such as a vehicle's navigation path. Therefore, the feature may not be displayed when user input is received, but it is still associated with the user input.

[0024] User input may include a single input or multiple inputs. Multiple inputs may be received by the system at different times. For example, the first user input may correspond to the user selecting a direction within one of a first or second area of ​​the display, and the second user input may correspond to a driving instruction or classification, wherein the data sent by the system to the vehicle includes the driving instruction or classification.

[0025] As briefly mentioned above, in some examples, instead of requiring user input in one view to indicate features to be displayed in another, features associated with one view can be automatically displayed in another view. For example, output data may include data associated with roads, sidewalks, intersections, traffic lights, such as lane markings or boundaries, which can be overlaid on the video view. Data associated with roads, sidewalks, intersections, traffic lights, etc., can be referred to as map data, which contains location information associated with one or more map features (such as roads). Therefore, output data may include map data that can be automatically (or upon user request) incorporated into the video view. In the example, map features may be visible in the model by default. In another example, output data may include vehicle route data, such as navigation paths or "passages," where a passage defines a safe area that a vehicle can navigate through, and route data can be overlaid on the video view. For example, route data may be generated by a vehicle's planning or prediction component. Route data can again be associated with location information. For example, a passage or navigation path may be defined by one or more points, each associated with a location in the environment. Therefore, route data can be automatically (or upon user request) incorporated into the video view. In the example, route data may be visible in the model by default. In the example, at least a portion of the route data is determined by the system (and therefore not necessarily received from the vehicle, although route data (such as routes) can be determined by the system using data received from the vehicle). In other examples, route data is received by the system from the vehicle.

[0026] As another example, if it is determined that a feature exists only in one of the video frame data or the output data (so the feature is visible in only one view, such as the feature being missing from the model view or hidden / occluded in the video view), then an indication of the feature can be automatically displayed in the view where the feature does not exist.

[0027] As another example, if the environment (such as output data) is detected to contain a priority agent / object (such as an emergency vehicle), an indication of the priority agent can be displayed in the video view (or vice versa). This can be useful for highlighting the priority agent to the operator when it is not visible or difficult to see within a view.

[0028] As another example, colors within the video view can be incorporated into the model view.

[0029] In this example, it's not necessarily necessary to display both views simultaneously (i.e., a representation and a model of the video frame data). For instance, an operator could provide user input in one view and then switch to another view displaying indications of the features. As another example, where no user input is required to display indications of the features (i.e., when this happens automatically, as mentioned above), it may not be necessary to display both views simultaneously. Similarly, as previously mentioned, a more general representation of sensor data could be displayed instead of a representation of the video frame data.

[0030] Therefore, in the examples described herein, a system is provided comprising one or more processors and one or more non-transitory computer-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform operations including: (i) receiving from an autonomous vehicle at the system: (a) sensor data captured by sensors of the vehicle, and (b) output data generated by the vehicle, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by a planning component of the vehicle for navigation in the environment, wherein the sensor data and the output data indicate the environment; (ii) causing one or more displays to display at least one of: (a) a representation of the sensor data or (b) a model of the environment based on the output data; (iii) determining the location of a feature within the sensor data or the model; (iv) causing the one or more displays to display an indication of the feature at an orientation corresponding to the location on the one or more displays; (v) receiving user input at the system; and (vi) sending data based on the user input from the system to the vehicle to cause the vehicle to take action. In the example, if both a representation of the sensor data and a model are displayed, features can be indicated in both. In other examples, features can be indicated in only one view. For example, if the location of a feature is determined in the model, the feature can be displayed in the representation of the sensor data, and if the location of a feature is determined in the sensor data, the feature can be displayed in the model. In the example, the sensor data is video frame data, such as video frame data captured by a video camera. In other examples, the sensor data is captured by another sensor instead of a video camera.

[0031] As mentioned, in the example, the location of features within sensor data (such as video frame data) or the model can be determined automatically without requiring user input to select or generate features within the displayed sensor data view or model view. Therefore, user input can be received after the indication of the displayed feature. In some cases, user input is received before the indication of the displayed feature, but the user input does not necessarily have to be user input that selects or generates features within the displayed sensor data view or model view. In some cases, user input is received before the indication of the displayed feature, and in others, user input is received after the indication of the displayed feature.

[0032] However, in other examples, user input is associated with a feature and a first orientation on one or more displays. For example, as discussed above, user input may include a first user input at a first orientation on one or more displays (such as within a displayed representation of sensor data or within a displayed model). The first orientation may correspond to the location of the feature within the sensor data or model (depending on where the first orientation is located), and based on the location, a corresponding orientation within the other of the displayed representation of sensor data or the displayed model can be determined, as previously discussed. Once the orientation has been determined, an indication of the feature can be displayed at the orientation on one or more displays, where the orientation corresponds to the determined location, and ultimately based on the first orientation indicated by the first user input from the operator.

[0033] More detailed examples of the systems, methods, and computer-readable media of this disclosure will now be presented with reference to the accompanying drawings.

[0034] Figure 1 An example implementation of this disclosure is depicted. As shown, an autonomous vehicle 110 is located within an environment. In this example, the vehicle 110 navigates throughout the environment at a specific speed; however, in other examples, the vehicle 110 may be stationary. The vehicle includes one or more sensors 112, which in this example include at least one image sensor, such as a video camera 112, that captures images / frames of the environment and (at least temporarily) stores the recorded frames as video frame data. Other sensors 112 may capture additional environmental data associated with the environment surrounding the vehicle 110. For example, other sensors 112 may include ultrasonic sensors, lidar sensors, radar sensors, etc., for acoustically detecting objects in the surrounding environment.

[0035] Figure 1View 116 depicts the environment as seen by an image sensor (such as video camera 112). For example, within the field of view 116, there may be one or more moving (i.e., non-stationary) objects, such as one or more pedestrians, vehicles or cyclists, and one or more stationary objects, such as one or more parked or stopped vehicles, buildings, signs, etc.

[0036] Remote system 100 can be communicatively coupled to vehicle 110 via one or more networks 114. Video frame data captured by one or more video cameras 112 can be transmitted to remote system 100 via network 114. Therefore, vehicle 110 may include a network interface to enable data transmission to remote system 100. In this example, the network interface includes a wireless antenna for sending video frame data (and any other data, such as output data) to remote system 100. Human operator 122 can monitor and / or control vehicle 110 via system 100.

[0037] Video frame data can include data associated with frames of the recorded video and indicate the environment around the vehicle. Video frame data can also include data corresponding to pixels within the frame.

[0038] Although video frame data is discussed in the following examples, it should be understood that the following discussion can generally be applied to sensor data in the same way.

[0039] Environmental data captured by one or more sensors 112 can be processed by vehicle 110 (such as by the perception components of vehicle 110) to generate perception output data. The perception output data can be used by the vehicle's planning components (in... Figure 17 (Discussed in more detail below), the planning component determines instructions for controlling the operation of vehicle 110 based at least in part on the sensing output data. The sensing output data may be sent to or otherwise transmitted to remote system 100 via network 114. The sensing output data may be transmitted to system 100 as output data, wherein the output data includes at least the sensing output data.

[0040] In some cases, environmental data captured by one or more sensors 112 may not be processed by vehicle 110 to generate output data, and may instead be transmitted to remote system 100. System 100 itself may process the environmental data to generate output data and / or a model of the environment, in some cases, in the same or similar manner as the vehicle would generate perception, prediction, and / or planning data.

[0041] Vehicle 110 can also determine or store prediction output data, wherein the prediction output data is generated by the prediction component of vehicle 110 (in Figure 17(Discussed in more detail below). The prediction component can generate one or more probability maps representing the predicted probabilities of the possible locations of one or more features / objects in the environment. The prediction output data can be sent to or otherwise transmitted to the remote system 100 via network 114. The prediction output data can be transmitted to system 100 as part of the output data.

[0042] Vehicle 110 can also determine or store map data, wherein the map data contains location information associated with one or more map features, such as roads, sidewalks, traffic lights, intersections, etc. The map data can be sent to or otherwise transmitted to remote system 100 via network 114. The map data can be transmitted to system 100 as part of the output data. In the example, at least some of the map data is accessible to system 100 without the need for vehicle 110 to send the map data (or in addition to vehicle 110 sending the map data). Vehicle 110 can also determine route data, wherein the route data contains location information associated with one or more route features, such as navigation paths or routes. The route data can be sent to or otherwise transmitted to remote system 100 via network 114. The route data can be transmitted to system 100 as part of the output data.

[0043] System 100 includes one or more processors 102 and one or more non-transitory computer-readable media 104 having instructions stored thereon that, when executed by the one or more processors 102, cause the one or more processors 102 to perform a specific operation, which will be discussed in more detail below.

[0044] Once system 100 receives the output data, system 100 can use the output data to generate an environment model, such as a graphical model.

[0045] System 100 may also include at least one display 106 (also referred to as a computer monitor) for displaying information to operator 122. System 100 may also include one or more input devices 120 for receiving user input from operator 122. For example, user input may cause instructions to be transmitted via network 114 to vehicle 110, such as following a specific navigation path in the environment. In this example, one or more displays 106 are themselves input devices 120. For example, one or more displays 106 may include a touchscreen display 106 that can accept user input associated with a specific orientation on the display 106.

[0046] Information displayed on one or more displays 106 or otherwise output may include a representation 108 of video frame data. Figure 1A representation 108 of video frame data received from vehicle 110 is depicted on one or more displays 106. Thus, one or more displays 106 can output a representation 108 of a view 116 of the environment seen by video camera 112.

[0047] Figure 1 A model 118 based on output data received from vehicle 110 is also depicted for the display environment of one or more displays 106. In an example, one or more displays 106 may display a representation 108 of video frame data while simultaneously displaying model 118. For example, as Figure 1 As shown, one display 106 can display a representation 108 of video frame data, and another display 106 can display a model 118. In other cases, both the representation 108 of video frame data and the model 118 can be displayed on the same display 106.

[0048] In other examples, one or more displays 106 may display representations 108 and models of video frame data at different times (and the operator may select between them as needed). In other examples, one or more displays 106 may display only one of model 118 or representations 108 of video frame data.

[0049] although Figure 1 Model 118 is shown as a 2D representation of the environment, but it should be understood that model 118 can be a 3D representation of the environment.

[0050] Typically, one or more displays 106 may display a representation 108 of video frame data in a specific area of ​​the display 106 (referred to herein as the “first area”), and one or more displays 106 may display a model 118 in another area of ​​the display 106 (referred to herein as the “second area”).

[0051] In this example, when vehicle 110 includes two or more video cameras 112 and video frame data from the two or more video cameras 112 is transmitted to system 100, one or more displays 106 can display representations of the video frame data from the two or more video cameras 112. For example, two or more representations of the video frame data can be displayed simultaneously on one or more displays 106, or one or more displays 106 can display representations at different times (and the operator can select between two or more representations of the video frame data as needed). In this example, one or more displays 106 can display representations of video frame data based on the video frame data captured by the two or more video cameras 112. For example, video data captured by two or more video cameras 112 can be combined / stitched together to provide a combined representation of the video frame data.

[0052] As previously discussed, when one or more displays 106 are displaying at least one of model 118 or representation 108 of video frame data, operator 122 may provide user input, such as via one or more user input devices 120. For example, a user may “select” a feature displayed in a first area (i.e., within the representation 108 of video frame data) or a second area (i.e., within model 118), or provide user input to modify or generate features, such as navigation paths. User input may be associated with a specific orientation (such as x, y coordinates on the display within the area).

[0053] Figures 2 to 15 A close-up of the output on one or more displays 106 viewed by operator 122 is depicted. It should be understood that the representation 108 of video frame data may be displayed on one display 106, and the model 118 may be displayed on another display 106 (at the same time or at different times), or both the model 118 and the representation 108 of video frame data may be displayed on the same display 106 (at the same time or at different times).

[0054] In any of the following examples, it should be understood that a reference to user input received in one view (i.e., within representation 108 or model 118 of video frame data) can be equally applied to examples where user input is received in another view. Similarly, a reference to an indication of a feature displayed in one view can be equally applied to examples where the indication can be displayed in another view. Therefore, the specific examples discussed herein can be reversed.

[0055] Figure 2 A close-up depiction of the output on one or more displays 106, viewed by operator 122, according to an example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 2 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0056] like Figure 2As shown, operator 122 can provide user input (such as first user input) via user input device 120 by moving and clicking pointer 202 (such as a mouse pointer) within a first area occupied by representation 108 of video frame data on one or more displays 106 or within a second area occupied by model 118 on one or more displays 106. In other examples, for example, pointer 202 may not be present, and user input may be received in another manner, such as via touch input on one of the displays 106. In this particular example, user input is provided within a first area occupied by representation 108 of video frame data on one or more displays 106. For example, operator can “click” or “touch” at a specific location on one or more displays 106 within the first or second area (i.e., within the displayed representation 108 or model 118 of video frame data). As mentioned, user input can be associated with features already displayed in the specific area (i.e., within the representation 108 or model 118 of video frame data). The location determined by user input on one or more displays 106 corresponds to a physical location within the environment. In this example, the features associated with the user input correspond to the location of a feature behind a vehicle on the road.

[0057] In the example, a first area occupied by a representation 108 of video frame data on one or more displays 106 can be a first area on one or more displays 106 having the same size as the representation 108 of video frame data. For example, if the representation 108 of video frame data is displayed on the display and the representation 108 of video frame data has a size of 500 × 300 pixels on the display, then the first area has a size corresponding to 500 × 300 pixels. Therefore, the first area on the display is occupied by the representation 108 of video frame data (i.e., the pixels of the first area display the representation 108 of video frame data). Other areas of one or more displays 106 can display other information, such as model 118. Similarly, a second area occupied by model 118 on one or more displays 106 can be a second area on one or more displays 106 having the same size as the displayed model 118. For example, if model 118 is displayed on the display and model 118 has a size of 500 × 300 pixels on the display, then the second area has a size corresponding to 500 × 300 pixels. Therefore, the second area on the display is occupied by model 118 (i.e., the pixels of the second area display model 118). Other areas of one or more displays 106 may display other information, such as a representation 108 of video frame data. In the example, the representation 108 of the video frame data and model 118 have different viewpoints. In other cases, they may be displayed as having the same viewpoint. In the example, both the representation 108 of the video frame data and model 118 display the same object / feature. In some examples, the field of view can allow one view to display a subset of features displayed in another view. For example, a feature may be outside the camera's field of view but can still be displayed in model 118.

[0058] After receiving user input associated with a feature, one or more displays 106 can display an indication of the feature at a specific orientation corresponding to the same physical location in the environment within another region (i.e., within the representation 108 of the video frame data or the other in model 118). Therefore, if the user input is received at a first orientation within a first region (i.e., within the representation 108 of the video frame data) or a second region (i.e., within model 118), a second corresponding orientation within the other region can be determined. In the example, this can be achieved by: (i) determining a first orientation (depending on where the user input was received) associated with the feature indicated by the user input within either the first or second region, and (ii) determining a second orientation of the feature within the other region based on the first orientation. For example, each orientation in the first region can be mapped to / converted to a corresponding orientation in the second region, where the corresponding orientation corresponds to the same physical location in the environment. A lookup table can be used to convert orientations between the two regions.

[0059] In the example, determining the second orientation can be achieved by: (i) determining a first orientation associated with the feature within a first or second region (depending on where user input is received), (ii) determining a location associated with the feature within the environment based on the first orientation, and (iii) determining a second orientation of the feature within the other of the first or second region based on the location. For example, each orientation in the first and second regions can be mapped to a corresponding physical location within the environment. The location within the environment can be associated with a coordinate system. The location within the environment can alternatively be referred to as the location within the video frame data or the model (where the video frame data and the feature within the model are associated with the location).

[0060] The mapping or transformation of orientation can be based on a calibration procedure that takes into account specific parameters of the video camera on vehicle 110 (such as lens distortion of the video camera), as well as the physical position and orientation of the video camera in 3D space. In some cases, one or more equations can be derived during the calibration of features, where a first orientation is input to a second orientation in another region on display 106 (or vice versa). In the example, determining the second orientation of a feature within the other of the first or second region based on the first orientation can be achieved by configuring a “virtual” camera in model 118. For example, a virtual camera can have the same parameters as a “real” video camera (such as field of view), position in 3D space, orientation, etc. Therefore, whenever input is received in the first or second region, the orientation can be transformed into a corresponding orientation in the other region, because each three-dimensional position in the environment (and therefore in model 118) is mapped to a corresponding x, y orientation in the representation of the video frame data.

[0061] Therefore, in the example, each “pixel” in the first region is mapped to a corresponding location in 3D space (i.e., a location within the model). Once the location within the model is known, an indicator can be displayed at the second location corresponding to that location.

[0062] In an example where user input is received in a first region, determining the second orientation of a feature within a second region based on the first orientation can be achieved by determining the position of the corresponding feature / object in model 118 and determining the second orientation based on the position of the corresponding feature. Determining the position of the corresponding feature may include: (a) determining a vector within the field of view of the video camera, where the vector extends from the video camera (such as the center of the field of view) to the first orientation; (b) determining a feature within the model along the vector (such as coinciding with the vector); and (c) determining the position of that feature. Once the position of the feature is known, the second orientation of the feature within the second region can be determined and thus indicated. The vector can be determined in 3D space (in the model) based on the known parameters of the video camera mentioned above. For example, if a user “clicks” a location on a road in the first region, a vector (e.g., a ray) extending from the “virtual” camera toward the location in space represented by that pixel in the first region can be determined, and the feature intersecting the vector (in this case, the road surface) is determined to be at the corresponding location clicked by the user. For example, the depth of a feature (such as a road surface) can be determined by a sensor system using sensor fusion and / or using a depth sensor (e.g., LiDAR, radar, ToF). In some examples, depth data can be directly fused from such sensors into the image space and used to determine the corresponding location in space for a user-selected feature using depth interpolation, segmentation, and / or other techniques.

[0063] In the example where user input is received in the second region, determining the second orientation of a feature within the first region based on the first orientation can be achieved by: (a) determining the position in 3D space (i.e., in the model) based on the first orientation, and determining the second orientation based on the position. This process can again utilize the calibration procedure or virtual camera discussed above.

[0064] In other examples, determining the second location of a feature within either a first or second region based on the first location can be achieved by identifying the object closest to the first location and determining the same corresponding object in the other region, where the corresponding object is displayed at the second location. For example, if the first location is on or near a car in a video frame view, the corresponding car could be located within the model.

[0065] In another example, video frame data can be associated with a first coordinate system, and the model (and therefore the output data) can be associated with a second coordinate system. Therefore, determining the second orientation can be achieved by: (i) determining a first orientation associated with a feature within a first or second region (depending on where user input is received), (ii) determining a first position of the feature within the first or second coordinate system based on the first orientation, and (iii) determining a second position of the feature within the other of the first or second coordinate system by transforming between the first and second coordinate systems, and (iv) determining the second orientation based on the second position. The position within the first coordinate system can alternatively be referred to as the position within the video frame data. Similarly, the position within the second coordinate system can alternatively be referred to as the position within the model.

[0066] After the second orientation has been determined, one or more displays 106 can display a feature indication 204 at the second orientation. Figure 2 In the example, because the user input is received in the first region (i.e., within the representation 108 of the video frame data), the second orientation corresponds to the orientation in the second region (i.e., within model 118). Therefore, Figure 2 The feature indicator 204 is shown, in this example, as an "x"-shaped mark displayed in a second position within model 118. Therefore, indicator 204 is displayed at a location corresponding to the same feature in the environment, in this example, behind a vehicle on the road. Thus, the operator 122 viewing model 118 can see where the same feature is located within model 118. Therefore, the operator 112 can more easily correlate corresponding features between the two views 108, 118. In some examples, in addition to displaying indicator 204, the area on the display in which indicator 204 is displayed can be adjusted. For example, the representation of model 118 or video frame data can be adjusted to "zoom in" on the feature (or more generally, the field of view can be adjusted to focus on the feature) to depict the feature more clearly to the operator.

[0067] Data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with, or before one or more displays 106 display an indication 204 of a feature in a second orientation. The data may be based on user input and cause vehicle 110 to take actions, such as moving or behaving in a particular manner. Data received from the system can be provided to the vehicle's planning components for processing. For example, operator 122 can provide driving instructions to vehicle 110, classifying features / objects and providing confirmation to vehicle 110 if vehicle 110 has already confirmed the operation request (such as "proceed?", "OK togo?", "is this feature / object a pedestrian?", etc.). In this example, data sent from system 100 to vehicle 110 may include the feature's location within the environment (where the location may be determined based on a first orientation), or may include other data associated with the feature to allow vehicle 110 to determine the location of the feature itself.

[0068] Figure 3 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 3 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0069] like Figure 3As shown, operator 122 can provide user input (such as first user input) via user input device 120 by moving and clicking pointer 202 (such as a mouse pointer) within a first area occupied by representation 108 of video frame data on one or more displays 106 or within a second area occupied by model 118 on one or more displays 106. In other examples, for example, pointer 202 may not be present, and user input may be received in another manner, such as via touch input on one of the displays 106. In this particular example, user input is provided within a first area occupied by representation 108 of video frame data on one or more displays 106. For example, operator can “click” or “touch” at a specific location on one or more displays 106. As mentioned, user input can be associated with features already displayed in a specific area (i.e., within representation 108 of video frame data or model 118). The location determined by user input on one or more displays 106 corresponds to a physical location within the environment. In this example, the feature associated with user input corresponds to a pedestrian.

[0070] After receiving user input associated with a feature, one or more displays 106 can display an indication of the feature at a specific orientation corresponding to the same physical location in the environment in another area (i.e., within the representation 108 of the video frame data or the other of model 118). Therefore, if the user input is received at a first orientation within a first area (i.e., within the representation 108 of the video frame data) or a second area (i.e., within model 118), a second corresponding orientation within the other area can be determined. Regarding Figure 2 Examples of how this can be achieved have been discussed, so for the sake of brevity, they will not be repeated here.

[0071] After the second orientation has been determined, one or more displays 106 can display a feature indication 204 at the second orientation. Figure 3 In the example, because the user input is received in the first region (i.e., within the representation 108 of the video frame data), the second orientation corresponds to the orientation in the second region (i.e., within model 118). Therefore, Figure 3 The indicator 204 is shown, in this example, as a mark in the form of a circle surrounding the corresponding feature displayed within model 118. Therefore, indicator 204 is displayed at a location corresponding to the same feature in the environment.

[0072] It should be understood that an indicator displaying a feature in a second location does not necessarily require the indicator to occupy a pixel exactly corresponding to the second location; rather, it can generally indicate the second location. For example, in this example, an indicator 204 in the form of a circle can generally indicate the second location, perhaps by centering on the second location. Other examples are possible.

[0073] Data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with, or before the indication 204 of a feature is displayed on one or more displays 106 in a second orientation. The data may be based on user input and cause vehicle 110 to take action, such as moving or behaving in a particular manner. Data received from the system can be provided to the vehicle's planning components for processing. For example, operator 122 may provide driving instructions to vehicle 110, or may classify the feature / object, or may provide confirmation to vehicle 110 if vehicle 110 has already acknowledged an operation request (such as "proceed?", "OK to go?", "Is this feature / object apedestrian?", etc.). In this example, the data sent from system 100 to vehicle 110 may include the feature's location within the environment (where the location may be determined based on a first orientation), or may include other data associated with the feature to allow vehicle 110 to determine the location of the feature itself. An example of a command to be sent to a vehicle can be found in patent application number 17 / 463,431, filed on August 31, 2021, entitled “COLLABORATIVE ACTION AMBIGUITY RESOLUTION FOR AUTONOMOUS VEHICLES”, the entire contents of which are hereby incorporated herein by reference for all purposes.

[0074] Figure 4 Depicting and Figure 3 A similar example exists, but instead of operator 122 providing user input (such as a first user input) within a first area of ​​one or more displays 106 occupied by a representation 108 of video frame data, the user input is received within a second area of ​​one or more displays 106 occupied by a model 118. Therefore, the feature indication 204 is displayed in the first area rather than the second area, as... Figure 3 As shown.

[0075] In some examples, in addition to or instead of indicator 204, indicator 208 may be displayed showing the predicted path associated with the feature / object. For example, operator 122 may click on a pedestrian in a second region, which causes the predicted path associated with the pedestrian to be overlaid in the first region (or vice versa). Thus, the first user input corresponds to the selection of an object visible in either representation 108 or model 108 of the video frame data, and the feature thus corresponds to the predicted path associated with the object in the environment.

[0076] In any of the examples discussed above or in the examples discussed herein, operator 122 may provide a second user input to provide driving instructions for the vehicle, wherein the first user input corresponds to the selection of a feature, as discussed above. Therefore, receiving user input may include receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of a first or second area of ​​one or more displays 106, and the second user input corresponding to driving instructions for the vehicle. Thus, data sent from system 100 to vehicle 110 may include driving instructions for the vehicle to navigate in its environment. In the example, data sent from system 100 to vehicle 110 may also include the location of the feature within the environment (where the location may be determined based on the first orientation). In this way, vehicle 110 may apply driving instructions based on location. For example, the vehicle may stop before, ignore, or bypass the location. In some cases, vehicle 110 may apply driving instructions without needing to receive the location of the feature. For example, vehicle 110 may apply instructions based on its own location information associated with the feature, or may apply instructions immediately without using location information.

[0077] Figure 5 An example is depicted showing how operator 122 can continue to provide driving instructions for the vehicle after the initial user input has been provided. For example, immediately following... Figure 4 The system can display prompt 206, such as a graphical menu. In some cases, prompt 206 can be displayed in response to receiving the first user input (i.e., selection of a feature, in this case, a pedestrian). In other cases, prompt 206 can be displayed in response to receiving additional user input.

[0078] like Figure 5 As shown, after prompt 206 has been displayed, operator 122 can provide user input (such as a second user input) via user input device 120 by moving pointer 202 and selecting a driving instruction. In this example, operator 122 selects one of the options displayed in prompt 206, thereby selecting a driving instruction. In this example, the driving instruction corresponds to "stop".

[0079] In other cases, no prompts are displayed, and user input indicating driving instructions may be provided in another manner, such as by verbal or typed commands. Other examples are possible. U.S. Patent Application Serial No. 16 / 852,116, filed April 17, 2020, entitled “Teleoperations For Collaborative Vehicle Guidance,” describes a vehicle that determines its navigation trajectory in an environment based on driving instructions received from an operator; the entire contents of that U.S. patent application are incorporated herein by reference for all purposes.

[0080] Figure 6 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 6 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0081] like Figure 6As shown, operator 122 can provide user input (such as one or more user inputs) via user input device 120 by moving and clicking pointer 202 (such as a mouse pointer) within a first area occupied by representation 108 of video frame data on one or more displays 106 or within a second area occupied by model 118 on one or more displays 106. In other examples, pointer 202 may be absent, and user input may be received in another manner, such as via touch input on one of the displays 106. In this particular example, user input is provided within the second area occupied by model 118 on one or more displays 106. For example, operator can “click”, “touch”, or “drag” at one or more specific locations on one or more displays 106. As mentioned, user input can be associated with features not yet displayed in a specific area (i.e., within representation 108 of video frame data or model 118). For example, user input may include user input corresponding to a navigation path 208 for generating vehicle 110. The navigation path may be associated with one or more driving instructions for vehicle 110, such as instructions to follow a specific route, path, or trajectory in the environment. In some cases, existing navigation routes may already be displayed in a specific area. Therefore, user input may include user input corresponding to modifying the existing navigation route 208 of vehicle 110. For example, points along the navigation route may be repositioned, or the navigation route may be lengthened or shortened.

[0082] Therefore, in this example, the features associated with user input are the navigation path rather than those within the environment. Figures 2 to 5 (The example) corresponds to the existing features.

[0083] The navigation path 208, modified or generated by user input, can be associated with one or more orientations determined by user input on one or more displays 106. Each orientation corresponds to a physical location within the environment.

[0084] After receiving user input associated with navigation path 208, one or more displays 106 can display indications of the navigation path at one or more specific orientations corresponding to the same physical location in the environment within another area (i.e., within the representation of video frame data 108 or the other in model 118). Therefore, if the user input is received at at least a first orientation within a first area (i.e., within the representation of video frame data 108) or a second area (i.e., within model 118), at least a second corresponding orientation within the other area can be determined. Regarding Figure 2 Examples of how this can be achieved have been discussed, so for the sake of brevity, they will not be repeated here.

[0085] After the corresponding second position has been determined, one or more displays 106 can display a feature indication 204 at the corresponding second position. Figure 6 In the example, because the user input is received in the second region (i.e., within model 118), the second orientation corresponds to the orientation in the first region (i.e., within the representation 108 of the video frame data). Therefore, Figure 6 An indication 204 is shown of the navigation path displayed in at least a second position within the representation 108 of the video frame data. Therefore, the operator 122, viewing model 118, can see the navigation path in the representation 108 of the video frame data 118. Thus, the operator 112 can determine whether the navigation path is appropriate.

[0086] After, simultaneously with, or before one or more displays 106 display the navigation route indication 204 in a second orientation, data can be sent from system 100 to or otherwise transmitted to vehicle 110. The data can be based on user input and cause vehicle 110 to take actions, such as moving or behaving in a particular manner. Data received from the system can be provided to the vehicle's planning components for processing. In this example, the data sent from system 100 to vehicle 110 includes data associated with the navigation route for vehicle 110 to follow.

[0087] In the example, the data associated with the navigation path sent by system 100 to vehicle 110 may include one or more locations within the environment associated with the navigation path (where the locations may be determined based on user input). In the example, the data associated with the navigation path sent by system 100 to vehicle 110 may include one or more driving instructions, such as "drive forward 5m", "turn left", "proceed to the next intersection", etc., rather than one or more locations within the environment associated with the navigation path, or processing said one or more locations.

[0088] Figure 7 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 7 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0089] like Figure 7As shown, operator 122 can provide user input (such as first user input) via user input device 120 by moving and clicking pointer 202 (such as a mouse pointer) within a first area occupied by representation 108 of video frame data on one or more displays 106 or within a second area occupied by model 118 on one or more displays 106. In other examples, for example, pointer 202 may not be present, and user input may be received in another manner, such as via touch input on one of the displays 106. In this particular example, user input is provided within the second area occupied by model 118 on one or more displays 106. For example, operator can “click” or “touch” at a specific location on one or more displays 106. As mentioned, user input can be associated with features already displayed in the specific area (i.e., within representation 108 of video frame data or model 118). The location determined by user input on one or more displays 106 corresponds to a physical location within the environment. In this example, the features associated with the user input correspond to mailboxes (as depicted in model 118), although the features actually correspond to pedestrians (as seen in representation 108 of the video frame data).

[0090] Therefore, in this example, vehicle 110 may have incorrectly classified an object / feature (i.e., a pedestrian) as another object / feature (i.e., a mailbox). In this case, operator 122 can provide user input to classify or reclassify a specific feature / object, and the result of the classification or reclassification is then sent to vehicle 110.

[0091] After receiving user input associated with a feature, one or more displays 106 can display an indication of the feature at a specific orientation corresponding to the same physical location in the environment in another area (i.e., within the representation 108 of the video frame data or the other of model 118). Therefore, if the user input is received at a first orientation within a first area (i.e., within the representation 108 of the video frame data) or a second area (i.e., within model 118), a second corresponding orientation within the other area can be determined. Regarding Figure 2 Examples of how this can be achieved have been discussed, so for the sake of brevity, they will not be repeated here.

[0092] After the second orientation has been determined, one or more displays 106 can display a feature indication 204 at the second orientation. Figure 7 In the example, because the user input is received in the second region (i.e., within model 118), the second orientation corresponds to the orientation in the first region (i.e., within the representation 108 of the video frame data). Therefore, Figure 7The indicator 204 for the feature is shown, in this example, as a circle surrounding the corresponding feature displayed within the representation 108 of the video frame data. It should be understood that displaying an indicator of a feature at a second location does not necessarily require the indicator to occupy a pixel exactly corresponding to the second location; rather, it can generally indicate the second location. For example, in this example, the circle-shaped indicator 204 could generally indicate the second location, perhaps by centering on the second location. Other examples are possible.

[0093] Data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with, or before the display of feature indication 204 on one or more displays 106 in a second orientation. The data can be based on user input and cause vehicle 110 to take actions, such as moving forward or behaving in a particular manner. Data received from the system can be provided to the vehicle's planning components for processing. In this example, the data sent from system 100 to vehicle 110 includes a classification of features / objects provided by user input.

[0094] Therefore, in any of the examples discussed above or in the examples discussed herein, operator 122 may provide a second user input to provide a classification for a feature, wherein the first user input corresponds to a selection of a feature, as discussed above. Thus, receiving user input may include receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of a first or second area of ​​one or more displays 106, and the second user input corresponding to a classification of a feature. Therefore, the data sent from system 100 to vehicle 110 may include classification. In the example, the data sent from system 100 to vehicle 110 may also include the location of the feature within the environment (where the location may be determined based on the first orientation). In this way, vehicle 110 may act on classification based on location. For example, the vehicle may behave in a particular way to consider the classification of objects. In some cases, vehicle 110 may receive classification without needing to receive the location of the feature. For example, vehicle 110 may determine the location of the feature based on its own location information associated with the feature.

[0095] Figure 8 An example is depicted showing how operator 122 can continue to provide classifications for features after the initial user input has been provided. For example, immediately following... Figure 7 The system can display prompt 206, such as a graphical menu. In some cases, prompt 206 can be displayed in response to receiving the first user input (i.e., selection of a feature, in this case, a pedestrian). In other cases, prompt 206 can be displayed in response to receiving additional user input.

[0096] like Figure 8As shown, after prompt 206 has been displayed, operator 122 can provide user input (such as a second user input) via user input device 120 by moving pointer 202 and selecting a category. In this example, operator 122 selects one of the options displayed in prompt 206, thereby selecting a category. In this example, the category corresponds to "pedestrians".

[0097] In other cases, no prompt is displayed, and user input indicating the category can be provided in another way, such as through verbal or typed commands. Other examples are possible.

[0098] In the example, vehicle 110 may have incorrectly classified the object / feature, and operator 122 provides user input indicating driving instructions (instead of providing a classification of the feature, or anything other than providing a classification of the feature). Therefore, Figure 9 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 9 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0099] like Figure 9 As shown, operator 122 can provide user input (such as first user input) via user input device 120 by moving and clicking pointer 202 (such as a mouse pointer) within a first area occupied by representation 108 of video frame data on one or more displays 106 or within a second area occupied by model 118 on one or more displays 106. In other examples, for instance, pointer 202 may not be present, and user input may be received in another manner, such as via touch input on one of the displays 106. In this particular example, user input is provided within a first area occupied by representation 108 of video frame data on one or more displays 106. For example, operator can “click” or “touch” at a specific location on one or more displays 106. As mentioned, user input can be associated with features already displayed in a specific area (i.e., within representation 108 of video frame data or model 118). The location determined by user input on one or more displays 106 corresponds to a physical location within the environment. In this example, the features associated with the user input correspond to trash or some other item (as seen in representation 108 of the video frame data), although in model 118 the features are depicted as pedestrians.

[0100] Therefore, in this example, vehicle 110 may have incorrectly classified an object / feature (i.e., garbage) as another object / feature (i.e., a pedestrian). In this case, operator 122 can provide user input to provide driving instructions for the incorrectly classified object, which are then sent to vehicle 110.

[0101] After receiving user input associated with a feature, one or more displays 106 can display an indication of the feature at a specific orientation corresponding to the same physical location in the environment in another area (i.e., within the representation 108 of the video frame data or the other of model 118). Therefore, if the user input is received at a first orientation within a first area (i.e., within the representation 108 of the video frame data) or a second area (i.e., within model 118), a second corresponding orientation within the other area can be determined. Regarding Figure 2 Examples of how this can be achieved have been discussed, so for the sake of brevity, they will not be repeated here.

[0102] After the second orientation has been determined, one or more displays 106 can display a feature indication 204 at the second orientation. Figure 9 In the example, because the user input is received in the first region (i.e., within the representation 108 of the video frame data), the second orientation corresponds to the orientation in the second region (i.e., within model 118). Therefore, Figure 9 The indicator 204 for the feature is shown, in this example, as a marker in the form of a circle surrounding the corresponding feature displayed within model 118. It should be understood that displaying an indicator of a feature in a second location does not necessarily require the indicator to occupy a pixel exactly corresponding to the second location; rather, it can generally indicate the second location. For example, in this example, the circular indicator 204 can generally indicate the second location, perhaps by centering on the second location. Other examples are possible.

[0103] Data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with, or before the display of feature indication 204 on one or more displays 106 in a second orientation. The data can be based on user input and cause vehicle 110 to take actions, such as moving forward or behaving in a particular manner. Data received from the system can be provided to the vehicle's planning components for processing. In this example, the data sent from system 100 to vehicle 110 includes driving instructions provided by user input. In this example, the data also includes a classification of features / objects provided by user input.

[0104] Therefore, as about Figure 4 and Figure 5The discussed receiving of user input may include receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of a first or second area of ​​one or more displays 106, and the second user input corresponding to driving instructions for the vehicle. Therefore, data sent from system 100 to vehicle 110 may include driving instructions for the vehicle to navigate in its environment. In the example, the data sent from system 100 to vehicle 110 may also include the location of a feature within the environment (where the location may be determined based on the first orientation). In this way, vehicle 110 may apply driving instructions based on location. For example, the vehicle may stop before, ignore, or bypass the location. In some cases, vehicle 110 may apply driving instructions without needing to receive the location of the feature. For example, vehicle 110 may apply instructions based on its own location information associated with the feature, or may apply instructions immediately without using location information.

[0105] Figure 10 An example is depicted showing how operator 122 can continue to provide driving instructions for the vehicle after the initial user input has been provided. For example, immediately following... Figure 9 The system can display prompt 206, such as a graphical menu. In some cases, prompt 206 can be displayed in response to receiving the first user input (i.e., selection of a feature, in this case, a pedestrian). In other cases, prompt 206 can be displayed in response to receiving additional user input.

[0106] like Figure 10 As shown, after prompt 206 has been displayed, operator 122 can provide user input (such as a second user input) via user input device 120 by moving pointer 202 and selecting a driving instruction. In this example, operator 122 selects one of the options displayed in prompt 206, thereby selecting a driving instruction. In this example, the driving instruction corresponds to "ignore".

[0107] In other cases, no prompt is displayed, and user input for indicating driving instructions can be provided in another way, such as through verbal or typed commands. Other examples are possible.

[0108] Figure 11 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 11As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0109] In the example, the feature (i.e., the pedestrian in this example) may be partially or completely occluded / hidden in one of the regions / views 118 and 108. For example, as Figure 11 As shown, a pedestrian may be at least partially obscured by another object in the environment, such as a vehicle. Therefore, it may be useful for operator 122 to provide user input to select a feature (i.e., a pedestrian) in the view where the feature is visible or most clearly visible. Thus, as described in the previous example, operator 122 may provide user input (such as a first user input) that allows one or more displays 106 to display an indication 204 of the feature at a specific orientation corresponding to the same physical location in the environment in another area (i.e., within the representation 108 of the video frame data or another of the models 118).

[0110] However, in other examples, instead of requiring user input in one area to display an indication of a feature in another area, an indication of a feature associated with one area can be automatically displayed in another area. As an example, this might occur if a feature is visible or partially visible in only one area 108, 118, such as if the feature is missing from model 118 or hidden / occluded in the representation 108 of the video frame data. Determining whether a feature is "missing" or "hidden / occluded" can be achieved by accessing the video frame data and the model (i.e., the output data). Therefore, in the example, if it is determined that at least a first feature among a plurality of features is not present in both the video frame data and the model, one or more displays 106 can display an indication of a feature in the representation 108 of the video frame data or the model 118, where the feature is "missing" or "hidden / occluded."

[0111] In examples where user input is not received at a specific location in either the first or second region, the location of the feature within the actual video frame data or model can be determined alternatively (i.e., not a transformation / conversion of the first location on display 106 to a second location on display 106, as discussed in the previous examples). For example, for each feature in the model, an examination can be performed to determine whether the video frame data includes data corresponding to the feature (or vice versa).

[0112] In model 118, if a feature exists (i.e., there are no missing or occluded features), its location can be determined within the model (or vice versa). If the model and video frame data are associated with different coordinate systems, the location can be transformed / transformed to the other in either the model or the video frame data. For example, the corresponding location of a feature in the video frame data can be determined based on its location in the model. If the model and video frame data are associated with the same coordinate system (so the location of a feature in the model corresponds to the same location in the video frame data), no transformation between coordinate systems is required. From this, the physical orientation on display 106 can be determined. To determine where the indication of a feature should be displayed on one or more displays 106 (i.e., where the indicated orientation should be displayed), one can follow the... Figure 2 The process is similar to the one described in the text.

[0113] In the example, the orientation of the indicated location can be determined by: (i) determining the position of the feature within the video frame data or model, and (ii) determining the orientation (on the display) corresponding to that position. For example, in Figure 11 In the case where the feature exists in the model but not in the video frame data, determining the orientation of the display indication can be achieved by: (i) determining the position of the feature within the model, and (ii) determining the position-based orientation (on the display) based on the position, where the orientation is within the representation of the video frame data (i.e., the first region). In other examples where the feature exists in the video frame data but not in the model, determining the orientation of the display indication can be achieved by: (i) determining the position of the feature within the video frame data, and (ii) determining the position-based orientation (on the display) based on the position, where the orientation is within the model (i.e., the second region).

[0114] In the example, the positions in the video frame data and model correspond to the same physical positions in the environment, and the positions within the environment can be associated with a coordinate system.

[0115] In other cases, the video frame data may be associated with a first coordinate system, and the model (and therefore the output data) may be associated with a second coordinate system. Therefore, in the example, step (ii) may initially include a transformation between the first and second coordinate systems. In some cases, each location in the video frame data (i.e., each location within the first coordinate system) may be mapped to a corresponding location in the model (i.e., each location within the second coordinate system), where the corresponding location corresponds to the same physical location in the environment.

[0116] In the example, each location in the video frame data and model can be mapped to an orientation in a first region (i.e., representation 108 of the video frame data) and a second region (i.e., model 118). Therefore, once the location in the video frame data or model is known, the corresponding orientation in the first or second region can be determined. A lookup table can be used to convert the locations in the video frame data and model into the two regions.

[0117] After the orientation on the display has been determined, one or more displays 106 can display a feature indicator 204 at the orientation. Figure 11 In the example, because the feature does not exist in the video frame data and is therefore not displayed in the first region (i.e., within representation 108 of the video frame data), the orientation corresponds to the orientation within the first region (i.e., within representation 108 of the video frame data). Therefore, Figure 11 The indicator 204 is shown, in this example, as a marker in the form of a circle surrounding the corresponding "missing," "hidden," or partially "hidden" feature displayed within the representation 108 of the video frame data. Therefore, indicator 204 shows the location corresponding to the same feature in the environment. Thus, the operator 122, viewing the representation 108 of the video frame data, can see where the same feature is located within the representation 108 of the video frame data. In this example, a graphical model of the feature (i.e., the pedestrian) can be generated instead of a marker.

[0118] As discussed in other examples, data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with or before one or more displays 106 display an indication 204 of a feature in orientation.

[0119] Figure 12 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 12 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0120] In the example, the feature (i.e., the pedestrian in this example) may not exist in either region / view 118 or 108. For example, as Figure 12As shown, pedestrians may be visible in representation 108 of the video frame data, but may be missing from model 118. Therefore, it may be useful for operator 122 to provide user input to select features (i.e., pedestrians) in a view where features are visible. Thus, as described in the previous example, operator 122 may provide user input (such as a first user input) that allows one or more displays 106 to display an indication 204 of a feature in a specific orientation corresponding to the same physical location in the environment in another area (i.e., within representation 108 of the video frame data or another of model 118).

[0121] However, in other examples, instead of requiring user input in one area to display an indication of a feature in another area, an indication of a feature associated with one area can be automatically displayed in another area. As an example, this might occur if a feature is visible or partially visible in only one area 108, 118, such as if the feature is missing from model 118 or hidden / occluded in the representation 108 of the video frame data. Determining whether a feature is "missing" or "hidden / occluded" can be achieved by accessing the video frame data and the model (i.e., the output data). Therefore, in the example, if it is determined that at least a first feature among a plurality of features is not present in both the video frame data and the model, one or more displays 106 can display an indication of a feature in the representation 108 of the video frame data or the model 118, where the feature is "missing" or "hidden / occluded."

[0122] Such as about Figure 11 In the example discussed, where user input is not received at a specific location within the first or second region, the location for displaying indicator 204 can be automatically determined. Once the location has been determined, one or more displays 106 can display the characteristic indicator 204 at that location. Figure 12 In the example, because the feature does not exist in the model and is therefore not displayed in the second region (i.e., within model 118), the orientation corresponds to the orientation in the second region (i.e., within model 118). Therefore, Figure 12 The indicator 204 is shown, in this example, as a marker displayed within model 118 in the form of a circle surrounding the corresponding "missing," "hidden," or partially "hidden" feature. Therefore, indicator 204 is displayed at the location corresponding to the same feature in the environment. Thus, the operator 122 viewing model 118 can see where the same feature is located within model 118. In this example, a graphical model of the feature (i.e., the pedestrian) can be generated instead of a marker.

[0123] As discussed in other examples, data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with or before one or more displays 106 display an indication 204 of a feature in orientation.

[0124] Regardless of whether user input for feature selection is received or whether an instruction is automatically displayed, the data sent from system 100 to vehicle 110 may include additional output data (such as additional perception output data) associated with the feature. This additional output data is particularly useful in cases where the model (and therefore the output data used by the vehicle's planning components) lacks data associated with the feature / object. This could happen if sensor 112 malfunctions or has a limited field of view of the environment, perhaps due to, for example, another object obscuring the sensor. Therefore, in this example, the additional output data may be generated by the system, where the additional output data is associated with the feature, and the data sent by the system to the vehicle includes the additional output data. In this example, the additional output data includes data associated with the feature (such as classification) and the feature's location within the environment. In cases where user input for feature selection is received, the feature's location within the environment may be based on a first orientation associated with the user input.

[0125] Figure 13 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 13 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0126] In this example, the output data received from vehicle 110 includes map data, which contains location information associated with one or more map features, such as roads, sidewalks, intersections, traffic lights, etc. In other examples, in addition to receiving map data from the vehicle, or even without receiving map data from the vehicle, system 100 can access map data that can be displayed as indications of map data (i.e., map features) in both the video frame data representation 108 and the model 118. In this example, the map data indications are always displayed in model 118, allowing operator 122 to know where lane boundaries, intersections, etc., are located in model 118. In this example, it may be useful to incorporate map data into the video frame data representation 108. This can be done automatically (or upon user request). If physical map features (such as lines drawn on roads) are found in the video frame data representation 108, the indications of the corresponding map features should be overlaid on the physical map features in the video frame data representation 108. If physical map features deviate from the indications of map features already shown in the representation 108 of the video frame data, this could be an indication of a problem, such as vehicle 110 incorrectly determining its location, or poor calibration between the video camera and other sensors 112.

[0127] Indications of map features can be automatically displayed in the representation 108 of the video frame data. As discussed with respect to 11, the orientation for displaying indication 204 can be automatically determined. For example, each map feature in the map data is associated with a location (and therefore the environment) within the model. From this, the orientation corresponding to the location can be determined (on the display). After the orientation has been determined, one or more displays 106 can display the feature indication 204 at the orientation in the representation 108 of the video frame data. It should be understood that some map features (such as lane markings) can be associated with multiple locations, and the map feature indication 204 can correspond to each of those locations. In other words, indications can be displayed at multiple orientations on one or more displays 106. In the same or similar manner, route data (such as navigation paths or routes) can be indicated in the representation 108 of the video frame data.

[0128] therefore, Figure 13 Indication 204, representing a map feature, is shown in this example as a lane marking displayed within representation 108 of the video frame data. Therefore, indication 204 is displayed at a location corresponding to the same physical map feature. Thus, the operator 122, viewing representation 108 of the video frame data, can see where the same feature is located within representation 108 of the video frame data.

[0129] As discussed in other examples, data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with or before one or more displays 106 display an indication 204 of a feature in orientation.

[0130] As another example, the colors within representation 108 of the video frame data can be incorporated into model 118. This can be done automatically (or upon user request). Therefore, the indication of a displayed feature can correspond to the indication of the color of the displayed feature within model 118, where the color is determined from representation 108 of the video frame data. In any example where the user inputs a selection of a specific feature within representation 108 or model 118 of the video frame data, the colors within representation 108 of the video frame data can also be incorporated into model 118.

[0131] Figure 14 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 14 As shown, the output on one or more displays 106 may also include a 2D or 3D model 118 of the environment based on environmental data. Model 118 may include models / renders of buildings, vehicles and pedestrians around the vehicles, as well as a model / render of the vehicle 110 itself.

[0132] As briefly mentioned above, in the example, if the environment (such as output data) is detected to contain a priority agent / object (such as an emergency vehicle), an indication of the priority agent can be displayed in the video view (or vice versa). This can be useful for emphasizing the priority agent to operator 122 when it is not visible or difficult to see in a view. For example, as Figure 14 As shown, vehicle 110 may be stuck in traffic, and several other vehicles may be very close to vehicle 110.

[0133] Vehicle 110 and / or system 100 can determine priority agent 210, such as an ambulance, based on output data. For example, sensors 112 on vehicle 112 can capture environmental data, including sirens, text (such as "ambulance"), and / or other data indicating that the vehicle is priority agent 210. Therefore, model 118 can render features as priority agent 210. For example, Figure 14An "A" is depicted on the model of priority agent 210 to notify operator 122 that the vehicle is an ambulance. In some cases, such as due to one or more obstacles, a particular field of view, lighting conditions, etc., priority agent 210 may not be very obvious in the representation 108 of the video frame data. Therefore, in this example, it may be useful to indicate priority agent 210 in the representation 108 of the video frame data. This can be done automatically, for example, without requiring user input to select features in one of views 108, 118.

[0134] As in Figure 11 In the examples discussed earlier, where the user input is not received at a specific orientation in the first or second region, the location of the feature within the model can be determined alternatively (i.e., not by converting / transforming the first orientation on display 106 to a second orientation on display 106, as discussed in the previous examples). For example, if priority agent 210 is detected in model 118, the location of the feature can be determined in the model. This location can be converted / transformed into a location in the video frame data. For example, the corresponding location of the feature in the video frame data can be determined based on the location in the model. From this, the physical orientation on display 106 can be determined, as in... Figure 11 As discussed in the article.

[0135] After the orientation on the display has been determined, one or more displays 106 can display a feature indicator 204 at the orientation. Figure 14 In the example, because the feature is detected in the model, the orientation corresponds to the orientation in the first region (i.e., within the representation 108 of the video frame data). Therefore, Figure 14 An indication 204 is shown within the representation 108 of the video frame data, representing a feature (priority agent 210), in this example, as a marker in the form of a circle surrounding the corresponding "missing," "hidden," or partially "hidden" feature. Thus, indication 204 is displayed at a location corresponding to the same feature in the environment. Therefore, an operator 122 viewing the representation 108 of the video frame data can see where the priority agent 210 is located within the representation 108 of the video frame data. In an example where the camera's field of view prevents the representation of the video frame data from including the priority agent, indication 204 can still be displayed, informing the operator 122 that the priority agent is nearby, such as just outside the view. For example, the indication could refer to the location of the priority agent within the environment. For example, the indication could include an arrow pointing in a specific direction to the priority agent.

[0136] As discussed in other examples, data can be sent from system 100 to or otherwise transmitted to vehicle 110 after, simultaneously with, or before one or more displays 106 display indication 204 of a feature at a location. For example, the data may include data associated with a priority agent, such as driving instructions.

[0137] In some examples, when one or more displays 106 display model 118, the one or more displays 106 display or render model 118 such that the model has a viewpoint substantially corresponding to the viewpoint of the representation 108 of the video frame data. Therefore, model 118 can be rendered from the same orientation in the environment and has the same field of view as the representation 108 of the video frame data. This can allow for easier rendering of features from one view (such as pedestrians) in another view. For example, if a pedestrian is obscured by a vehicle, it can be rendered in a way similar to the above description. Figure 11 The description is similar to that used in the representation of video frame data 108, which displays a 3D rendering of pedestrians.

[0138] therefore, Figure 15 A close-up depiction of the output on one or more displays 106 viewed by operator 122, according to another example. As shown, the output on one or more displays 106 includes a representation 108 of video frame data. Also... Figure 15 As shown, the output on one or more displays 106 includes an environment-based model 118 of the environment. Model 118 is rendered with the same viewpoint as the representation 108 of the video frame data.

[0139] Figure 16 A flowchart of example method 300 is shown. Example method 300 may be implemented by one or more components of system 100. In the example, method 300 may be encoded and stored as instructions on one or more non-transitory computer-readable media, which, when executed by one or more processors 102 of system 100, cause system 100 to implement method 300. In the example, the method is a computer-implemented method.

[0140] As in Figure 16As can be seen, method / process 300 may include receiving, at step 302, from the autonomous vehicle: (i) sensor data captured by the vehicle's sensors, and (ii) output data, wherein the output data is generated at least in part based on environmental data captured by one or more of the vehicle's sensors and used by the vehicle's planning components for navigation in the environment, wherein the sensor data and output data indicate the environment. At step 304, the method includes causing one or more displays to display at least one of: (i) a representation of the sensor data, or (ii) a model of the environment based on the output data. In some examples, step 304 includes causing one or more displays to display both: (i) a representation of the sensor data, or (ii) a model of the environment based on the output data. For example, the representation of the sensor data may be displayed in a first area, and the model may be displayed in a second area. At step 306, the method includes determining the location of a feature within the sensor data or model (this may correspond to determining its location within the environment, such as within a coordinate system). At step 308, the method includes causing one or more displays to display an indication of the feature at an orientation corresponding to the location on one or more displays. For example, if the location of the feature is determined in the sensor data, an indication can be displayed in the model, and if the location of the feature is determined in the model, an indication can be displayed in the representation of the sensor data. At step 310, the method includes receiving user input at the system. At step 312, the method includes the system sending data based on the user input to the vehicle to cause the vehicle to take action.

[0141] As mentioned, in some examples, sensor data can correspond to video data that includes video frame data.

[0142] As mentioned, in some cases, step 306 is triggered based on receiving user input (such as a first user input). Therefore, step 310 may occur before step 306, although one or more additional user inputs may be received after step 306. Thus, in the example, steps 306 through 308 alternatively include determining a first orientation associated with the feature within a first or second region (as discussed in previous examples), determining a second orientation of the feature within the other of the first or second region based at least in part on the first orientation, and determining at least in part on whether the first orientation is within the first or second region of one or more displays, such that one or more displays display an indication of the feature at the second orientation within the other of the first or second region.

[0143] In other cases, step 306 is triggered based on the object / feature having a low classification probability (such as below a threshold probability). For example, if an object in the model data is identified as a pedestrian, but a vehicle has low confidence in this (the object could actually be a traffic cone), the object can be indicated in the representation of the video frame data. Therefore, the low classification probability feature is then indicated / highlighted to the operator.

[0144] In some examples, the user can (via input) indicate that features belonging to a certain category should be represented in the representation of the video frame data. For example, it might be useful to indicate all pedestrians (determined in the model) in the video representation of the video frame data.

[0145] In the example, steps 310 and 312 are omitted.

[0146] Figure 17 A block diagram of an example system 400 implementing the techniques discussed above and in this paper is shown. Figure 17 It can represent Figure 1 The vehicle 110 and system 100. In some cases, example system 400 may include vehicle 402, which may represent Figure 1 Vehicle 110 is mentioned. In some cases, vehicle 402 may be an autonomous vehicle configured to operate according to a Level 5 classification issued by the National Highway Traffic Safety Administration (NHTSA), which describes a vehicle capable of performing all safety-critical functions throughout the journey without expecting the driver (or occupant) to control the vehicle at any time. However, in other examples, vehicle 402 may be a fully or partially autonomous vehicle with any other level or classification. Furthermore, in some cases, the techniques described herein may also be usable by non-autonomous vehicles.

[0147] Vehicle 402 may include vehicle computing unit 404, sensor 406, transmitter 408, network interface 410, and / or drive system 412. Sensor 406 may represent sensor 112 discussed above.

[0148] In some cases, sensor 406 may include lidar sensors, radar sensors, ultrasonic transducers, sonar sensors, position sensors (e.g., Global Positioning System (GPS), compasses, etc.), inertial sensors (e.g., inertial measurement units (IMUs), accelerometers, magnetometers, gyroscopes, etc.), image sensors (e.g., red-green-blue (RGB), infrared (IR), intensity, depth, time-of-flight cameras, etc.), microphones, wheel encoders, environmental sensors (e.g., thermometers, hygrometers, light sensors, pressure sensors, etc.), etc. Sensor 406 may include multiple instances of each of these or other types of sensors. For example, radar sensors may include individual radar sensors located at corners, front, rear, sides, and / or top of vehicle 402. As another example, cameras may include multiple cameras located at various locations near the exterior and / or interior of vehicle 402. Sensor 406 may provide input to vehicle computing unit 404 and / or computing unit 432.

[0149] Data captured by sensors can be referred to as sensor data or environmental data. In the example, sensor data captured by a camera can be referred to as sensor data or video frame data.

[0150] Vehicle 402 may also include a transmitter 408 for emitting light and / or sound, as described above. Transmitter 408 may include internal audio and visual transmitters for communicating with passengers of vehicle 402. Internal transmitters may include speakers, lights, signs, displays, touchscreens, haptic transmitters (e.g., vibration and / or force feedback), mechanical actuators (e.g., seatbelt tensioners, seat positioners, headrest positioners, etc.), and so on. Transmitter 408 may also include external transmitters. External transmitters may include lights (e.g., indicator lights, signs, light arrays, etc.) for signaling other indications of direction of travel or vehicle movement, and one or more audio transmitters (e.g., speakers, speaker arrays, horns, etc.) for audible communication with pedestrians or other nearby vehicles, one or more of which include beam steering technology.

[0151] Vehicle 402 may also include a network interface 410 that enables communication between vehicle 402 and one or more other local or remote computing devices. Network interface 410 may facilitate communication with other local computing devices and / or drive components 412 on vehicle 402. Network interface 410 may additionally or alternatively allow the vehicle to communicate with other nearby computing devices (e.g., other nearby vehicles, traffic lights, etc.). Network interface 410 may additionally or alternatively enable vehicle 402 to communicate with computing device 432 via network 438. In some examples, computing device 432 may include one or more nodes of a distributed computing system (e.g., a cloud computing architecture). Computing device 432 corresponds to system 100 discussed above.

[0152] Vehicle 402 may include one or more drive components 412. In some cases, vehicle 402 may have a single drive component 412. In some cases, drive component 412 may include one or more sensors to detect the condition of drive component 412 and / or the surrounding environment of vehicle 402. By way of example and not limitation, sensors for drive component 412 may include: one or more wheel encoders (e.g., rotary encoders) for sensing the rotation of the wheels of drive component; inertial sensors (e.g., inertial measurement units, accelerometers, gyroscopes, magnetometers, etc.) for measuring the orientation and acceleration of drive component; cameras or other image sensors; ultrasonic sensors for acoustically detecting objects in the surrounding environment of drive component; lidar sensors; radar sensors, etc. Some sensors (such as wheel encoders) may be unique to drive component 412. In some cases, sensors on drive component 412 may be superimposed on or complement a corresponding system of vehicle 402 (e.g., sensor 406).

[0153] Drive assembly 412 may include a number of vehicle systems, including a high-voltage battery, a motor for propelling the vehicle, an inverter for converting direct current from the battery into alternating current for use by other vehicle systems, a steering system including a steering motor and a steering rack (which may be electric), a braking system including hydraulic or electric actuators, a suspension system including hydraulic and / or pneumatic components, a stability control system for distributing braking force to mitigate traction loss and maintain control, an HVAC system, lighting (e.g., headlights / taillights for illuminating the external surroundings of the vehicle), and one or more other systems (e.g., cooling systems, safety systems, on-board charging systems, other electrical components such as DC / DC converters, high-voltage connectors, high-voltage cables, charging systems, charging ports, etc.). Additionally, drive assembly 412 may include a drive assembly controller that can receive and preprocess data from sensors and control the operation of various vehicle systems. In some cases, the drive assembly controller may include one or more processors and a memory communicatively coupled to one or more processors. The memory may store one or more components to perform various functionalities of drive assembly 412. In addition, the drive component 412 may also include one or more communication connections that enable the respective drive component to communicate with one or more other local or remote computing devices.

[0154] Vehicle computing device 404 may include processor 414 and memory 416 communicatively coupled to one or more processors 414. Computing device 432 may also include processor 434 and / or memory 436. Processor 414 and / or 434 may be any suitable processor capable of executing instructions to process data and perform operations as described herein. By way of example and not limitation, processor 414 and / or 434 may include one or more central processing units (CPUs), graphics processing units (GPUs), integrated circuits (e.g., application-specific integrated circuits (ASICs)), gate arrays (e.g., field-programmable gate arrays (FPGAs)), and / or any other means or part of means for processing electronic data to convert such electronic data into other electronic data that can be stored in registers and / or memory.

[0155] Memory 416 and / or 436 may be examples of non-transitory computer-readable media. Memory 416 and / or 436 may store an operating system and one or more software applications, instructions, programs, and / or data to implement the methods described herein and the functions attributed to various systems. In various implementations, any suitable memory technology may be used to implement the memory, such as static random access memory (SRAM), synchronous dynamic RAM (SDRAM), non-volatile / flash memory, or any other type of memory capable of storing information. The architectures, systems, and individual elements described herein may include many other logical, programming, and physical components; those shown in the figures are merely examples relevant to the discussion herein.

[0156] In some cases, memory 416 and / or memory 436 may store sensing component 418, positioning component 420, planning component 422, map 424, driving log data 426, prediction component 428 and / or system controller 430, wherein zero or more parts of any of them may be hardware, such as GPU, CPU and / or other processing units.

[0157] The perception component 418 can detect objects in the environment surrounding vehicle 402 (e.g., identify the presence of objects), classify objects (e.g., determine the object type associated with the detected objects), segment sensor data and / or other representations of the environment (e.g., identify portions of the sensor data and / or environmental representations as associated with the detected objects and / or object types), determine characteristics associated with the objects (e.g., tracks, said tracks identifying current, predicted, and / or previous orientation, heading, speed, and / or acceleration associated with the objects), etc. The data determined by the perception component 418 is referred to as perception output data. The perception component 418 can be configured to associate boundary regions (or other indications) with the identified objects. The perception component 418 can be configured to associate a confidence score associated with the classification of the identified objects with the identified objects. In some examples, objects may be colored based on their perceived category when rendered via a display. The object classification determined by the perception component 418 can distinguish different object types, such as, for example, passenger cars, pedestrians, cyclists, drivers, delivery trucks, semi-trucks, traffic signs, etc.

[0158] In at least one example, the localization component 420 may include hardware and / or software to receive data from sensor 406 to determine the azimuth, velocity, and / or orientation (e.g., one or more of x-azimuth, y-azimuth, z-azimuth, roll, pitch, or yaw) of vehicle 402. For example, the localization component 420 may include and / or request / receive a map 424 of the environment and may continuously determine the position, velocity, and / or orientation of the autonomous vehicle 402 within map 424. In some cases, the localization component 420 may utilize SLAM (simultaneous azimuth and mapping), CLAMS (simultaneous calibration, azimuth, and mapping), relative SLAM, bundle adjustment, nonlinear least squares optimization, etc., to receive image data, lidar data, radar data, IMU data, GPS data, wheel encoder data, etc., to accurately determine the position, attitude, and / or velocity of the autonomous vehicle. In some cases, the localization component 420 may provide data to various components of vehicle 402 to determine the initial azimuth of the autonomous vehicle for trajectory generation and / or for determining map data, as discussed herein. In some examples, the positioning component 420 may provide the sensing component 418 with the position and / or orientation of the vehicle 402 relative to the environment and / or associated sensor data.

[0159] The planning component 422 may receive position and / or orientation information of the vehicle 402 from the positioning component 420, and / or receive sensing data from the sensing component 418, and may determine instructions for controlling the operation of the vehicle 402 based at least in part on any such data. In some examples, determining the instructions may include determining the instructions based at least in part on a format associated with the system to which the instructions are associated (e.g., a first instruction for controlling the motion of the autonomous vehicle may be formatted such that the system controller 430 and / or drive component 412 can parse / make implement a first message and / or signal format (e.g., analog, digital, aerodynamic, kinematic), and a second instruction for the transmitter 408 may be formatted according to a second format associated therewith).

[0160] Driving log data 426 may include sensor data, perception data, and / or scene tags collected / determined by vehicle 402 (e.g., by perception component 418), as well as any other messages generated and / or sent by vehicle 402 during operation, including but not limited to control messages, error messages, etc. In some examples, vehicle 402 may transmit driving log data 426 to computing device 432.

[0161] Prediction component 428 can generate one or more probability maps representing predicted probabilities of the possible locations of one or more objects in the environment. For example, prediction component 428 can generate one or more probability maps for vehicles, pedestrians, animals, etc., within a threshold distance from vehicle 402. In some examples, prediction component 428 can measure the tracks of objects and generate discretized prediction probability maps, heatmaps, probability distributions, discretized probability distributions, and / or object trajectories based on observed and predicted behavior. In some examples, one or more probability maps can represent the intentions of one or more objects in the environment. In some examples, planning component 422 can be communicatively coupled to prediction component 428 to generate predicted trajectories of objects in the environment. For example, prediction component 428 can generate one or more predicted trajectories for objects within a threshold distance from vehicle 402. In some examples, prediction component 428 can measure the tracks of objects and generate object trajectories based on observed and predicted behavior. Although prediction component 428 is shown on vehicle 402 in this example, prediction component 428 can also be provided elsewhere, such as in a remote computing device. In some examples, prediction components can be provided at both the vehicle and the remote computing device. These components can be configured to operate according to the same or similar algorithms. Data generated by the prediction component 428 can be provided to the computing device 432 as prediction output data. More generally, output data can be provided to the computing device 432, and the output data may include both perceived output data and prediction output data. The output data can be used by the planning component 422 for navigation in the environment.

[0162] Memory 416 and / or 436 may additionally or alternatively store mapping systems, planning systems, riding management systems, etc. Although perception components 418 and / or planning components 422 are shown as stored in memory 416, perception components 418 and / or planning components 422 may include processor-executable instructions, machine learning models (e.g., neural networks), and / or hardware.

[0163] As described herein, the localization component 420, perception component 418, planning component 422, and / or other components of system 400 may include one or more ML models. For example, the localization component 420, perception component 418, and / or planning component 422 may each include different ML model pipelines. In some examples, the ML model may include a neural network. An exemplary neural network is a biologically inspired algorithm that passes input data through a series of connected layers to produce an output. Each layer in the neural network may also include another neural network, or may include any number of layers (whether convolutional or not). As will be understood in the context of this disclosure, neural networks may utilize machine learning, which can refer to a broad class of algorithms in which outputs are generated based on learned parameters.

[0164] Although discussed in the context of neural networks, any type of machine learning consistent with this disclosure can be used. For example, machine learning algorithms can include, but are not limited to, regression algorithms (e.g., ordinary least squares regression (OLSR), linear regression, logistic regression, stepwise regression, multivariate adaptive regression splines (MARS), local estimation scatter plot smoothing (LOESS)), instance-based algorithms (e.g., ridge regression, least absolute value convergence and selection operator (LASSO), elastic networks, least angular regression (LARS)), and decision tree algorithms (e.g., classification and regression trees (CART), iterative binary trees). (ID3), Chi-squared Automatic Interaction Detection (CHAD), Decision Stump, Conditional Decision Tree), Bayesian Algorithms (e.g., Naive Bayes, Gaussian Naive Bayes, Multinomial Naive Bayes, Average Single Dependency Estimator (AODE), Bayesian Belief Network (BNN), Bayesian Network), Clustering Algorithms (e.g., k-means, k-median, Expectation-Maximum (EM), Hierarchical Clustering), Association Rule Learning Algorithms (e.g., Perceptron, Backpropagation, Hopfield Network, Radial Basis Function Network (RBFN)), Deep Learning Algorithms (e.g., Deep Boltzmann Machine (DBM), Deep Belief Network (DBN), Convolutional Neural Network (CNN), Stacked Autoencoder), Dimensionality Reduction Algorithms (e.g., Principal Component Analysis (PCA), Principal Component Regression (PCR), Partial Least Squares Regression (PLSR), Sammon Mapping, Multidimensional Scaling (MDS), Projective Pursuit, Linear Discriminant Analysis (LDA), Mixture Discriminant Analysis (MDA), Quadratic Discriminant Analysis (QDA), Flexible Discriminant Analysis (FDA)), Ensemble Algorithms (e.g., Boosting, Bootstrapped) Aggregation (Bagging), AdaBoost, stacked generalization (hybrid), Gradient Boosting Machine (GBM), Gradient Boosting Regression Tree (GBRT), Random Forest, SVM (Support Vector Machine), supervised learning, unsupervised learning, semi-supervised learning, etc. Additional examples of architectures include neural networks such as ResNet-50, ResNet-101, VGG, DenseNet, PointNet, etc. In some examples, the ML models discussed herein may include PointPillars, SECOND, top-down feature layers (e.g., see U.S. Patent Application Serial No. 15 / 963,833, which is incorporated herein in its entirety), and / or VoxelNet. Architecture-delayed optimizations may include MobilenetV2, ShuffleNet, ChannelNet, PeleeNet, etc. In some examples, ML models may include residual blocks, such as Pixor.

[0165] The memory 420 may additionally or alternatively store one or more system controllers 430, which may be configured to control the steering, propulsion, braking, safety, transmitter, communication, and other systems of the vehicle 402. These system controllers 430 may communicate with and / or control corresponding systems of the drive assembly 412 and / or other components of the vehicle 402.

[0166] It should be noted that, although Figure 17 While shown as a distributed system, in alternative examples, components of vehicle 402 may be associated with computing device 432, and / or components of computing device 432 may be associated with vehicle 402. That is, vehicle 402 may perform one or more functions associated with computing device 432, and vice versa.

[0167] Example Terms A system comprising: One or more processors; and One or more non-transitory computer-readable media having instructions stored thereon, which, when executed by the one or more processors (or the system), cause the one or more processors (or the system) to perform operations, including: Received from the autonomous vehicle at the system: Video frame data captured by the vehicle's video camera; and Perception output data generated by the vehicle, wherein the perception output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components for navigation in the environment, wherein the video frame data and the perception output data indicate the environment; This causes a representation of the video frame data to be displayed in a first area of ​​one or more displays; The second area of ​​the one or more displays displays a model of the environment based on the perceived output data; The system receives user input, which is associated with a feature within one of the first or second regions. Determine a first orientation associated with the feature within the first region or the second region; The second orientation of the feature within the other of the first region or the second region is determined at least in part based on the first orientation; At least in part based on whether the first orientation is within the first or second region of the one or more displays, the one or more displays display an indication of the feature at the second orientation in either the first or second region; and The system sends data, at least in part based on the user input, to the vehicle to cause the vehicle to take action.

[0168] According to the system described in Clause 1, wherein: The feature corresponds to the vehicle's navigation path; The user input includes user input corresponding to modifying or generating the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

[0169] According to the system described in Clause 1, wherein: The feature is visible within the displayed representation of the video frame data; Instructions to cause the one or more displays to display the feature at the second orientation in either the first region or the second region include instructions to cause the one or more displays to display the feature at the second location in the second region; The operation also includes generating additional perception output data for use by the vehicle's planning component to navigate in the environment, wherein the additional perception output data is associated with the feature; and The data sent by the system to the vehicle includes the additional sensing output data, and the additional sensing output data includes: Data associated with the aforementioned feature; and The feature is located within the environment, the location being based on the first orientation.

[0170] According to the system described in Clause 1, wherein: The feature corresponds to an object within the environment; Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second regions of the one or more displays, and the second user input corresponding to a classification of the feature; and The data sent by the system to the vehicle includes the classification.

[0171] According to the system described in Clause 1, wherein: Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second areas of the one or more displays, and the second user input corresponding to a driving instruction for the vehicle; and The data sent from the system to the vehicle includes driving instructions for the vehicle to navigate in the environment.

[0172] According to any of the foregoing clauses, the system wherein: The video frame data and the model are associated with a coordinate system; Determining the second orientation of the feature within the other of the first region or the second region, at least in part, based on the first orientation, includes: Determine the position of the feature in the coordinate system based on the first orientation; and The second orientation is determined based on the stated location.

[0173] According to the system described in Clause 1, wherein: the feature corresponds to a predicted path associated with an object within the environment; and The user input corresponds to the selection of the object visible in either the displayed representation or the displayed model of the video frame data.

[0174] A method, the method comprising: Received from the autonomous vehicle at the system level: Sensor data captured by the vehicle's sensors; and Output data, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components to navigate in the environment, wherein the sensor data and the output data indicate the environment; The display shall cause one or more displays to show at least one of the following: (i) a representation of the sensor data or (ii) a model of the environment based on the output data; Determine the location of the feature within the sensor data or the model; This causes the one or more displays to show an indication of the feature at a position corresponding to the location on the one or more displays; The system receives user input; and The system sends data based on the user input to the vehicle so that the vehicle can take action.

[0175] According to the method described in Clause 8, determining the location of the feature within the sensor data or the model includes automatically determining the location of the feature within the sensor data or the model.

[0176] According to the method of Clause 9, said feature is a first feature of a plurality of features in the environment, and said method includes: Determining that at least the first feature of the plurality of features does not exist in either the sensor data or the model; and Based on the determination that at least the first feature among the plurality of features does not exist in either the sensor data or the model: This causes the one or more displays to show the indication of the feature.

[0177] According to the method described in Clause 9, wherein: The aforementioned features correspond to priority agents; The one or more displays are configured to display at least one of the following: (i) a representation of the sensor data or (ii) the model, and the one or more displays are configured to display an indication of the feature, including: Cause the one or more displays to display the representation of at least the sensor data; and This causes the one or more displays to show an indication of the priority agent within the representation of the sensor data; and The data sent from the system to the vehicle includes data associated with the priority agent.

[0178] The method according to any one of Clauses 8 to 11, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; and The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model comprising: The model is displayed on one or more displays such that the model has a viewing angle that substantially corresponds to the viewing angle of the representation of the sensor data.

[0179] According to the method described in Clause 8, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; Receiving the user input at the system includes receiving at least a first user input, the first user input being associated with the feature and a first orientation within the displayed representation or displayed model of the sensor data on the one or more displays; Receiving the first user input enables determining the location of the feature within the sensor data or the model based on the first orientation; and Indications that cause the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays include: At least in part, based on whether the first orientation is within the displayed representation or the displayed model of the sensor data, the one or more displays display the indication of the feature at the orientation based on the first orientation in either the displayed representation or the displayed model of the sensor data.

[0180] According to the method described in Clause 13, wherein: The feature corresponds to the vehicle's navigation path; The first user input corresponds to the modification or generation of the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

[0181] According to the method described in Clause 13, wherein: The feature corresponds to the predicted path associated with the object within the environment; and The first user input corresponds to the selection of the object visible in either the displayed representation or the displayed model of the sensor data.

[0182] According to the method described in Clause 13, wherein: Receiving the user input at the system includes receiving at least the first user input and the second user input, the second user input corresponding to one of the following: Driving instructions for the vehicle; or The classification of the features; and The data sent by the system to the vehicle includes one of the following: The driving instructions; or The classification.

[0183] The method according to any one of Clauses 8 to 16, wherein the one or more displays display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to show the representation of at least the sensor data; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the model; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the location within the representation of the sensor data.

[0184] The method according to any one of Clauses 8 to 16, wherein the one or more displays display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to display at least the model; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the sensor data; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the specified location within the model.

[0185] The method according to any one of Clauses 8 to 18, wherein the sensor data is video frame data.

[0186] According to the method described in Clause 8, wherein: The feature is visible within the displayed representation of the sensor data; Instructions that cause the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays include instructions that cause the one or more displays to display the feature at the orientation in the model; The method further includes generating additional output data for use by the vehicle's planning component to navigate in the environment, wherein the additional output data is associated with the feature; and The data sent by the system to the vehicle includes the additional output data, and the additional output data includes: Data associated with the aforementioned feature; and The feature is located at the specified location within the environment.

[0187] According to the method described in Clause 8, wherein: The sensor data and the model are associated with a coordinate system; Determining the position of a feature within the sensor data or the model includes determining the position of the feature within the coordinate system.

[0188] One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a system, cause the system to perform operations, the operations including: Received from the autonomous vehicle at the system: Sensor data captured by the vehicle's sensors; and Output data, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components to navigate in the environment, wherein the sensor data and the output data indicate the environment; The display shall cause one or more displays to show at least one of the following: (i) a representation of the sensor data or (ii) a model of the environment based on the output data; Determine the location of the feature within the sensor data or the model; This causes the one or more displays to show an indication of the feature at a position corresponding to the location on the one or more displays; The system receives user input; and The system sends data based on the user input to the vehicle so that the vehicle can take action.

[0189] One or more non-transitory computer-readable media as described in Clause 18, wherein at least one of the following: Determining the location of the feature within the sensor data or the model includes automatically determining the location of the feature within the sensor data or the model; and Receiving the user input at the system includes receiving at least a first user input associated with the feature and a first orientation within the displayed representation or model of the sensor data on the one or more displays, wherein receiving the first user input causes the position of the feature within the sensor data or the model to be determined based on the first orientation.

[0190] One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a system, cause the system to perform the method according to any one of clauses 8 to 21.

[0191] A system comprising: One or more processors; and One or more non-transitory computer-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of clauses 8 to 21.

[0192] A method, the method comprising: Received from the autonomous vehicle at the system level: Video frame data captured by the vehicle's video camera; and Perception output data generated by the vehicle, wherein the perception output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components for navigation in the environment, wherein the video frame data and the perception output data indicate the environment; This causes a representation of the video frame data to be displayed in a first area of ​​one or more displays; The second area of ​​the one or more displays displays a model of the environment based on the perceived output data; The system receives user input, which is associated with a feature within one of the first or second regions. Determine a first orientation associated with the feature within the first region or the second region; The second orientation of the feature within the other of the first region or the second region is determined at least in part based on the first orientation; At least in part based on whether the first orientation is within the first or second region of the one or more displays, the one or more displays display an indication of the feature at the second orientation in either the first or second region; and The system sends data, at least in part based on the user input, to the vehicle to cause the vehicle to take action.

[0193] According to the method described in Clause 26, wherein: The feature corresponds to the vehicle's navigation path; The user input includes user input corresponding to modifying or generating the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

[0194] According to the method described in Clause 26, wherein: The feature is visible within the displayed representation of the video frame data; Instructions to cause the one or more displays to display the feature at the second orientation in either the first region or the second region include instructions to cause the one or more displays to display the feature at the second location in the second region; The method further includes generating additional perception output data for use by the vehicle's planning component to navigate in the environment, wherein the additional perception output data is associated with the feature; and The data sent by the system to the vehicle includes the additional sensing output data, and the additional sensing output data includes: Data associated with the aforementioned feature; and The feature is located within the environment, the location being based on the first orientation.

[0195] According to the method described in Clause 26, wherein: The feature corresponds to an object within the environment; Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second regions of the one or more displays, and the second user input corresponding to a classification of the feature; and The data sent by the system to the vehicle includes the classification.

[0196] According to the method described in Clause 26, wherein: Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second areas of the one or more displays, and the second user input corresponding to a driving instruction for the vehicle; and The data sent from the system to the vehicle includes driving instructions for the vehicle to navigate in the environment.

[0197] The method according to any one of Clauses 26 to 30, wherein: The video frame data and the model are associated with a coordinate system; Determining the second orientation of the feature within the other of the first region or the second region, at least in part, based on the first orientation, includes: Determine the position of the feature in the coordinate system based on the first orientation; and The second orientation is determined based on the stated location.

[0198] According to the method described in Clause 26, wherein: the feature corresponds to a predicted path associated with an object within the environment; and The user input corresponds to the selection of the object visible in either the displayed representation or the displayed model of the video frame data.

[0199] A system comprising one or more processors and one or more non-transitory computer-readable media having instructions stored thereon, the instructions, when executed by the one or more processors (or the system), causing the one or more processors (or the system) to perform operations, the operations including: Received from the autonomous vehicle at the system level: Sensor data captured by the vehicle's sensors; and Output data, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components to navigate in the environment, wherein the sensor data and the output data indicate the environment; The display shall cause one or more displays to show at least one of the following: (i) a representation of the sensor data or (ii) a model of the environment based on the output data; Determine the location of the feature within the sensor data or the model; This causes the one or more displays to show an indication of the feature at a position corresponding to the location on the one or more displays; The system receives user input; and The system sends data based on the user input to the vehicle so that the vehicle can take action.

[0200] According to the system described in Clause 33, determining the location of the feature within the sensor data or the model includes automatically determining the location of the feature within the sensor data or the model.

[0201] According to the system described in Clause 34, said feature is a first feature among a plurality of features in said environment, and said operation further includes: Determining that at least the first feature of the plurality of features does not exist in either the sensor data or the model; and Based on the determination that at least the first feature among the plurality of features does not exist in either the sensor data or the model: This causes the one or more displays to show the indication of the feature.

[0202] According to the system described in Clause 34, wherein: The aforementioned features correspond to priority agents; The one or more displays are configured to display at least one of the following: (i) a representation of the sensor data or (ii) the model, and the one or more displays are configured to display an indication of the feature, including: Cause the one or more displays to display the representation of at least the sensor data; and This causes the one or more displays to show an indication of the priority agent within the representation of the sensor data; and The data sent from the system to the vehicle includes data associated with the priority agent.

[0203] The system according to any one of clauses 33 to 36, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; and The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model comprising: The model is displayed on one or more displays such that the model has a viewing angle that substantially corresponds to the viewing angle of the representation of the sensor data.

[0204] According to the system described in Clause 33, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; Receiving the user input at the system includes receiving at least a first user input, the first user input being associated with the feature and a first orientation within the displayed representation or displayed model of the sensor data on the one or more displays; Receiving the first user input enables determining the location of the feature within the sensor data or the model based on the first orientation; and Indications that cause the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays include: At least in part, based on whether the first orientation is within the displayed representation or the displayed model of the sensor data, the one or more displays display the indication of the feature at the orientation based on the first orientation in either the displayed representation or the displayed model of the sensor data.

[0205] According to the system described in Clause 38, wherein: The feature corresponds to the vehicle's navigation path; The first user input corresponds to the modification or generation of the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

[0206] According to the system described in Clause 38, wherein: The feature corresponds to the predicted path associated with the object within the environment; and The first user input corresponds to the selection of the object visible in either the displayed representation or the displayed model of the sensor data.

[0207] According to the system described in Clause 38, wherein: Receiving the user input at the system includes receiving at least the first user input and the second user input, the second user input corresponding to one of the following: Driving instructions for the vehicle; or The classification of the features; and The data sent by the system to the vehicle includes one of the following: The driving instructions; or The classification.

[0208] A system according to any one of clauses 33 to 41, wherein the one or more displays are configured to display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to show the representation of at least the sensor data; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the model; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the location within the representation of the sensor data.

[0209] A system according to any one of clauses 33 to 41, wherein the one or more displays are configured to display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to display at least the model; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the sensor data; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the specified location within the model.

[0210] In any one of the provisions 33 to 43, the sensor data is video frame data.

[0211] According to the system described in Clause 33, wherein: The feature is visible within the displayed representation of the sensor data; Instructions that cause the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays include instructions that cause the one or more displays to display the feature at the orientation in the model; The operation also includes generating additional output data for use by the vehicle's planning component to navigate in the environment, wherein the additional output data is associated with the feature; and The data sent by the system to the vehicle includes the additional output data, and the additional output data includes: Data associated with the aforementioned feature; and The feature is located at the specified location within the environment.

[0212] According to the system described in Clause 33, wherein: The sensor data and the model are associated with a coordinate system; Determining the position of a feature within the sensor data or the model includes determining the position of the feature within the coordinate system.

[0213] One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a system, cause the system to perform the method according to any one of clauses 26 to 32.

[0214] A system comprising: One or more processors; and One or more non-transitory computer-readable media having instructions stored thereon, which, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of clauses 26 to 32.

[0215] According to the method described in Clause 8, wherein: The output data includes map data corresponding to map features; The features correspond to map features; Determining the location of the feature within the sensor data or the model includes determining the location of the feature within the model; and Instructions that cause the one or more displays to display the feature at the orientation corresponding to the location on the one or more displays include instructions that cause the one or more displays to display the feature at the orientation in at least the representation of the sensor data.

[0216] According to the method described in Clause 8, wherein: The output data includes route data corresponding to the vehicle's navigation path or route; The feature corresponds to the navigation path or the channel; Determining the location of the feature within the sensor data or the model includes determining the location of the feature within the model; and Instructions that cause the one or more displays to display the feature at the orientation corresponding to the location on the one or more displays include instructions that cause the one or more displays to display the feature at the orientation in at least the representation of the sensor data.

[0217] While the example clauses described above pertain to a particular implementation, it should be understood that, within the context of this document, the content of the example clauses may also be implemented by means of methods, apparatus, systems, computer-readable media, and / or other implementations. Furthermore, any one of example clauses 1 to 50 may be implemented alone or in combination with any other one or more of the example clauses.

[0218] in conclusion While one or more examples of the techniques described herein have been described, various modifications, additions, arrangements, and equivalents thereof are included within the scope of the techniques described herein.

[0219] In the description of the examples, reference is made to the accompanying drawings, which form part of this document, illustrating specific examples of the claimed subject matter by way of illustration. It should be understood that other examples may be used, and variations or modifications, such as structural changes, may be made. Such examples, variations, or modifications do not necessarily deviate from the scope of the established claimed subject matter. While the steps in this document may be presented in a certain order, in some cases the order may be changed so that certain inputs are provided at different times or in a different order without altering the functionality of the described system and method. The disclosed procedures may also be performed in different orders. Furthermore, the various calculations in this document need not be performed in the disclosed order, and other examples using different orders of calculation can be readily implemented. Besides reordering, these calculations may also be decomposed into sub-computations, yielding the same results.

[0220] Although the subject matter has been described in language specific to structural features and / or methodological actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described. Rather, specific features and actions are disclosed as exemplary forms of implementing the claims.

[0221] The components described herein represent instructions that can be stored in any type of computer-readable medium and can be implemented in software and / or hardware. All the methods and processes described above can be embodied in software code components and / or computer-executable instructions executed by one or more computers or processors, hardware, or some combination thereof, and are fully automated via these software code components and / or computer-executable instructions. Some or all of the methods described may alternatively be embodied in dedicated computer hardware.

[0222] At least some of the processes discussed herein are illustrated as logic flowcharts, where each operation represents a sequence of operations that can be implemented in hardware, software, or a combination thereof. In the context of software, an operation represents computer-executable instructions stored on one or more non-transitory computer-readable storage media, which, when executed by one or more processors, cause a computer or autonomous vehicle to perform the operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific abstract data type. The order in which operations are described is not intended to be restrictive, and any number of the operations can be combined in any order and / or in parallel to implement the process.

[0223] Unless otherwise stated, conditional languages ​​such as “can,” “able,” “may,” or “possibly” are understood in context to imply that some examples include certain features, elements, and / or steps that are not included in other examples. Therefore, such conditional languages ​​are generally not intended to imply that certain features, elements, and / or steps are required in any way for one or more examples, or that one or more examples must include logic for determining whether or not certain features, elements, and / or steps are included or will be performed in any particular example, with or without user input or prompts.

[0224] Unless otherwise stated, conjunctive language such as the phrase “at least one of X, Y, or Z” should be understood to mean that an item, term, etc., can be X, Y, or Z, or any combination thereof, including multiples of each element. Unless explicitly stated as singular, “a” means both singular and plural.

[0225] Any routine description, element, or block depicted in the flowcharts described herein and / or in the accompanying drawings should be understood as potentially representing a module, segment, or portion of code comprising one or more computer-executable instructions for implementing a particular logical function or element in the routine. Alternative implementations are included within the scope of the examples described herein, where elements or functions may be removed or performed in a non-shown or non-discussed order, including substantially synchronously, in reverse order, with appended operations, or with omitted operations, depending on the functionality involved, as will be understood by those skilled in the art. Note that the term can substantially indicate scope. For example, substantially simultaneous can indicate that two activities occur within each other's time scope, substantially the same dimension can indicate that two elements have dimensions within each other's scope, etc.

[0226] Many variations and modifications can be made to the above examples, and their elements should be understood as existing in other acceptable examples. All such modifications and variations are intended to be included within the scope of this disclosure and protected by the appended claims.

Claims

1. A system comprising: One or more processors; as well as One or more non-transitory computer-readable media having instructions stored thereon, the instructions, when executed by the one or more processors, causing the one or more processors to perform operations, the operations including: Received from the autonomous vehicle at the system: Video frame data captured by the vehicle's video camera; and Perception output data generated by the vehicle, wherein the perception output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components for navigation in the environment, wherein the video frame data and the perception output data indicate the environment; This causes a representation of the video frame data to be displayed in a first area of ​​one or more displays; The second area of ​​the one or more displays displays a model of the environment based on the perceived output data; The system receives user input, which is associated with a feature within one of the first or second regions. Determine a first orientation associated with the feature within the first region or the second region; The second orientation of the feature within the other of the first region or the second region is determined at least in part based on the first orientation; At least in part based on whether the first orientation is within the first or second region of the one or more displays, the one or more displays display an indication of the feature at the second orientation in either the first or second region; and The system sends data, at least in part based on the user input, to the vehicle to cause the vehicle to take action.

2. The system according to claim 1, wherein: The feature corresponds to the vehicle's navigation path; The user input includes user input corresponding to modifying or generating the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

3. The system according to claim 1, wherein: The feature is visible within the displayed representation of the video frame data; Instructions to cause the one or more displays to display the feature at the second orientation in either the first region or the second region include instructions to cause the one or more displays to display the feature at the second location in the second region; The operation also includes generating additional perception output data for use by the vehicle's planning component to navigate in the environment, wherein the additional perception output data is associated with the feature; and The data sent by the system to the vehicle includes the additional sensing output data, and the additional sensing output data includes: Data associated with the aforementioned feature; as well as The feature is located within the environment, the location being based on the first orientation.

4. The system according to claim 1, wherein: The feature corresponds to an object within the environment; Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second regions of the one or more displays, and the second user input corresponding to a classification of the feature; and The data sent by the system to the vehicle includes the classification.

5. The system according to claim 1, wherein: Receiving the user input at the system includes receiving at least a first user input and a second user input, the first user input being associated with a first orientation within one of the first or second areas of the one or more displays, and the second user input corresponding to a driving instruction for the vehicle; and The data sent from the system to the vehicle includes driving instructions for the vehicle to navigate in the environment.

6. The system according to claim 1, wherein: The video frame data and the model are associated with a coordinate system; Determining the second orientation of the feature within the other of the first region or the second region, at least in part, based on the first orientation, includes: The position of the feature in the coordinate system is determined based on the first orientation. as well as The second orientation is determined based on the stated location.

7. A method, the method comprising: Received from the autonomous vehicle at the system level: Sensor data captured by the vehicle's sensors; as well as Output data, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components to navigate in the environment, wherein the sensor data and the output data indicate the environment; The display shall cause one or more displays to show at least one of the following: (i) a representation of the sensor data or (ii) a model of the environment based on the output data; Determine the location of the feature within the sensor data or the model; This causes the one or more displays to show an indication of the feature at a position corresponding to the location on the one or more displays; The system receives user input. as well as The system sends data based on the user input to the vehicle so that the vehicle can take action.

8. The method of claim 7, wherein determining the location of the feature within the sensor data or the model comprises automatically determining the location of the feature within the sensor data or the model.

9. The method of claim 8, wherein the feature is a first feature of a plurality of features in the environment, and wherein the method comprises: Determine that at least the first feature among the plurality of features does not exist in either the sensor data or the model; as well as Based on the determination that at least the first feature among the plurality of features does not exist in either the sensor data or the model: This causes the one or more displays to show the indication of the feature.

10. The method according to claim 8, wherein: The aforementioned features correspond to priority agents; The one or more displays are configured to display at least one of the following: (i) a representation of the sensor data or (ii) the model, and the one or more displays are configured to display an indication of the feature, including: Cause the one or more displays to display the representation of at least the sensor data; and This causes the one or more displays to show an indication of the priority agent within the representation of the sensor data; and The data sent from the system to the vehicle includes data associated with the priority agent.

11. The method according to claim 7, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; and The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model comprising: The model is displayed on one or more displays such that the model has a viewing angle that substantially corresponds to the viewing angle of the representation of the sensor data.

12. The method according to claim 7, wherein: The one or more displays shall display at least one of the following: (i) the representation of the sensor data or (ii) the model comprising: The one or more displays shall display both of the following: (i) the representation of the sensor data and (ii) the model; Receiving the user input at the system includes receiving at least a first user input, the first user input being associated with the feature and a first orientation within the displayed representation or displayed model of the sensor data on the one or more displays; Receiving the first user input enables determining the location of the feature within the sensor data or the model based on the first orientation; and Indications that cause the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays include: At least in part, based on whether the first orientation is within the displayed representation or the displayed model of the sensor data, the one or more displays display the indication of the feature at the orientation based on the first orientation in either the displayed representation or the displayed model of the sensor data.

13. The method according to claim 12, wherein: The feature corresponds to the vehicle's navigation path; The first user input corresponds to the modification or generation of the navigation path for the vehicle; and The data sent from the system to the vehicle includes data associated with the navigation path so that the vehicle can follow the navigation path.

14. The method of claim 12, wherein: The feature corresponds to the predicted path associated with the object within the environment; and The first user input corresponds to the selection of the object visible in either the displayed representation or the displayed model of the sensor data.

15. The method according to claim 12, wherein: Receiving the user input at the system includes receiving at least the first user input and a second user input, the second user input corresponding to one of the following: Driving instructions for the vehicle; or The classification of the features; and The data sent by the system to the vehicle includes one of the following: The driving instructions; or The classification.

16. The method of claim 7, wherein the one or more displays are made to display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to show the representation of at least the sensor data; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the model; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the location within the representation of the sensor data.

17. The method of claim 7, wherein the one or more displays are made to display at least one of: (i) the representation of the sensor data or (ii) the model comprising: This causes the one or more displays to display at least the model; Determining the location of the feature within the sensor data or the model includes: Determine the location of the feature within the sensor data; and The indication that causes the one or more displays to display the feature at an orientation corresponding to the location on the one or more displays includes: This causes the one or more displays to show the indication of the feature at the specified location within the model.

18. The method of claim 7, wherein the sensor data is video frame data.

19. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of a system, cause the system to perform operations, the operations including: Received from the autonomous vehicle at the system: Sensor data captured by the vehicle's sensors; as well as Output data, wherein the output data is generated at least in part based on environmental data captured by one or more sensors of the vehicle and used by the vehicle's planning components to navigate in the environment, wherein the sensor data and the output data indicate the environment; The display shall cause one or more displays to show at least one of the following: (i) a representation of the sensor data or (ii) a model of the environment based on the output data; Determine the location of the feature within the sensor data or the model; This causes the one or more displays to show an indication of the feature at a position corresponding to the location on the one or more displays; The system receives user input. as well as The system sends data based on the user input to the vehicle so that the vehicle can take action.

20. One or more non-transitory computer-readable media according to claim 18, wherein at least one of the following: Determining the location of the feature within the sensor data or the model includes automatically determining the location of the feature within the sensor data or the model; and Receiving the user input at the system includes receiving at least a first user input associated with the feature and a first orientation within the displayed representation or model of the sensor data on the one or more displays, wherein receiving the first user input causes the position of the feature within the sensor data or the model to be determined based on the first orientation.

Citation Information

Patent Citations

  • Collaborative action ambiguity resolution for autonomous vehicles

    US12005925B1

  • Data Segmentation Using Masks

    US20190332118A1

  • Teleoperations for collaborative vehicle guidance

    US20210323573A1