System for Scene-Aware Interaction

By using multimodal sensors to perceive static and dynamic objects around the vehicle in real time and generating context-aware driving instructions, the problem of inaccurate navigation in complex environments is solved, and the safety of the navigation system and the driver's navigation experience are improved.

CN115038936BActive Publication Date: 2025-07-08MITSUBISHI ELECTRIC CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080095350.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-02-06
Filing Date
2020-12-17
Publication Date
2025-07-08
Estimated Expiration
2040-12-17

AI Technical Summary

Technical Problem

When providing route guidance, existing navigation systems fail to effectively utilize real-time information of static and dynamic objects near the vehicle, making it difficult for drivers to identify the correct steering in complex environments, increasing safety risks.

Method used

By combining multimodal sensor information, such as cameras, microphones, LiDAR and GPS, static and dynamic objects around the vehicle are sensed in real time, context-based driving instructions are generated, and a scene-aware interaction system is used to provide intuitive route guidance.

Benefits of technology

Improves the accuracy of the navigation system and driver safety, and reduces misleading and distraction by providing clear and intuitive driving instructions, and enhances drivers' navigation capabilities in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115038936B_ABST
    Figure CN115038936B_ABST
Patent Text Reader

Abstract

A navigation system is provided that is configured to provide driving instructions to a driver of a moving vehicle based on a real-time description of objects in a scene that are relevant to the driving vehicle. The navigation system includes: an input interface configured to receive a dynamic map for a route for driving the vehicle, a state of the vehicle on the route at the current moment, and a set of significant objects related to the vehicle's route at the current moment, wherein at least one significant object is an object sensed by a measurement system of the vehicle moving on a route between a current position at the current moment and a future position at a future moment, and wherein the set of significant objects includes one or more static objects and one or more dynamic objects; a processor configured to generate driving instructions based on a description of significant objects in the dynamic map derived from a driver's perspective specified by the state of the vehicle; and an output interface configured to present the driving instructions to the driver of the vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to systems for providing scene-aware interaction systems, and more particularly to a scene-aware interaction navigation system for providing route guidance to a driver of a vehicle based on real-time unimodal or multimodal information about static and dynamic objects near the vehicle. Background Art

[0002] Navigation assistance for a driver operating a vehicle is typically provided by a system such as a GPS receiver, which can provide voice route guidance to the driver. The route guidance takes the form of steering instructions, most commonly indicating the distance to a turning point, the direction of the turn, and possibly some additional information to clarify where to turn, such as "turn right into Johnson Street at the second intersection within 100 feet". However, in some cases, this method of providing route guidance to the driver may confuse the driver, for example, when the driver does not know and cannot easily identify that the name of the street to turn onto is "Johnson Street", or when there are multiple streets and paths in close proximity. Then, the driver may not be able to identify the correct street to turn onto, miss the turn, become confused, and potentially cause a dangerous situation.

[0003] Alternatively, there are route guidance systems that can use stored point-of-interest-related information from a map to indicate a turning point, such as "turn within 100 feet at the post office". However, in some cases, this method may confuse the driver, for example when a tree or vehicle obscures the post office or makes it difficult to identify, or when the stored information is outdated and there is no longer a post office at that turning point.

[0004] Alternatively, there are experimental route guidance systems that can accept real-time camera images captured by the driver and overlay graphical elements, such as arrows, indicating a specific route to follow on the real-time image. However, this method does not provide spoken descriptive statements and requires the driver to take their eyes off the road to see the route guidance. Summary of the Invention

[0005] Scene-aware interaction systems can be applied to a variety of applications, such as in-vehicle infotainment and home appliances, interaction with service robots in building systems, and monitoring systems. GPS is just one positioning method for a navigation system, and other positioning methods can be used for other applications. Hereinafter, the navigation system is described as one example application of scene-aware interaction.

[0006] At least one insight of the present disclosure is that existing methods are different from the guidance that a hypothetical passenger who knows where to turn would provide to a driver. A passenger who knows the route and provides guidance to the driver generally does not consider both static and dynamic objects to formulate driving instructions that they consider to be the most intuitive, natural, relevant, easy to understand, clear, etc., in order to help the driver safely follow the intended route.

[0007] At least one other realization of the present disclosure is that existing methods do not utilize real-time information about dynamic objects (such as other vehicles) near the vehicle to identify reference points for providing route guidance. At least one other realization of the present disclosure is that existing methods do not utilize real-time information to consider the current situation so that the driver can easily identify the current situation (such as other objects that obstruct the view, such as vehicles or trees), and the current situation may change or affect the appropriate way to describe static objects near the vehicle. The appearance of the static object is different from the appearance of the static object stored in the static database, for example, due to construction or renovation, or directly because the static object no longer exists, thus making it irrelevant to the reference point for providing route guidance.

[0008] The purpose of some embodiments is to provide route guidance to a vehicle driver based on real-time unimodal or multimodal information about static and dynamic objects near the vehicle. For example, the purpose of some embodiments is to provide context-based driving instructions, such as "turn right before the brown brick building" or "follow the white car", as a supplement or alternative to GPS-based instructions such as "turn right onto Johnson Street at the second intersection within 100 feet". Such context-based driving instructions can be generated based on the real-time perception of the scene near the vehicle. For this purpose, context-based navigation is referred to as scene-aware navigation herein.

[0009] Some embodiments are based on the understanding that at different time points, different numbers or types of objects can be relevant to the route for driving a vehicle. All of these relevant objects are potentially useful for scene-aware navigation. However, compared to autonomous driving where driving decisions are made by a computer, when driving instructions are made for too many different objects or objects that may not be easily recognizable by a human driver, the human driver may be confused and / or distracted. Therefore, since different objects can be more or less relevant to context-based driving instructions, the purpose of some embodiments is to select objects from a set of significant objects relevant to the driver's route and generate driving instructions based on the description of the significant object.

[0010] The route guidance system of the present invention can receive information from multiple sources, including static maps, planned routes, the current position of the vehicle determined by GPS or other methods, and real-time sensor information from a series of sensors, the series of sensors including but not limited to one or more cameras, one or more microphones, and one or more distance detectors including radar and LiDAR. The real-time sensor information is processed by a processor capable of detecting from the real-time sensor information a set of significant static and dynamic objects near the vehicle and a set of object attributes, the set of object attributes including, for example: the category of each object, such as a car, a truck, a building; and the color, size, and position of the object. For dynamic objects, the processor can also determine the trajectory of the dynamic object. In the case of sound information obtained by a microphone, the processor can detect the category of the object by identifying the type of the sound, and the object attributes can include the direction and distance of the object from the vehicle, the movement trajectory of the object, and the intensity of the sound. The set of significant objects and their set of attributes are hereinafter referred to as the dynamic map.

[0011] The route guidance system uses many methods, such as rule-based methods or machine learning-based methods, to process the dynamic map in order to identify significant objects from the set of significant objects based on the route for use as selected significant objects for providing route guidance.

[0012] The conveyance of route guidance information can include highlighting significant objects using bounding rectangles or other graphical elements on a display (e.g., an LCD display in an instrument cluster or a central console). Alternatively, the method of conveyance can include using, for example, rule-based methods or machine learning-based methods to generate statements that include a set of descriptive attributes of the significant objects. The generated statements can be conveyed to the driver on the display. Alternatively, the generated statements can be converted into spoken sounds that the driver can hear through speech synthesis.

[0013] Another object of the present invention is that significant objects can be determined by considering the distance of the vehicle from the route turning point. In particular, multiple different significant objects can be selected in various distance ranges such that in each distance range, the selected significant objects provide the driver with the maximum information about the planned route. For example, at a long distance from the turning point, large static objects such as buildings approaching the turning point can be determined as significant objects because the turning point cannot be clearly seen yet, while at a short distance from the turning point, dynamic objects such as another vehicle that has traveled along the planned route can be determined as significant objects because it can be clearly seen and is distinctive enough for route guidance.

[0014] Another object of the present invention is that, according to a planned route, the route guidance system of the present invention can provide descriptive warnings about other objects near the vehicle in a certain form. For example, if the next step of the planned route is a turn and the route guidance system detects an obstacle on the planned route, a descriptive warning message can be transmitted to the driver to warn them of the presence of the object. More specifically, if a person is crossing or appears to be about to cross the street at a point near the vehicle along the planned route, the route guidance system can provide a descriptive warning message. For example, the route guidance system can generate and speak out a statement such as: "Warning, there is someone on the crosswalk on your left."

[0015] Another object of the present invention is to provide the possibility of two-way interaction between the driver and the route guidance system of the present invention, which enables the driver to seek clarification regarding the location, attributes or other information of significant objects and request different significant objects. The two-way interaction can include one or more interaction mechanisms, and the interaction mechanisms include spoken dialogue, where an automatic speech recognizer enables the route guidance system to obtain the text uttered by the driver, so that the system can process the text to understand and adapt to the driver's response to the system. The interaction can also include information captured by one or more cameras, which receive images of the driver and are input into a computer vision subsystem, and the computer vision subsystem can extract information about the driver, including but not limited to the driver's posture, such as the direction pointed by the driver's hand or the direction of the driver's gaze. The interaction can also include manual input from the driver, including pressing one or more control buttons, which can be arranged in a manner accessible to the driver, such as on the steering wheel, instrument cluster or center console.

[0016] The above problems are solved by the subject matter according to the independent claims. According to some embodiments, the navigation system is configured to provide driving instructions to the driver of the vehicle based on a real-time description of objects related to the driving vehicle in a scene. The navigation system may include: an input interface configured to accept a route for driving the vehicle, the state of the vehicle on the route at the current moment, and a dynamic map of a set of significant objects related to the route of the vehicle at the current moment, wherein at least one significant object is an object sensed by a measurement system of the vehicle moving on the route between the current position at the current moment and the future position at a future moment, and wherein the set of significant objects includes one or more static objects and one or more dynamic objects; a processor configured to generate driving instructions based on a description of significant objects in the dynamic map derived from the driver's perspective specified by the state of the vehicle; and an output interface configured to present the driving instructions to the driver of the vehicle.

[0017] Some embodiments of the present disclosure are based on the recognition that scenario-aware interaction with a user (operator) can be performed based on attention multi-modal fusion, which analyzes multi-modal sensing information and provides more natural and intuitive interaction with humans through context-dependent natural language generation.

[0018] In some cases, the multi-modal sensing information can be image / video captured by a camera, audio information obtained by a microphone, and position information estimated by a distance sensor such as LiDAR or radar.

[0019] Integrating attention multi-modal fusion into scene understanding technology and context-based natural language generation enables a powerful scene-aware interaction system to interact with users more intuitively based on objects and events in the scene. Scene-aware interaction technology can be widely applied to a variety of applications, including the human-machine interface (HMI) of in-vehicle infotainment and household appliances, interaction with service robots in building systems, and monitoring systems.

[0020] The presently disclosed embodiments will be further explained with reference to the accompanying drawings. The drawings shown are not necessarily to scale, but generally focus on illustrating the principles of the presently disclosed embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1A

[0022] Figure 1A A block diagram and illustration of a navigation system according to some embodiments of the present disclosure are shown;

[0023] Figure 1B

[0024] Figure 1B A block diagram and illustration of a navigation system according to some embodiments of the present disclosure are shown;

[0025] Figure 1C

[0026] Figure 1C A block diagram and illustration of a navigation system according to some embodiments of the present disclosure are shown;

[0027] Figure 1D

[0028] Figure 1D A block diagram and illustration of a navigation system according to some embodiments of the present disclosure are shown;

[0029] Figure 2

[0030] Figure 2 ​​​​​​​​​​Schematic diagram of a route guidance system according to some embodiments of the present disclosure, which shows the information flow from the external scene near the vehicle to the output of driving instructions;

[0031] Figure 3

[0032] Figure 3 Block diagram of a computer that receives inputs from multiple sources and sensors and outputs information to a display or speaker according to some embodiments of the present disclosure;

[0033] Figure 4

[0034] Figure 4 Block diagram showing a multi-modal attention method according to an embodiment of the present disclosure;

[0035] Figure 5

[0036] Figure 5 Block diagram showing an example of a multi-modal fusion method (multi-modal feature fusion method) for sentence generation according to an embodiment of the present disclosure;

[0037] Figure 6A

[0038] Figure 6A Flowchart showing the training of a parameter function of a navigation system according to an embodiment of the present disclosure, the parameter function being configured to generate driving instructions based on the state of the vehicle and a dynamic map;

[0039] Figure 6B

[0040] Figure 6B Flowchart showing the training of a parameter function of a navigation system according to an embodiment of the present disclosure, a first parameter function being configured to determine the attributes and spatial relationships of a set of significant objects of a dynamic map based on the state of the vehicle to obtain a transformed dynamic map, and a second parameter function being configured to generate driving instructions based on the transformed dynamic map;

[0041] Figure 6C

[0042] Figure 6C Flowchart showing the training of a parameter function of a navigation system according to an embodiment of the present disclosure, a first parameter function being configured to determine the state of the vehicle and a dynamic map based on measurements from a scene, and a second parameter function being configured to generate driving instructions based on the vehicle state and the dynamic map;

[0043] Figure 6D

[0044] Figure 6D ​​​​​​​​​​​​​​is a flowchart showing end-to-end training of a parameter function of a navigation system according to an embodiment of the present disclosure, the navigation system being configured to generate driving instructions based on measurements from a scene;

[0045] Figure 6E

[0046] Figure 6E is a flowchart showing training of a parameter function of a navigation system according to an embodiment of the present disclosure, a first parameter function being configured to determine a state of a vehicle and a dynamic map based on measurements from a scene, a second parameter function being configured to determine attributes and spatial relationships of a set of prominent objects of the dynamic map based on the state of the vehicle to obtain a transformed dynamic map, a third parameter function being configured to select a subset of prominent objects from the transformed dynamic map, and a fourth parameter function being configured to generate driving instructions based on the selected prominent objects;

[0047] Figure 6F

[0048] Figure 6F is a flowchart showing multi-task training of a parameter function of a navigation system according to an embodiment of the present disclosure, a first parameter function being configured to determine a state of a vehicle and a dynamic map based on measurements from a scene, a second parameter function being configured to determine attributes and spatial relationships of a set of prominent objects of the dynamic map based on the state of the vehicle to obtain a transformed dynamic map, a third parameter function being configured to select a subset of prominent objects from the transformed dynamic map, and a fourth parameter function being configured to generate driving instructions based on the selected prominent objects;

[0049] Figure 7

[0050] Figure 7 shows example prominent objects in a dynamic map according to some embodiments of the present disclosure, as well as attributes of the objects and values of these attributes;

[0051] Figure 8

[0052] Figure 8 shows a set of prominent objects and their relative spatial relationships according to some embodiments of the present disclosure;

[0053] Figure 9

[0054] Figure 9 shows a set of prominent objects and corresponding relevance scores for generating route guidance statements at different time instances according to some embodiments of the present disclosure;

[0055] Figure 10 ​​​​​​​​​​​​

[0056] Figure 10 Shows a set of significant objects according to some embodiments of the present disclosure and corresponding relevance scores for generating route guidance statements at different time instances;

[0057] Figure 11

[0058] Figure 11 Shows an example of a conversation between a route guidance system and a driver according to some embodiments of the present disclosure;

[0059] Figure 12

[0060] Figure 12 Is a flowchart of a specific embodiment of a route guidance system according to some embodiments of the present disclosure, the route guidance system using a rule-based object sorter within a statement generator.

[0061] Although the above figures illustrate embodiments of the present disclosure currently, as noted in the discussion, other embodiments are also contemplated. The present disclosure presents illustrative embodiments in a presenting rather than limiting manner. Those skilled in the art can design many other modifications and embodiments that fall within the scope and spirit of the principles of the embodiments of the present disclosure currently. Detailed Description

[0062] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with a description for enabling one or more exemplary embodiments to be implemented. It is contemplated that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the disclosed subject matter set forth in the appended claims.

[0063] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, those of ordinary skill in the art can understand that the embodiments can be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter can be shown as components in block diagrams so as not to obscure the embodiments with unnecessary details. In other examples, well-known processes, structures, and technologies can be shown without unnecessary details so as not to obscure the embodiments. Additionally, the same reference numerals and markings in the various figures indicate the same elements.

[0064] ​​​​Moreover, each embodiment can be described as a process depicted as a flow chart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flow chart may describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Additionally, the order of the operations can be rearranged. A process can terminate when its operations are completed, but can have additional steps not discussed or included in the figure. Furthermore, not all operations in any particular described process will occur in all embodiments. A process can correspond to a method, a function, a program, a subroutine, a subprogram, etc. When a process corresponds to a function, the termination of the function can correspond to the function returning to the calling function or the main function.

[0065] In addition, embodiments of the disclosed subject matter can be implemented, at least in part, manually or automatically. Execution or at least assistance in implementing embodiments can be performed by using a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, program code or code segments for performing the necessary tasks can be stored in a machine-readable medium. A processor can perform the necessary tasks.

[0066] Figures 1A through 1D A block diagram and illustration of a navigation system in accordance with some embodiments of the present disclosure are shown. In some instances, a navigation system can be referred to as a route guidance system, and a route guidance system can be referred to as a navigation system.

[0067] Figure 1A A block diagram of a navigation system showing features of some embodiments is presented. A set of salient objects in a dynamic map can be identified and described based on sensing information sensed by a measurement system 160 of a vehicle, the sensing information including information from one or more modalities, such as audio information from a microphone 161, visual information from a camera 162, depth information from a distance sensor (i.e., a depth sensor) such as a LiDAR 163, and positioning information from a global positioning system (GPS) 164. The system outputs a driving instruction 105 based on a description of one or more of the salient objects from the set of salient objects. In some embodiments, a processor generates the driving instruction 105 by submitting measurements from the measurement system 160 to a parameter function 170 that has been trained to generate driving instructions based on the measurements. In other embodiments, multi-modal sensing information obtained by the measurement system is used to determine the state of the vehicle (which we also refer to as the vehicle state in this document) and the dynamic map. The processor is configured to submit the state of the vehicle and the dynamic map to the parameter function 170, which is configured to generate the driving instruction 105 based on a description of salient objects in the dynamic map derived from the driver's perspective specified by the state of the vehicle.

[0068] Figure 1B , Figure 1C and Figure 1D illustrate diagrams of a navigation system according to some embodiments of the present invention. The system has obtained a route for driving a vehicle and has information about the state of the vehicle on the driving route 110 at the current moment. It should be understood that the route consists of a series of segments and turns, where each segment has a determined length and position, and each turn is in a specific direction leading from one segment or turn to another segment or turn. In some embodiments, the segments and turns are parts of a road that are connected to provide a path that a vehicle can follow to travel from one position to another. The route is represented by the part of the driving route 110 that will soon be traversed by the vehicle, as indicated by the arrow overlaid on the road. In some embodiments, the state of the vehicle includes the position and orientation of the vehicle relative to a dynamic map that contains a set of significant objects related to driving the vehicle on the route. The significant objects include one or more static objects (i.e., objects that are always stationary), such as buildings 130, signs 140, or mailboxes 102, and one or more dynamic objects (i.e., objects with the ability to move), such as other vehicles 120, 125, or pedestrians 106. In some embodiments, a dynamic object that is not currently moving but has the ability to move (e.g., a parked car or a pedestrian standing still currently) is considered a dynamic object (although the speed is equal to zero). The system includes a processor configured to generate driving instructions 105 that are presented to the driver of the vehicle via an output interface such as a speech synthesis system 150.

[0069] In some embodiments, the driving instructions include significant objects (102, 125, 126) in the dynamic map derived from the driver's perspective specified by the state of the vehicle. For example, in Figure 1B , the driving instruction 105 is based on the description "red mailbox" of the significant object 102, and the significant object 102 is selected from the set of significant objects in the dynamic map based on the driver's perspective. In some embodiments, the driver's perspective includes the current position of the vehicle relative to the dynamic map and the part of the route 110 that is relevant based on the current position and orientation of the vehicle. For example, the "red mailbox" part is selected because it is in the direction of the upcoming turn in the route 110. In an alternative scenario where the route 110 is a left turn ( Figure 1B not shown in the figure), the driving instruction will be based on a different object 130, and the description "blue building" of the different object 130 is used in a driving instruction such as "turn left before the blue building" because from the perspective of the driver who is about to turn left, the blue building 130 is more relevant than the red mailbox 102.

[0070] In Figure 1CIn this case, the driving instruction 105 is based on the description of the prominent object 125 in the dynamic map as "a silver car turning right". In Figure 1D In this case, the driving instruction 105 is a warning based on the description of a set of prominent objects (pedestrian 106 and crosswalk) in the dynamic map as "a pedestrian in the crosswalk". From the driver's perspective, these objects are important because they are visible to the driver of the vehicle and they are on the next part of the vehicle's route 110.

[0071] Figure 2 is a schematic diagram of a proposed route guidance system showing the information flow from the external scene 201 near the vehicle to the output driving instruction 213. The vehicle is equipped with multiple real-time sensor modalities that provide real-time sensor information 203 to the route guidance system. The object detection and classification module 204 uses a parametric function to process the real-time sensor information 203 in order to extract information about the objects near the vehicle, where the information about the objects near the vehicle includes their positions relative to the vehicle and their category types, and the category types include at least buildings, cars, trucks, and pedestrians. The object attribute extraction module 205 performs additional operations to extract a set of object attributes for each detected object, where the set of attributes includes at least color, distance from the vehicle, and size, and for some specific categories of objects may also include trajectory information such as speed and direction of movement. Those of ordinary skill in the art should understand that for each different category of object, there may be different sets of attributes. For example, a truck may have truck-type attributes, which may take a value from, for example, van, semi-trailer, dump, etc. as needed, so that the statement generator 212 can generate highly descriptive driving instruction statements 213 for route guidance. The dynamic map 206 receives information from the object detection and classification module 204, the object attribute extraction module 205, as well as the planned driving route 202 and the viewing volume 211 determined by the vehicle state 209. The dynamic map 206 uses the driving route information 202 to identify a subset of the detected objects that are prominent with respect to the planned route. Prominent objects are those that are relevant to the driving route, for example, by being at the same corner as a route turning point or just past the turning point on the planned route. The dynamic map consists of a set of prominent static and dynamic objects, including their category types and attributes, which are candidate objects to be used for providing route guidance to the driver.

[0072] The vehicle state 209 includes one or a combination of vehicle position, speed, and orientation. In some embodiments, the driver's perspective 210 is the observation position of a given driver at the seat height and within the angular range (e.g., + / - 60 degrees around the front direction of the vehicle) that the driver can reasonably see without excessive head movement. The driver's perspective 210 is used to determine the visibility volume 211, which is a subset of the space that the driver can see. This is useful because one or more real-time sensors can be mounted on the vehicle in such a way that they can see objects that the driver cannot. For example, a LiDAR mounted on the roof may be able to detect a first object that is beyond a second, closer object, but the first object is occluded by the second object when viewed from the driver's perspective 210. This makes the first object not a suitable prominent object at that moment because it cannot be seen. Thus, the dynamic map 206 can also use the visibility volume 211 to determine the set of prominent objects. Alternatively, prominent objects that cannot be seen from the driver's perspective 210 may be important for providing driving instruction statements 213. For example, an ambulance may approach from behind the vehicle and thus be hidden from the driver's direct view. The statement generation module 212 can generate a driving instruction statement 213 that provides a warning to the driver about the approaching ambulance. It should be understood that the dynamic map is continuously updated based on the real-time sensor information 203, and the state of the visibility volume 211 can change at any time.

[0073] The statement generation module 212 performs the operation of generating a driving instruction statement 213 given the driving route 202, the visibility volume 211, and the dynamic map 206. The statement generation module 212 uses a parametric function to select a small subset of the most prominent objects from the set of static prominent objects 207 and dynamic prominent objects 208 in the dynamic map 206 for generating the driving instruction statement 213. Broadly speaking, the most prominent objects tend to be larger and have a more distinctive color or position so that the driver can quickly observe it.

[0074] The statement generation module 212 can be implemented by multiple different parametric functions. A possible parametric function for implementing the statement generation module 212 is by using template-based driving commands, also simply referred to as driving commands. An example of a template-based driving command is "Follow the <attribute> <significant object> that turns <direction> ahead". In the foregoing example, <attribute>, <significant object>, and <direction> are template slots that the statement generation module 212 fills to generate the driving instruction statement 213. In this case, <attribute> is one or more attributes of the significant object, and <direction> is the next turning direction in the driving route 202. A specific example of this type of template-based driving command is "Follow the large brown van that turns left ahead". In this specific example, "large", "brown", and "van" are the attributes of the "truck" that has "turned left" in the same direction as the next turning direction of the driving route 202. Many possible template-based driving commands are feasible, including for example "Turn <direction> before the <attribute> <significant object>", "Turn <direction> after the <attribute> <significant object>", "Merge into <direction>", "Drive towards the <attribute> <significant object>", "Stop at the <attribute> <significant object>", "Park near the <attribute> <significant object>". The use of the words "before", "after", "near" indicates the relative spatial relationship between the significant object and the route. For example, "Turn right before the large green sculpture". It should be understood that the foregoing list is not comprehensive, and many additional variations of template-based driving commands are possible, including some variations that provide a driving instruction statement 213 that includes more than one significant object.

[0075] Figure 3is a block diagram of a route guidance system 300 of the present invention. The route guidance system is implemented in a computer 305, which may interface with one or more peripheral devices as needed to implement functions. A driver control interface 310 interfaces the computer 305 to one or more driver controllers 311, which may include, for example, buttons on a vehicle steering wheel and enable a driver to provide a form of input to the route guidance system 300. A display interface 350 interfaces the computer 305 to one or more display devices 355, which may include, for example, a display mounted on an instrument cluster or a display mounted on a center console, and enables the route guidance system to display visual output to the driver. A camera interface 360 interfaces the computer 305 to one or more cameras 365, one of which is positioned to receive light from the vicinity of the front exterior of the vehicle. Another camera 365 may be positioned to receive light from the interior of the vehicle, enabling the route guidance system 300 to observe the driver's face and movements, thereby enabling another form of input. A distance sensor interface 370 interfaces the computer 305 to one or more distance sensors 375, which may include, for example, forward-looking, side-looking, or rear-looking radar and LiDAR facing the exterior, enabling the route guidance system to obtain 3D information about the vicinity of the vehicle, including distances to nearby objects. Additionally, the distance sensors 375 may include one or more radar sensors and LiDAR facing the interior, enabling the route guidance system to obtain 3D information about the driver's movements, thereby enabling another form of input to the system 300. A GPS interface 376 interfaces the computer 305 to a GPS receiver 377, which is capable of receiving GPS signals providing the current real-time position of the vehicle. A microphone interface 380 interfaces the computer 305 to one or more microphones 385 that may be located, for example, outside the vehicle to enable reception of sound signals from outside the vehicle and one or more microphones 385 located inside the vehicle to enable reception of sound signals from inside the vehicle, including the driver's voice. A speaker interface 390 interfaces the computer 305 to one or more speakers 395 to enable the system 300 to emit audible output to the driver, which may include, for example, driving instructions 213 presented in audible form by a speech synthesizer. Collectively, the driver controllers 311, cameras 365, distance sensors 375, GPS receiver 377, and microphones 385 constitute real-time sensors providing the real-time information 203 as described above.

[0076] The computer 305 may be equipped with a network interface controller (NIC) 312, which enables the system 300 to exchange information from a network 313, which may include, for example, the Internet. The exchanged information may include network-based maps and other data 314, such as the locations and attributes of static objects near the vehicle. The computer is equipped with a processor 320 that executes the actual algorithms required to implement the route guidance system 300 and a storage unit 330, which is a form of computer memory. For example, it may be dynamic random access memory (DRAM), a hard disk drive (HDD), or a solid state drive (SSD). The storage unit 330 may be used for many purposes, including but not limited to storing an object detection and classification module 331, an object attribute extraction module 332, a dynamic map 333, a route 334, a route guidance module 335, and a statement generation module 336. Additionally, the computer 305 has a working memory 340 for storing temporary data used by various modules and interfaces.

[0077] Multimodal Attention Method

[0078] A statement generator with a multimodal fusion model may be constructed based on the multimodal attention method. Figure 4 is a block diagram showing a multimodal attention method according to an embodiment of the present disclosure. In addition to the feature extractors 1 to K, the attention estimators 1 to K, the weighted sum processors 1 to K, the feature transformation modules 1 to K, and the sequence generator 450, the multimodal attention method further includes a modality attention estimator 455 and a weighted sum processor 445, instead of using a simple sum processor (not shown). The multimodal attention method is performed in combination with a sequence generation model (not shown), a feature extraction model (not shown), and a multimodal fusion model (not shown). In both methods, the sequence generation model may provide the sequence generator 450, and the feature extraction model may provide the feature extractors 1 to K (411, 421, 431). Additionally, the feature transformation modules 1 to K (414, 424, 434), the modality attention estimator 455, the weighted sum processors 1 to K (413, 423, 433), and the weighted sum processor 445 may be provided by the multimodal fusion model.

[0079] Given multimodal video data including K modalities, such that K ≥ 2 and some modalities may be the same, the modality-1 data is converted into a fixed-dimension content vector using a feature extractor 411, an attention estimator 412, and a weighted summation processor 413 of the data, where the feature extractor 411 extracts multiple feature vectors from the data, the attention estimator 412 estimates each weight of each extracted feature vector, and the weighted summation processor 413 outputs (generates) a content vector calculated as the weighted sum of the extracted feature vectors using the estimated weights. The modality-2 data is converted into a fixed-dimension content vector using a feature extractor 421, an attention estimator 422, and a weighted summation processor 423 of the data. Up to the modality-K data, K fixed-dimension content vectors are obtained, where the feature extractor 431, the attention estimator 432, and the weighted summation processor 433 are used for the modality-K data. Each of the modality-1, modality-2, …, modality-K data may be successive data in a successive order with intervals in time or other predetermined order with a predetermined time interval.

[0080] Then, each of the K content vectors is transformed (converted) into an N-dimensional vector by each feature transformation module 414, 424, and 434, and K transformed N-dimensional vectors are obtained, where N is a predetermined positive integer.

[0081] The K transformed N-dimensional vectors are summed into a single N-dimensional content vector in a simple multimodal method, while in the Figure 4 multimodal attention method, the vectors are converted into a single N-dimensional content vector using a modality attention estimator 455 and a weighted summation processor 445, where the modality attention estimator 455 estimates each weight of each transformed N-dimensional vector, and the weighted summation processor 445 outputs (generates) an N-dimensional content vector calculated as the weighted summation of the K transformed N-dimensional vectors using the estimated weights.

[0082] A sequence generator 450 receives the single N-dimensional content vector and predicts a single label corresponding to a word of a statement describing the video data.

[0083] To predict the next word, the sequence generator 450 provides context information of the statement, such as a vector representing a previously generated word, to the attention estimators 412, 422, 432, and the modality attention estimator 455 for estimating attention weights to obtain an appropriate content vector. The vector may be referred to as a pre-step (or pre-step) context vector.

[0084] The sequence generator 450 predicts starting with a sentence start token “ <sos>”the next word starting from, and iteratively predicting the next word (the predicted word) until the prediction corresponds to a special symbol for "end of statement" <eos>” to generate descriptive statements. In other words, the sequence generator 450 generates a sequence of words from the multimodal input vector. In some cases, the multimodal input vector can be received via different input / output interfaces such as an HMI and an I / O interface (not shown) or one or more I / O interfaces (not shown).

[0085] In each generation process, a predicted word with the highest probability among all possible words given by the weighted content vector and the pre-step context vector is generated. In addition, the predicted words can be accumulated in the memory 340, the storage device 330, or more storage devices (not shown) to generate a sequence of words, and this accumulation process can continue until a special symbol (the end of the sequence) is received. The system 300 can send the predicted words generated from the sequence generator 450 via the NIC and network, the HMI and I / O interface, or one or more I / O interfaces, so that the data of the predicted words can be used by other computers (not shown) or other output devices (not shown).

[0086] When each of the K content vectors comes from different modal data and / or is generated by different feature extractors, the modal or feature fusion with the weighted sum of the K transformed vectors enables better prediction of each word by focusing on different modalities and / or different features according to the context information of the statement. Therefore, this multimodal attention method can comprehensively or selectively use attention weights on different modalities or features to infer each word of the description using different features.

[0087] In addition, the multimodal fusion model in the system 300 can include a data distribution module (not shown), which receives a plurality of temporally successive data via the I / O interface, distributes the received data into Modality-1, Modality-2, …, Modality-K data, divides the successive data of each distribution according to one or more mined intervals, and then provides Modality-1, Modality-2, …, Modality-K data to the feature extractors 1 to K respectively.

[0088] In some cases, the multiple time - successive data can be a video signal captured by a camera and an audio signal recorded by a microphone. When the time - successive depth images obtained by a distance sensor are used as the modality data, the system 300 uses the feature extractors 411, 421, and 431 (with K = 3) in the figure. The real - time multi - modal information can include images (frames) from at least one camera, signals from a measurement system, communication data from at least one adjacent vehicle, or sound signals through at least one microphone arranged in the vehicle. The real - time multi - modal information is provided to the feature extractors 411, 421, and 431 in the system 300 through the camera interface 360, the distance - sensor interface 370, or the microphone interface 380. The feature extractors 411, 421, and 431 can extract image data, audio data, and depth data respectively as modality - 1 data, modality - 2 data, and modality - 3 data (e.g., K = 3). In this case, the feature extractors 411, 421, and 431 receive the modality - 1 data, modality - 2 data, and modality - 3 data respectively from the data streams of the real - time images (frames) according to the first interval, the second interval, and the third interval.

[0089] In some cases, when image features, motion features, or audio features can be captured at different time intervals, the data distribution module can divide the multiple time - successive data at predetermined different time intervals respectively.

[0090] In some cases, one or a combination of an object detector, an object classifier, a motion - trajectory estimator, and an object - attribute extractor can be used as one of the feature extractors, which receives the time - successive data with a predetermined time interval via the camera interface 360, the distance - sensor interface 370, or the microphone interface 380, and generates a sequence of feature vectors including information of the detected objects such as object positions, object categories, object attributes, object motions, and intersection positions.

[0091] Examples of multi - modal fusion models

[0092] The method of statement generation can be based on multi - modal sequence - to - sequence learning. Embodiments of the present disclosure provide an attention model for processing the fusion of multi - modalities, where each modality has its own sequence of feature vectors. For statement generation, multi - modal inputs such as image features, motion features, and audio features are available. In addition, the combination of multi - features from different feature - extraction methods usually effectively improves the statement quality.

[0093] Figure 5 is a block diagram showing an example of a multi - modal fusion method (multi - modal feature - fusion method) for statement generation assuming K = 2. The input image / audio sequence 560 can be in a time - successive order with a predetermined time interval. The input sequence of feature vectors is obtained using one or more feature extractors 561.

[0094] Given the input image / audio sequence 560, where one can be the image sequence X1 = x 11 , x 12 , …, x 1L , and the other can be the audio signal sequence X2 = x 21 , x 22 , …, x 2L’ . Each image or audio signal is first fed to the corresponding feature extractor 561 for the image or audio signal. For images, the feature extractor can be a pre-trained convolutional neural network (CNN), such as GoogLeNet, VGGNet, or C3D, where each feature vector can be obtained by extracting the activation vector of the fully connected layer of the CNN for each input image. The image feature vector sequence X’1 is shown as x’ Figure 5 in 11 , x’ 12 , …, x’ 1L . For audio signals, the feature extractor can be the mel-frequency analysis method, which generates Mel-frequency cepstral coefficients (MFCCs) as feature vectors. The audio feature vector sequence X’2 is shown as x’ Figure 5 in 21 , x’ 22 , …, x’ 2L’ .

[0095] The multimodal fusion method can employ an encoder based on bidirectional long short-term memory (BLSTM) or gated recurrent unit (GRU) to further transform the feature vector sequence so that each vector contains its context information. However, in the real-time image captioning task, the CNN-based features can be used directly, or one or more feedforward layers can be added to reduce the dimension.

[0096] If the BLSTM encoder is used after feature extraction, the activation vector (i.e., the encoder state) can be obtained as follows:

[0097]

[0098] where h t (f) and h t (b) are the forward and backward hidden activation vectors:

[0099]

[0100]

[0101] The hidden state of the LSTM is given by:

[0102] h t = LSTM(h t-1 , x' t ; λ E ), (4)

[0103] where λ can be the encoder network of a forward LSTM network or a backward LSTM network The LSTM function of λ E is calculated as:

[0104] LSTM(h t-1 , x t ; λ) = o t tanh(c t ), (5)

[0105] where

[0106]

[0107]

[0108]

[0109] where σ() is the element-wise sigmoid function, and i t , f t , to, and c t are the input gate, forget gate, output gate, and cell activation vector of the t-th input vector, respectively. The weight matrices W zz (λ) and the bias vectors b Z (λ) are identified by the subscript z ∈ {x, h, i, f, o, c}. For example, W hi is the hidden input gate matrix, and W xo is the input-output gate matrix. Peephole connections are not used in this program.

[0110] If a feedforward layer is used, the activation vector is calculated as

[0111] h t = tanh(W p x' t + b p ), (10)

[0112] where W p is the weight matrix and b p is the bias vector. Additionally, when directly using CNN features, it is assumed that h t = x t .

[0113] a. The attention mechanism 562 is implemented by applying attention weights to the hidden activation vectors over the entire input sequence 560 or the sequence of feature vectors extracted by the feature extractor 561. These weights enable the network to emphasize features from those time steps that are most important for predicting the next output word.

[0114] b. Let α i,t be the attention weight between the i-th output word and the t-th input feature vector. For the i-th output, a context vector c i is obtained as a weighted sum of the hidden unit activation vectors:

[0115]

[0116] The attention weights can be computed as

[0117]

[0118] And

[0119]

[0120] where W A and V A are matrices, w A and b A are vectors, and e i,t is a scalar.

[0121] In Figure 5 , the attention mechanism is applied to each modality, where c 1,i and c 2,i represent the context vectors obtained from the first modality and the second modality, respectively.

[0122] The attention mechanism is further applied to multimodal fusion. Using the multimodal attention mechanism, based on the previous decoder state s i-1 , the LSTM decoder 540 can selectively attend to specific modalities (or specific feature types) of the input to predict the next word. The attention-based feature fusion according to an embodiment of the present disclosure can be performed using the following formula

[0123]

[0124] to generate a multimodal context vector g i , 580, where is a matrix, and d k,i is the transformed context vector 570 obtained from the k-th context vector c k,i corresponding to the k-th feature extractor or modality, which is computed as

[0125]

[0126] Among them, is a matrix, is a vector. Then, the multi-modal context vector g i , 580 is fed into the LSTM decoder 540. The multi-modal attention weight β k,i is obtained in a manner similar to the temporal attention mechanism of equations (11), (12), and (13):

[0127]

[0128] where

[0129]

[0130] and, W B and V Bk are matrices, w B and b Bk are vectors, and v k,i is a scalar.

[0131] The LSTM decoder 540 employs an LSTM-based decoder network λ D , which utilizes the multi-modal context vector g i (i = 1, …, M + 1) to generate the output word sequence 590. The decoder starts with the sentence start token " <sos>”Start predicting the next word iteratively until the end marker of its prediction statement" <eos>” up to this point. The statement start marker can be referred to as the start tag, and the statement end marker can be referred to as the end tag.

[0132] Given the decoder state s i-1 and the multimodal context vector g i , the LSTM decoder 540 infers the next word probability distribution as

[0133]

[0134] where, and are matrices, is a vector.

[0135] And the decoder predicts the next word y with the highest probability according to the following formula i :

[0136] y i = argmax y∈U P(y|s i-1 , g i ), (19)

[0137] where, U represents the vocabulary. The LSTM network of the decoder updates the decoder state to

[0138] s i = LSTM(s i-1 , y' i ; λ D ), (20)

[0139] where, y’ i is the word embedding vector of y i , and the initial state s0 is set to the zero vector, and y’0 is given as the start tag <sos>word embedding vectors. During the training phase, given Y = y1, …, y M As a reference, to determine the matrices and vectors represented by W, V, w, and b in equations (1) to (20). However, during the testing phase, the best word sequence needs to be found based on the following formula:

[0140]

[0141] Therefore, the beam search method in the testing phase can be used to maintain multiple states and hypotheses with the highest cumulative probability at each i-th step, and select the best hypothesis from those hypotheses that have reached the end-of-statement marker.

[0142] Example description of statement generation for a scene-aware interactive navigation system

[0143] To design a scene-aware interactive navigation system, according to some embodiments of the present invention, real-time images captured by a camera on a vehicle can be used to generate navigation statements for a human driver of the vehicle. In this case, the object detection and classification module 331, the object attribute extraction module 332, and the motion trajectory estimation module (not shown) can be used as feature extractors of the statement generator.

[0144] The object detection and classification module 331 can detect multiple significant objects from each image, where a bounding box and an object category are predicted for each object. The bounding box indicates the position of the object in the image, which is represented as a four-dimensional vector (x1, y1, x2, y2), where x1 and y1 represent the coordinate points of the upper left corner of the object in the image, and x2 and y2 represent the coordinate points of the lower right corner of the object in the image.

[0145] The object class identifier is an integer indicating a predetermined object class such as a building, a sign, a pole, a traffic light, a tree, a person, a bicycle, a bus, and a car. The object class identifier can be represented as a one-hot vector. The object attribute extraction module 332 can estimate the attributes of each object, wherein the attributes can be the shape, color, and state of the object, such as high, wide, large, red, blue, white, black, walking, and standing. The attributes are predicted as attribute identifiers, which are predetermined integers for each attribute. The attribute identifier can be represented as a one-hot vector. The motion trajectory estimation module (not shown) can use the previously received images to estimate the motion vector of each object, wherein the motion vector can include the direction and speed of the object in the 2D image and can be represented as a 2-dimensional vector. The motion vector can be estimated using the difference in the position of the same object in the previously received images. The object detection and classification module 331 can also detect road intersections and provide a four-dimensional vector representing the bounding box of the road intersection. Using these aforementioned modules, a feature vector for each detected object can be constructed by concatenating the bounding box vectors of the object and road intersections, the one-hot vectors of the object category identifiers and attribute identifiers, and the motion vector of the object.

[0146] For multiple objects, the feature vectors can be considered as different vectors from different feature extractors. Multimodal fusion methods can be used to fuse feature vectors from multiple objects. In this case, a sequence of feature vectors for each object is constructed by assigning the feature vector of the object detected in the current image to the object with the highest degree of overlap detected in the previous image. The degree of overlap between two objects can be calculated using the Intersection-over-Union (IoU) measure:

[0147]

[0148] where |A∩B| and |A∪B| are the intersection and union regions of two objects A and B, respectively. If the bounding box of A is And the bounding box of B is Then the intersection and union can be calculated as:

[0149]

[0150]

[0151] Assume I1,I2,…,I t represents the temporal succession of image data up to the current time frame t. For object o, the τ-length feature vector sequence It can be obtained as follows.

[0152]

[0153]

[0154] Among them, O(I) represents the set of detected objects from image I, and F(o) is a function that extracts feature vectors from object o. If there are no overlapping objects in the previous image, a zero vector can be used, that is, if then where β is a predetermined threshold for ignoring very small overlaps, and d represents the dimension of the feature vector. According to equations (21) and (22), multiple sequences of feature vectors can be obtained by starting each sequence with each object in O(I t-τ+1 ).

[0155] For example, Faster R-CNN (Ren Shaoqing et al., "Faster R-CNN: Towards real-time object detection with region proposal networks", Advances in neural information processing systems, 2015) is a known prior art method that can be used in the object detection and classification module 331 and the object attribute extraction module 332.

[0156] When the given route indicates that the vehicle will turn right at the next intersection, the system can generate a statement such as "Turn right at the intersection before the black building". To generate such a statement, information about the turning direction can be added to the feature vector. In this case, the direction can be represented as a 3-dimensional one-hot vector such that

[0157] (1, 0, 0) = turn left,

[0158] (0, 1, 0) = go straight,

[0159] (0, 0, 1) = turn right,

[0160] (0, 0, 0) = no intersection.

[0161] This direction vector can be concatenated to each feature vector extracted from each object.

[0162] In addition, in order to configure the voice dialogue system to accept voice requests from the driver and output voice responses to the driver, a speech recognition system and a text-to-speech synthesis system can be included in the system. In this case, the text statement given as the result of speech recognition of the driver's request can be fed into the statement generation module 336, where the statement can be used for Figure 4 One of the multimodal inputs of the multimodal fusion method (i.e., Modal-k data, where k is an integer such that 1 ≤ k ≤ K). Each word in the sentence can be converted into a word embedding vector of a fixed dimension, so the text sentence can be represented as a sequence of feature vectors. Using the sequence of feature vectors extracted from the driver's request and the detected objects, the multimodal fusion method can generate a reasonable sentence as a response to the driver. The generated sentence can be further converted into a voice signal by a text-to-speech synthesis system to output the signal via some audio speakers.

[0163] Settings for training the sentence generation model

[0164] To learn an attention-based sentence generator with a multimodal fusion model, scene-aware interaction data is created, where 21,567 images are obtained from a camera attached to the car dashboard. Then the images are annotated by human subjects, where 36,935 object intersection pairs are labeled with object names, attributes, bounding boxes, and sentences to provide navigation to the car driver. The data includes 2,658 unique objects and 8,970 unique sentences.

[0165] The sentence generation model, i.e., the decoder network, is trained using the training set to minimize the cross-entropy criterion. The image features are fed to a BLSTM encoder before the decoder network. The encoder network has two BLSTM layers with 100 units each. The decoder network has one LSTM layer with 100 units. Each word is embedded into a 50-dimensional vector when fed into the LSTM layer. We adopt the AdaDelta optimizer (M.D. Zeiler. ADADELTA: An adaptive learning rate method. CoRR, abs / 1212.5701, 2012.) to update the parameters, which is widely used to optimize the attention model. The LSTM and attention model are implemented using PyTorch (Paszke, Adam, et al., "PyTorch: An imperative style, high-performance deep learning library." Advances in Neural Information Processing Systems, 2019).

[0166] Figure 6A It is a flowchart showing the training of a parameter function 635 of a navigation system 600A according to an embodiment of the present disclosure. The parameter function 635 is configured to generate a driving instruction 640 based on a vehicle state 610 and a dynamic map 611. For example, the parameter function 635 can be implemented as a neural network including parameters included in a parameter set 650, or can be implemented as a rule-based system that also involves the parameters included in the parameter set 650. The training can be performed by considering a training set 601 of training data examples, which include a combination of an observed vehicle state 610, an observed dynamic map 611, and a corresponding driving instruction 602. The training data examples can be collected by driving a vehicle under various conditions, recording the observed vehicle state 610 and the observed dynamic map 611, and collecting the corresponding driving instruction 602 as a label by asking a human to give examples of driving instructions that they think are relevant to guiding a driver in the case corresponding to the current vehicle state and dynamic map. To help the driver safely follow the intended route that the navigation system is trying to guide the driver on, multiple people can be asked to provide examples of one or more driving instructions that they think are particularly suitable as a hypothesized driving instruction in the current situation, based on how intuitive, natural, relevant, easy to understand, clear, etc. the driving instruction can be considered. The corresponding driving instruction can be collected by a passenger while the vehicle is being driven, or can be collected offline by showing examples of the vehicle state and dynamic map to a human labeler, who annotates them with the corresponding driving instruction. For example, if the vehicle collecting the training data encounters an intersection where a black car is turning right in front of the vehicle on the intended route that the navigation system is trying to guide the driver on, a video clip from the vehicle dashboard camera showing the moment the black car turns right and an indication that the intended route means to turn right at this intersection can be shown to the human labeler, and the human labeler labels this moment with a corresponding driving instruction such as "Follow the turning black car". For example, if the human labeler notices a potential hazard that may affect the ability to turn right safely, such as a pedestrian trying to cross the street on the future path of the vehicle, the labeler can label this moment with a corresponding driving instruction (e.g., "Watch out for the pedestrian trying to cross the street"). The objective function calculation module 645 calculates the objective function by calculating an error function between the generated driving instruction 640 and the training driving instruction 602. The error function can be based on a similarity metric, a cross-entropy criterion, etc. The training module 655 can use the objective function to update the parameters 650. In the case where the parameter function 635 is implemented as a neural network, the training module 655 is a network training module, and the parameters 650 include network parameters.In a case where the parameter function 635 is implemented as a rule-based system, the parameter 650 includes parameters of the rule-based system such as weights and thresholds, and the parameters of the rule-based system can be modified using the training module 655 to minimize or reduce the objective function 645 with respect to the training set 601.

[0167] Figure 6B is a flowchart showing the training of the parameter functions of the navigation system 600B according to an embodiment of the present disclosure. The first parameter function 615 is configured to determine the attributes and spatial relationships of a set of significant objects of the dynamic map 611 based on the vehicle state 610 to obtain the transformed dynamic map 620, and the second parameter function 635 is configured to generate a driving instruction 640 based on the transformed dynamic map 620. For example, the first parameter function 615 can be implemented as a neural network having parameters included in the parameter set 650, or implemented as a rule-based system that also involves the parameters included in the parameter set 650, and the second parameter function 635 can be implemented as a neural network having parameters included in the parameter set 650, or implemented as a rule-based system that also involves the parameters included in the parameter set 650. Training can be performed in a similar manner to Figure 6A the system where the parameter 650 also includes the parameters of the first parameter function 615 here, which can also be trained using the training module 655 based on the objective function 645 obtained by comparing the generated driving instruction 640 with the training driving instruction 602.

[0168] Figure 6C is a flowchart showing the training of the parameter functions of the navigation system 600C according to an embodiment of the present disclosure. The first parameter function 605 is configured to determine the vehicle state 610 and the dynamic map 611 based on the measurement 603 from the scene, and the second parameter function 635 is configured to generate a driving instruction 640 based on the vehicle state 610 and the dynamic map 611. For example, the first parameter function 605 can be implemented as a neural network having parameters included in the parameter set 650, or implemented as a rule-based system that also involves the parameters included in the parameter set 650. Training can be performed by considering the training set 601 of the collected training data examples, where the training data examples include combinations of observed measurement values 603 and corresponding driving instructions 602. Training data examples can be collected in a similar manner to Figure 6A the system in [reference] by driving the vehicle under various conditions, recording the measurements 603 of the observed scenes, and collecting the corresponding driving instructions 602. Given the training set 601, it can be in a similar manner to Figure 6A Perform training in a similar manner in the system, where the parameter 650 further includes the parameters of the first parameter function 605 here, which can also be trained by the training module 655 based on the objective function 645 obtained by comparing the generated driving instruction 640 with the training driving instruction 602.

[0169] Figure 6D is a flowchart showing the end-to-end training of the parameter function 635 of the navigation system 600D according to an embodiment of the present disclosure. The navigation system 600D is configured to generate a driving instruction 640 based on measurements 603 from a scene. Training can be performed by considering a training set 601 of collected training data examples, which include combinations of observed measurements 603 and corresponding driving instructions 602. The objective function calculation module 645 calculates the objective function by calculating the error function between the generated driving instruction 640 and the training driving instruction 602. The training module 655 can use the objective function to update the parameter 650.

[0170] Figure 6E is a flowchart showing the training of the parameter functions of the navigation system 600E according to an embodiment of the present disclosure. The first parameter function 605 is configured to determine the vehicle state 610 and the dynamic map 611 based on measurements 603 from a scene. The second parameter function 615 is configured to determine the attributes and spatial relationships of a set of significant objects of the dynamic map 611 based on the vehicle state 610 to obtain a transformed dynamic map 620. The third parameter function 625 is configured to select a subset 630 of significant objects from the transformed dynamic map 620, and the fourth parameter function 635 is configured to generate a driving instruction 640 based on the selected significant objects 630. For example, each parameter function can be implemented as a neural network or a rule-based system, with parameters included in the parameter set 650. Training can be performed by considering a training set 601 of collected training data examples, which include combinations of observed measurements 603 and corresponding driving instructions 602. The objective function calculation module 645 calculates the objective function by calculating the error function between the generated driving instruction 640 and the training driving instruction 602. The training module 655 can use the objective function to update the parameter 650.

[0171] Figure 6F It is a flowchart showing the multi-task training of parameter functions of a navigation system 600E according to an embodiment of the present disclosure. The first parameter function 605 is configured to determine a vehicle state 610 and a dynamic map 611 based on measurements 603 from a scene. The second parameter function 615 is configured to determine attributes and spatial relationships of a set of significant objects of the dynamic map 611 based on the vehicle state 610 to obtain a transformed dynamic map 620. The third parameter function 625 is configured to select a subset 630 of significant objects from the transformed dynamic map 620, and the fourth parameter function 635 is configured to generate a driving instruction 640 based on the selected significant objects 630. For example, the parameter function can be implemented as a neural network having parameters included in a parameter set 650. Training can be performed by considering a training set 601 of collected training data examples, which includes a combination of observed measurements 603 and corresponding labeled data 602. The labeled data 602 includes one or a combination of a vehicle state label, a dynamic map label, a transformed dynamic map label, a selected significant object, and a driving instruction. The objective function calculation module 645 calculates an objective function by calculating a weighted sum of one or a combination of an error function between the determined vehicle state 610 and a training vehicle state from the labeled data 602, an error function between the determined dynamic map 611 and a training dynamic map 602 from the labeled data 602, an error function between the selected significant objects 630 and a training selected significant object from the labeled data 602, and an error function between the generated driving instruction 640 and a training driving instruction from the labeled data 602. The training module 655 can use the objective function to update the parameters 650.

[0172] Figure 7 Illustrates example significant objects 710, 720, 730, 740, 750, 760 in a dynamic map according to some embodiments of the present disclosure, as well as corresponding attributes and their values 711, 721, 731, 741, 751, 761 of the objects.

[0173] The types of attributes that a prominent object in a dynamic map may possess include: category, color, dynamics (i.e., motion), shape, size, location, appearance, and depth. The attribute category refers to the type of the object. For example, for a prominent object 760, the attribute 761 category has a value of intersection, indicating that the object is an intersection between two roads. Other possible values of the category attribute include car 711, building 721, pedestrian 741, and sound types such as siren 751. Another attribute, color, refers to the color of the object and can have values such as the brown color of building 721, the white color of building 731, or the black color of car 711. Another attribute used in some embodiments is the dynamic state of the object, i.e., information about the motion of the object, which can take values such as the direction of travel of the object (e.g., a car 711 turning right), its speed (e.g., 15 km / h of car 711), or its lack of motion (if the object is a dynamic object such as a car or a pedestrian that is currently stationary). Other attributes used in some embodiments include: the shape of buildings 721, 731; location, such as the depth of car 711 relative to vehicle 701 or its position relative to the reference frame of the dynamic map; the size of the entire prominent object; and the size of the visible portion of the prominent object in cases where the processor determines that only a portion of the object is visible from the driver's perspective.

[0174] It should be noted that in some embodiments in accordance with the present disclosure, a prominent object does not need to be currently visible or perceivable by the driver in order to be relevant from the driver's perspective. For example, an ambulance approaching the current or future location of a vehicle can be relevant to be included in a driving instruction such as "Warning: Ambulance approaching from behind" or "Warning: Ambulance approaching from behind the blue building on the left", even if the driver of the vehicle cannot currently see or hear the ambulance.

[0175] The spatial relationships between prominent objects in a dynamic map are also used for generating driving instructions. The spatial relationships can indicate the relative 3D position of one or more objects relative to another object or a set of objects. The relative position is expressed as being located to the left, right, front, rear, above, below, etc. Depth or distance information estimated from a camera or directly obtained from a distance sensor such as a LiDAR or radar sensor (i.e., a depth sensor) is used in determining the relative 3D position.

[0176] Figure 8 Shows example prominent objects 801, 802, 803, 804 in a dynamic map, as well as pairwise selected spatial relationships 812, 834. In this example, prominent object 801 has a spatial relationship 812 with prominent object 802, indicating that prominent object 801 is 5 meters to the left of prominent object 802. Similarly, prominent object 803 has a spatial relationship 834 with prominent object 804, indicating that prominent object 804 is 20 meters in front of and 15 meters to the right of prominent object 803.

[0177] The movement trajectories and sounds from prominent objects can also be used to generate driving instructions. A movement trajectory indicating the movement of a prominent object over a predetermined amount of time is determined for each prominent object. Sounds associated with the prominent objects are directly obtained using a microphone.

[0178] As Figure 9 and Figure 10 shown, a movement trajectory 916 is estimated based on the movement of prominent object 906 over a predetermined amount of time, and a movement trajectory 914 is estimated based on the movement of prominent object 904 over a predetermined amount of time. The scene also includes static objects 902, 903, and an occlusion object 905 that emits a unique sound that can be sensed by the measurement system of vehicle 901.

[0179] At a specific moment, the navigation system can compare the attributes of prominent objects perceived from the driver's perspective to estimate a relevance score for each prominent object, which indicates the relevance of the prominent object for inclusion in the generated driving instructions. Then, the navigation system selects prominent objects from the set of prominent objects based on the indication of the relevance scores of the prominent objects for inclusion in the generated driving instructions. The navigation system estimates the relevance score of each prominent object based on one or a combination of a function of the distance of the prominent object to the vehicle, a function of the distance of the prominent object to the next turn on the route, and a function of the distance of the vehicle to the next turn on the route.

[0180] For Figure 9 and Figure 10 the example shown, the route of vehicle 901 indicates that the vehicle should turn right 950 at the upcoming intersection. As Figure 9 shown, when vehicle 901 is 100 meters from the intersection, the prominent object with the highest relevance score 930 is the prominent object 906 with movement trajectory 916, and the generated driving instruction is "Follow the black car turning right". As Figure 10 shown, when vehicle 901 is 50 meters from the intersection, the prominent objects with the highest relevance score 1030 include the prominent object 904 with movement trajectory 914 and the prominent object 905. In this case, the generated driving instruction 1040 is "Be careful of pedestrians crossing the street from the left and an ambulance approaching from the left".

[0181] These examples illustrate the adaptation of a navigation system based on a set of significant objects and their attributes related to the vehicle's route at the current moment and the vehicle's state.

[0182] The navigation system generates driving instructions in the form of statements according to linguistic rules, such that an output interface is connected to a speaker configured to read out the linguistic statement. The navigation system also supports a voice dialogue system configured to receive voice requests from the driver and output voice responses to the driver, such that the linguistic statement uses the operation history of the voice dialogue system. The voice dialogue system is used to clarify the generated driving instructions or provide other interaction means between the driver and the scenario as well as the driving instructions.

[0183] Figure 11 A scenario is shown with a set of vehicle 1101 and significant objects 1102, 1103, 1104 in a dynamic map. The first generated driving instruction 1105 is "Follow the black car turning right". In the case where there are two black cars 1102, 1104 in the scenario, the driver can request clarification, "Which black car?" 1106. The second generated driving instruction 1107 is "The black car in front of the low dark building".

[0184] Figure 12 It is a flowchart showing a specific embodiment of the path guidance system of the present invention. In this embodiment, the system receives real-time sensor information from one or more audio sensors 1211, one or more cameras 1212, one or more LiDAR distance sensors 1214, GPS position 1201, and route direction 1210. The object detector and classifier 1220 outputs a set of all detected objects that should be understood to include the attributes of the objects. The significant object detector 1222 uses the route direction 1210 to determine the dynamic map 1224 as previously discussed. In this embodiment, significant objects follow two different processing paths, depending on whether they are static objects (such as buildings) or dynamic objects (such as cars). Information about dynamic objects is processed by the significant dynamic object trajectory estimator 1240, which estimates the trajectory of an object consisting of the moving speed and direction of the object. There are many possible implementations of the significant dynamic object trajectory estimator 1240, including an implementation that compares the position of an object in a first camera image with the position of the object in a second camera image to estimate the object trajectory.

[0185] Thereafter, the significant dynamic object attribute extractor 1241 extracts the attributes of significant dynamic objects to produce a set 1242 of significant dynamic objects with attributes. The significant static object attribute extractor 1231 extracts the attributes of the set of significant static objects to produce a set 1232 of significant static objects with attributes. The significant static object attribute extractor 1231 also receives as input local map data 1203 obtained from a map server 1202 using the vehicle GPS location 1201. This enables the significant static object attribute extractor 1231 to include additional attributes of the significant static objects, such as the name of the object. For example, if the significant static object is an enterprise, the attributes of the object can include the enterprise name.

[0186] There are many possible implementations of the existence statement generation module 1243. A very powerful implementation is a parametric function of a neural network implemented to be trained using a dataset of corresponding statements and significant objects provided by human labelers.

[0187] Figure 12 The specific implementation of the statement generation module 1243 shown in uses a rule-based object sorter 1245 that uses a set of manually generated rules to sort the significant objects in order to output the selected significant objects 1250. The rules can be used to compare sets of significant objects based on the data and attributes of the significant objects to sort the significant objects to identify the selected significant objects 1250. For example, the rules can favor dynamic objects moving in the same direction as the vehicle. These rules can favor larger objects over smaller objects, or favor bright colors such as red or green over darker colors such as brown or black.

[0188] As a specific example of an implementation of the object sorter 1245, we mathematically define two bounding boxes, one for the intersection of the next turning point and the second for the object,

[0189] Intersection:

[0190] Object:

[0191] where, and are the upper left x and y camera image coordinates of the bounding box of the intersection of the next turning point respectively, and are the lower right x and y camera image coordinates of the bounding box of the intersection respectively, and x1 and y1 are the upper left x and y camera image coordinates of the bounding box b of the object, and x2 and y2 are the lower right x and y camera image coordinates of the bounding box b of the object.

[0192] For each salient object O, we compute a set of metrics related to its data and attributes. For example:

[0193]

[0194] is a number that measures the area of an object in a camera image. It is larger for larger objects.

[0195]

[0196] is a number that is a measure of the prevalence of a salient object of a given class among all salient objects, where N(O) is the number of salient objects of the same class as O, and H is the total number of salient objects. If there are fewer objects of the same class as O, then this number is larger.

[0197]

[0198] is a number that is a measure of the prevalence of a salient object of a given color among all salient objects, where N(c) is the number of salient objects of the same color as o, and H is the total number of salient objects. If there are fewer salient objects of the same color as salient object O, then this number is larger.

[0199] Each of these metrics can be combined into a salient object score by the following formula:

[0200] S = W A F A + W o F o + W C F C

[0201] where W A 、W o and W c are manually determined weights used to define the relative importance of each metric. For example, we can choose W A = 0.6, W o = 0.3 and W c = 0.1, respectively. Then, the object sorter 1245 computes S for all salient objects and sorts them from largest to smallest. Then the selected salient object 1250 can be determined to be the salient object with the largest score S. The exact values of the weights are usually adjusted manually until the system works well. It should be understood that there are many possible salient object metrics and ways to combine them into a score, and the above disclosure is only one possible implementation.

[0202] In addition, in this embodiment, the object sorter 1245 of the statement generator 1243 receives as input any detected driver speech 1261 from the driver that has been detected by the automatic speech recognition module 1260 from the audio input 1211. The dialogue system 1262 provides an output for adjusting the function of the object sorter 1245. For example, a previous driving instruction for the first prominent object is presented to the driver, but the driver does not see the prominent object being referred to. As a result, the driver indicates by voice that they do not see the prominent object. Therefore, the object sorter should reduce the score of the previous prominent object in order to select an alternative prominent object as the selected prominent object 1250.

[0203] Furthermore, another embodiment of the present invention is based on the recognition that a method for providing route guidance to a driver in a vehicle can be implemented by the following steps: obtaining multimodal information; analyzing the obtained multimodal information; identifying one or more prominent objects based on the route; and generating a statement providing route guidance based on the one or more prominent objects. The method may include the step of outputting the generated statement using one or more of a speech synthesis module or a display. In this case, the route is determined based on the current location and the destination, the statement is generated based on the obtained multimodal information and the prominent objects, and the multimodal information includes information from one or more imaging devices. The analysis may be achieved by including one or a combination of the following steps: detecting and classifying multiple objects; associating multiple attributes with the detected objects; detecting the position of an intersection in the forward direction of the vehicle based on the route; estimating the movement trajectories of a subset of the objects; and determining the spatial relationship between the detected subsets of the objects, where the spatial relationship indicates the relative position and orientation between the objects. In some cases, the step of detecting and classifying multiple objects may be performed by using a machine learning-based system, and the attributes may also include one or a combination of a primary color and a depth relative to the current position of the vehicle, and the classified object categories may include one or more of pedestrians, vehicles, bicycles, buildings, traffic signs. In addition, the generated statement may provide route guidance including driving instructions related to the prominent objects, and the generated statement indicates a warning based on the result of the analysis.

[0204] In some cases, the imaging device can be one or more cameras, one or more distance sensors, or a combination of one or more cameras and one or more distance sensors. In some cases, at least one distance sensor can be LiDAR (Light Detection and Ranging) or radar, etc., and one or more of the imaging devices can capture information from the vehicle's surrounding environment. Additionally, the multimodal information can include signals obtained in real time when the vehicle is being driven and / or sound signals obtained by one or more microphones, and in some cases, the sound signal can be the user's voice, which allows the navigation system using this method to interact with the user (driver) and generate more information for the user. The multimodal information can be the history of the interaction between the user and the system and includes map information. The interaction can include one or more of user voice input and previously generated statements. The analysis can also include locating the vehicle in the map. In this case, the map information can include multiple points of interest, and one or more significant objects are selected from the points of interest based on the result of the analysis.

[0205] Above, the navigation system is described as one example application of the scene-aware interaction system. However, the present invention is not limited to the navigation system. For example, some embodiments of the present invention can be used in in-vehicle infotainment and household appliances, interaction with service robots in building systems, and monitoring systems. GPS is only one of the positioning methods for the navigation system, and other positioning methods can also be used for other applications.

[0206] According to another embodiment of the present disclosure, the scene-aware interaction system can be implemented by changing the driver control interface 310 and the driver controller 311 into a robot control interface (not shown) and a robot control interface. In this case, the GPS / positioner interface 376 and the GPS / positioner 377 can be used according to the system design of the service robot, and the training data set can be changed.

[0207] In addition, according to the embodiments of the present disclosure, an effective method for executing the multimodal fusion model is provided. Therefore, the methods and systems using the multimodal fusion model can reduce the use of the central processing unit (CPU), power consumption, and / or network bandwidth usage.

[0208] The above embodiments of the present disclosure can be implemented in any of a variety of ways. For example, the embodiments can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or a set of processors provided either in a single computer or distributed among multiple computers. Such a processor can be implemented as an integrated circuit having one or more processors in the integrated circuit components. However, the processor can be implemented using any appropriate format of circuitry.

[0209] In addition, the various methods or processes outlined herein can be encoded as software that can be executed on one or more processors employing any one of a variety of operating systems or platforms. Additionally, such software can be written using any of a number of suitable programming languages and / or programming or scripting tools, and can also be compiled into intermediate code or executable machine language code that is executed on a framework or virtual machine. In general, in various embodiments, the functionality of program modules can be combined or distributed as desired.

[0210] In addition, embodiments of the present disclosure can be implemented as a method, and examples of the method have been provided. The acts performed as part of the method can be ordered in any suitable way. Accordingly, embodiments can be constructed in which the acts are performed in an order different than that shown, which order can include performing some acts simultaneously, although shown as sequential acts in illustrative embodiments. Further, the use of ordinal terms such as "first", "second", etc. in the claims to modify a claim element itself does not mean any precedence, priority, or order of one claim element with respect to another or the temporal order of performing method acts, but is merely used as a label to distinguish one claim element having a certain name from another element having the same name (except for the use of the ordinal term) to distinguish claim elements.

[0211] Although the present disclosure has been described with reference to certain preferred embodiments, it should be understood that various other adaptations and modifications can be made within the spirit and scope of the present disclosure. Accordingly, aspects of the appended claims cover all such variations and modifications that fall within the true spirit and scope of the present disclosure.< / sos> < / eos> < / sos> < / eos> < / sos>

Claims

1. A navigation system configured to provide driving instructions to a driver of a vehicle based on a real-time description of objects related to the driving vehicle in a scene, the navigation system comprising: An input interface configured to receive a dynamic map for a route for driving the vehicle, a state of the vehicle on the route at a current moment, and a set of significant objects related to the vehicle's route at the current moment, wherein the dynamic map includes the set of significant objects detected by multimodal sensing information and a set of attributes of the objects, the multimodal sensing information including images or videos captured by a camera, audio information obtained by a microphone, and position information estimated by a distance sensor, wherein at least one significant object is an object sensed by a measurement system of the vehicle, wherein the vehicle moves on the route between a current position at the current moment and a future position at a future moment, and wherein the set of significant objects includes one or more static objects and one or more dynamic objects; A processor configured to generate driving instructions based on a description of significant objects in the dynamic map, wherein the significant objects in the dynamic map are derived from a driver's perspective specified by the state of the vehicle, and wherein the state of the vehicle includes a position and an orientation of the vehicle relative to the dynamic map; and An output interface configured to present the driving instructions to the driver of the vehicle, wherein The processor is configured to submit the state of the vehicle and the dynamic map to a parameter function configured to generate the driving instructions, The dynamic map includes features indicating values of attributes of the significant objects and spatial relationships between the significant objects, wherein the processor determines the attributes of the significant objects and the spatial relationships between the significant objects, updates the attributes and the spatial relationships, and submits the updated attributes and spatial relationships to the parameter function to generate the driving instructions, The driving instructions include driving commands selected from a set of predetermined driving commands, wherein each of the predetermined driving commands is modified based on one or more significant objects and is associated with a score indicating a level of clarity of the modified driving command to the driver, and wherein the parameter function is trained to generate the driving instructions including the modified driving commands with higher scores.

2. The navigation system according to claim 1, wherein, The parameter function is trained using training data including a combination of vehicle states, dynamic maps, and corresponding driving instructions related to the driver's perspective.

3. The navigation system according to claim 1, wherein The attributes of the significant objects include one or a combination of a category of the significant object, a dynamic state of the significant object, a shape of the significant object, a size of the significant object, a size of a visible part of the significant object, a position of the significant object, and a color of the significant object, wherein the spatial relationship includes one or a combination of a type of relative position, height, distance, angle, and occlusion. Cause the processor to update the attribute and the spatial relationship based on the state of the vehicle.

4. The navigation system according to claim 1, wherein the navigation system further comprises: A communication interface configured to receive measurements of the scene at the current moment from the measurement system, wherein the measurement is received from at least one sensor, and the at least one sensor includes one or a combination of a camera, a depth sensor, a microphone, a GPS of the vehicle, a GPS of an adjacent vehicle, a distance sensor, and a sensor of a roadside unit RSU.

5. The navigation system according to claim 4, wherein, The processor executes a first parametric function that is trained to extract features from the measurement to determine the state of the vehicle and the dynamic map.

6. The navigation system according to claim 5, wherein, The processor executes a second parametric function that is trained to generate the driving instruction from the extracted features generated by the first parametric function, wherein the first parametric function and the second parametric function are jointly trained.

7. The navigation system according to claim 4, wherein, The processor executes a parametric function that is trained to generate the driving instruction from the measurement.

8. The navigation system according to claim 1, wherein, The set of predetermined driving commands includes a follow driving command, a subsequent turn driving command, and a previous turn driving command.

9. The navigation system according to claim 4, wherein, The processor is configured to: execute a first parametric function that is trained to extract features from the measurement to determine the state of the vehicle and the dynamic map; execute a second parametric function that is trained to transform the dynamic map based on the state of the vehicle to produce a transformed dynamic map that specifies the attributes and spatial relationships of the prominent objects from the driver's perspective; execute a third parametric function that is trained to select one or more prominent objects from the set of prominent objects based on the attributes and the spatial relationships of the prominent objects in the transformed dynamic map; and execute a fourth parametric function that is trained to generate the driving instruction based on the attributes and the spatial relationships of the selected prominent objects.

10. The navigation system according to claim 1, wherein, The driving instruction is generated in the form of a sentence according to linguistic rules, wherein the output interface is connected to a speaker configured to read out the linguistic sentence.

11. The navigation system according to claim 10, wherein the navigation system further comprises: A voice dialogue system configured to receive a voice request from the driver and output a voice response to the driver, wherein the processor uses the operation history of the voice dialogue system to generate the linguistic sentence.

12. The navigation system according to claim 1, wherein, The processor is configured to compare the attributes of the prominent objects perceived from the driver's perspective to estimate a relevance score for each prominent object, the relevance score indicating the relevance of the prominent object for inclusion in the generated driving instruction; and select prominent objects from the set of prominent objects based on the value of the relevance score of the prominent object to include the prominent object in the generated driving instruction.

13. The navigation system according to claim 12, wherein, Estimate the relevance score of each significant object based on one or a combination of a function of the distance from the significant object to the vehicle, a function of the distance from the significant object to the next turn on the route, and a function of the distance from the vehicle to the next turn on the route.

14. The navigation system according to claim 1, wherein, The description of the significant object includes a driving command, values of attributes of the significant object, and a label of a category of the significant object.

Citation Information

Patent Citations

  • Systems and Methods for Using Real-Time Imagery in Navigation

    US20170314954A1

  • System and method of object-based navigation

    US20190325746A1

  • System and method for using context in navigation dialog

    US7831433B1