Determining a relationship between an object and a human-provided representation associated with the object

The system determines relationships between objects and human-provided representations using image and audio analysis, addressing the limitations of existing ADAS systems by integrating human feedback for enhanced object recognition and classification.

JP2026021242APending Publication Date: 2026-02-10TOYOTA JIDOSHA KK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025080278
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-29
Filing Date
2025-05-13
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing systems struggle to determine relationships between objects and human-provided representations without predetermining relationships in a database schema or using textual information, limiting the integration of human feedback in advanced driver assistance systems (ADAS).

Method used

A system and method that utilize image and audio recordings to determine relationships between objects and human-provided representations, such as gestures or comments, by analyzing the spatial and temporal correlation between camera-generated images and human feedback, enabling database queries and annotation based on these relationships.

Benefits of technology

Enables effective integration of human feedback into ADAS systems, enhancing object recognition and classification by incorporating human opinions and perspectives, improving the accuracy and relevance of ADAS functionalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021242000001_ABST
    Figure 2026021242000001_ABST
Patent Text Reader

Abstract

To provide a system.SOLUTION: A system for determining a relationship between an object and a human-provided representation associated with the object may include a processor and a memory. The memory may store an image receiving module, a record receiving module, a relationship determining module, and a database querying module. The image receiving module may receive an image that includes a depiction of the object but lacks a depiction of a human-provided indication associated with the object. The record receiving module may receive a record that includes a representation of the human-provided indication but lacks a representation of the object. The relationship determination module may determine the existence of a relationship between the object and the human-provided representation. The database query module may cause the database to generate information about the object based on the determination of the existence of the relationship in response to a query about the subject matter of the human-provided representation.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The techniques of this disclosure are directed to determining relationships between objects and human-provided representations associated with the objects. [Background technology]

[0002] The development of advanced driver assistance systems (ADAS) technology has led to the inclusion of forward-facing cameras in many vehicles. Forward-facing cameras are required by the U.S. Department of Transportation to be installed in all production vehicles starting in 2018. Forward-facing cameras can be used to support ADAS technologies such as forward collision warning systems, lane departure warning systems, lane centering systems, lane keep assist systems, adaptive cruise control systems, traffic sign recognition systems, and the like. Furthermore, some vehicles are configured so that images from the forward-facing cameras can be presented on one or more displays installed in the vehicle. Having the ability to present images from the forward-facing cameras on one or more displays installed in the vehicle can be used, for example, to support ADAS technologies, to record evidence of a traffic collision, to improve visibility for the vehicle operator (e.g., when the forward-facing camera provides the operator with an extended view of one or more objects in front of the vehicle), and the like. Summary of the Invention

[0003] In embodiments, a system for determining a relationship between an object and a human-provided representation associated with the object may include a processor and a memory. The memory may store an image receiving module, a record receiving module, a relationship determination module, and a database query module. The image receiving module may include instructions that, when executed by the processor, cause the processor to receive an image that includes a representation of the object but lacks a representation of a human-provided representation associated with the object. The record receiving module may include instructions that, when executed by the processor, cause the processor to receive a record that includes a representation of the human-provided representation but lacks a representation of the object. The relationship determination module may include instructions that, when executed by the processor, cause the processor to determine the existence of a relationship between the object and the human-provided representation. The database query module may include instructions that, when executed by the processor, cause the processor to generate information about the object from a database based on a determination of the existence of a relationship in response to a query regarding the subject of the human-provided representation.

[0004] In another embodiment, a method for determining a relationship between an object and a human-provided representation associated with the object may include receiving, by a processor, an image including a representation of the object but lacking a representation of a human-provided representation associated with the object. The method may include receiving, by a processor, a record including a representation of the human-provided representation but lacking a representation of the object. The method may include determining, by the processor, the existence of a relationship between the object and the human-provided representation. The method may include causing, by the processor, a database to generate information about the object based on the determination of the existence of the relationship in response to a query regarding the subject of the human-provided representation.

[0005] In another embodiment, a non-transitory computer-readable medium for determining a relationship between an object and a human-provided representation associated with the object may include instructions that, when executed by one or more processors, cause the one or more processors to receive an image that includes a representation of the object but lacks a representation of a human-provided representation associated with the object. The non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause the one or more processors to receive a record that includes a representation of the human-provided representation but lacks a representation of the object. The non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause the one or more processors to determine the existence of a relationship between the object and the human-provided representation. The non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause the one or more processors to generate information about the object in a database based on a determination of the existence of a relationship in response to a query regarding the subject of the human-provided representation. [Brief explanation of the drawings]

[0006] The accompanying drawings, which are incorporated herein and constitute a part of this specification, illustrate various systems, methods, and other embodiments of the present disclosure. It will be understood that the boundaries of elements shown in the figures (e.g., boxes, groups of boxes, or other shapes) represent one embodiment of the boundaries. In some embodiments, one element may be designed as multiple elements, or multiple elements may be designed as one element. In some embodiments, an element shown as an internal component of another element may be implemented as an external component, and vice versa. Additionally, elements may not be drawn to scale.

[0007] [Figure 1] FIG. 1 includes a diagram illustrating an example of an environment at a prior time for determining relationships between objects and human-provided representations associated with the objects in accordance with techniques of this disclosure. [Figure 2]FIG. 2 includes a diagram illustrating an example of an environment at a later time for determining relationships between objects and human-provided representations associated with the objects in accordance with techniques of this disclosure. [Figure 3] FIG. 3 is a block diagram illustrating an example system for determining relationships between objects and human-provided representations associated with the objects, consistent with techniques of this disclosure. [Figure 4A] FIG. 4A includes a flow diagram illustrating an example method associated with determining a relationship between an object and a human-provided representation associated with the object, in accordance with techniques of this disclosure. [Figure 4B] FIG. 4B includes a flow diagram illustrating an example method associated with determining a relationship between an object and a human-provided representation associated with the object, in accordance with the techniques of this disclosure. [Figure 5] FIG. 5 includes a block diagram illustrating an example of elements located in a vehicle in accordance with the techniques of this disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0008] The techniques of this disclosure are directed to determining relationships between objects and human-provided representations associated with the objects. The techniques of this disclosure may improve database technology because the existence of a relationship between an object and a human-provided representation associated with the object may be determined without (1) having to predetermine the existence of the relationship in a database schema or (2) having to use textual information (e.g., via a keyboard interface, speech-to-text technology, or the like).

[0009] An image may be received that includes a description of an object but lacks a description of a human-provided indication associated with the object. For example, the human-provided indication may include one or more of a hand gesture, a gaze, an audible comment, or the like. For example, the image may have been generated by a camera. For example, when the image was generated, the object may be within the field of view of the camera, but the human-provided indication may be one or more of being outside the field of view of the camera or otherwise imperceptible by the camera (e.g., the human-provided indication is an audible comment). A recording may be received that includes a description of a human-provided indication but lacks a description of the object.

[0010] The existence of a relationship between an object and a human-provided representation may be determined. For example, (1) the image may have been generated by a camera at a first time, (2) the recording may have been generated at a second time, (3) the human-provided representation may refer to a particular direction, (4) the location of the object at the second time may be in a particular direction from the human who generated the human-provided representation, and (5) based on (a) information about the location of the object at the second time and (b) information about the relative movement between the camera and the object between the first time and the second time, it may be determined that the location of the object at the first time corresponds to a depiction of the object included in the image. For example, relationship information between objects and human-provided representations may be stored in a database. Information about the object may be generated from the database in response to a query about the subject of the human-provided representations.

[0011] FIG. 1 includes a diagram 100 illustrating an example of an environment at a previous time 101 for determining relationships between objects and human-provided representations associated with the objects, according to the techniques of this disclosure. For example, diagram 100 may include a first street 102 (positioned along a latitude line) and an avenue A 103 (positioned along a longitude line). For example, diagram 100 may include a road junction 104 (e.g., an intersection) of first street 102 and avenue A 103. For example, first street 102 may include a right westbound lane 105, a left westbound lane 106, a left eastbound lane 107, and a right eastbound lane 108. For example, avenue A 103 may include a right southbound lane 109, a left southbound lane 110, a left northbound lane 111, and a right northbound lane 112. For example, diagram 100 may include bar 113 at the southeast corner of road junction 104. For example, diagram 100 may include cloud computing platform 114. For example, cloud computing platform 114 may include communication device 115 and data storage 116.

[0012] For example, diagram 100 may include a first vehicle 117, a second vehicle 118, a third vehicle 119, and a fourth vehicle 120. For example, second vehicle 118 may include one or more of a driver's seat 121, a front passenger seat 122, a processor 123, memory 124, data storage 125, a communication device 126, a forward-facing camera 127, a rear-facing camera 128, a cabin-view camera 129, a microphone 130, or a display 131. For example, Asher 132 may be in the driver's seat 121 and Bryce 133 may be in the front passenger seat 122. For example, third vehicle 119 may include one or more of driver's seat 134, front passenger seat 135, processor 136, memory 137, data storage 138, communication device 139, forward-facing camera 140, rear-facing camera 141, cabin-view camera 142, microphone 143, or display 144. For example, Chad 145 may be in driver's seat 134 and Dylan 146 may be in front passenger seat 135. For example, fourth vehicle 120 may include one or more of driver's seat 147, front passenger seat 148, processor 149, memory 150, data storage 151, communication device 152, forward-facing camera 153, rear-facing camera 154, cabin-view camera 155, microphone 156, or display 157. For example, Evan 158 may be in driver's seat 147 and Forrest 159 may be in front passenger seat 148. For example, the diagram 100 may include a pedestrian 160, Grace Worthington.

[0013] For example, diagram 100 may include a first pair of rays 161 and 162, a second pair of rays 163 and 164, and a third pair of rays 165 and 166. For example, first pair of rays 161 and 162 may define a field of view 167 of rear-facing camera 128 of second vehicle 118. For example, second pair of rays 163 and 164 may define a field of view 168 of forward-facing camera 140 of third vehicle 119. For example, third pair of rays 165 and 166 may define a field of view 169 of forward-facing camera 153 of fourth vehicle 120.

[0014] For example, at the previous time 101, (1) a first vehicle 117 may be located in the left westbound lane 106 immediately west of road junction 104 and traveling in a westward direction, (2) a second vehicle 118 may be located in the left eastbound lane 107 immediately west of road junction 104 and immediately south of first vehicle 117 and traveling in an eastward direction, (3) a third vehicle 119 may be located in the right northbound lane 112 approximately 30 meters south of road junction 104 and traveling in a northward direction, (4) a fourth vehicle 120 may be located in the right northbound lane 112 approximately 60 meters south of road junction 104 and parked, and (5) a pedestrian 160 may be located on the sidewalk east of the right northbound lane 112 approximately 5 meters north of fourth vehicle 120 and traveling in a southward direction. For example, diagram 100 may include a depiction 170 of the location of first vehicle 117 before the previous time 101 using dashed lines.

[0015] FIG. 2 includes a diagram 200 illustrating an example environment at a later time 201 for determining relationships between objects and human-provided representations associated with the objects in accordance with techniques of this disclosure. For example, at a later time 201, (1) a first vehicle 117 may be located in the left westbound lane 106 approximately 60 meters west of road junction 104 and traveling in a westward direction, (2) a second vehicle 118 may be located in the left eastbound lane 107 within road junction 104 and traveling in an eastward direction, (3) a third vehicle 119 may be located in the right northbound lane 112 just south of road junction 104 and just west of bar 113 and traveling in a northward direction, (4) a fourth vehicle 120 may be located in the right northbound lane 112 approximately 60 meters south of road junction 104 and parked, and (5) a pedestrian 160 may be located on the sidewalk just east of the fourth vehicle 120 and east of the right northbound lane 112 and traveling in a southward direction.

[0016] 3 is a block diagram illustrating an example system 300 for determining relationships between objects and human-provided representations associated with the objects, in accordance with techniques of this disclosure. System 300 may include, for example, a processor 302 and a memory 304. Memory 304 may be communicatively coupled to processor 302. For example, memory 304 may store an image receiving module 306, a record receiving module 308, a relationship determination module 310, and a database query module 312.

[0017] For example, the image receiving module 306 may include instructions that function to control the processor 302 to receive an image that includes a depiction of an object but lacks a depiction of a human-provided indicia associated with the object.

[0018] For example, the recording receiving module 308 may include instructions that operate to control the processor 302 to receive a recording that includes a depiction of a human-provided representation but lacks a depiction of an object.

[0019] For example, the image may be generated by a camera. For example, the recording may be an audio recording. Alternatively or additionally, the image may be a first image and the recording may be a second image. For example, the first image may have been generated by a first camera and the second image may have been generated by a second camera. For example, the first camera may be one or more of a forward-facing camera disposed on the vehicle or a rear-facing camera disposed on the vehicle. For example, the second camera may be a cabin-view camera disposed on the vehicle.

[0020] For example, the image may have been generated at a first time and the recording may have been generated at a second time. For example, the second time may be later than the first time. Alternatively, for example, the second time may be earlier than the first time.

[0021] For example, the human-provided indication may include one or more of a hand gesture, a gaze, an audible comment, or the like. For example, the hand gesture may be a gesture pointing in a particular direction, the gaze may be a particular direction, the audible comment may include information signifying a particular direction, or the like. Additionally or alternatively, the hand gesture may represent the perspective of the human who generated the hand gesture, the gaze may represent the perspective of the human who generated the gaze, the audible comment may represent the perspective of the human who generated the audible comment, or the like.

[0022] For example, the relationship determination module 310 may include instructions that control the processor 302 to determine the existence of a relationship between an object and a human-provided representation. For example, (1) the image may have been generated by a camera at a first time, (2) the recording may have been generated at a second time, (3) the human-provided representation may refer to a particular direction, (4) the location of the object at the second time may be in a particular direction from the human who generated the human-provided representation, and (5) the instructions for determining the existence of a relationship may include instructions for determining that the location of the object at the first time corresponds to a representation of the object included in the image based on (a) information about the location of the object at the second time and (b) information about relative movement between the camera and the object between the first time and the second time. For example, the relative movement may include one or more of movement of a camera (e.g., included in a vehicle) or movement of the object.

[0023] Further, for example, the memory 304 may further store a hand gesture module 314. For example, the hand gesture module 314 may include instructions that function to control the processor 302 to effect operation of a hand gesture technique in response to a human-provided indication that includes a hand gesture. For example, the hand gesture technique may include (1) operating gesture recognition technology to determine that the hand gesture is a gesture that points in a particular direction, and (2) generating a hand gesture vector in the particular direction. For example, the origin of the hand gesture vector may be the hand positioned to generate the hand gesture.

[0024] Alternatively or additionally, for example, the memory 304 may further store a gaze module 316. For example, the gaze module 316 may include instructions that function to control the processor 302 to effect operation of a gaze technique in response to a human-provided indication including a gaze. For example, the gaze technique may include (1) operating an eye gaze point tracking technique to determine that the gaze point of the eye is in a particular direction, and (2) generating a gaze vector in the particular direction. For example, the origin of the gaze vector may be the eye.

[0025] 1 and 2, for example, with respect to the third vehicle 119, (1) the image may be a first image generated at an earlier time 101 by the forward-facing camera 140 that includes a depiction of the bar 113, (2) the recording may be a second image generated at a later time 201 by the cabin-view camera 142 that includes a depiction of Dylan 146 generating a hand gesture with his right hand that is a gesture pointing in the direction of the bar 113, (3) a hand gesture vector 202 may be generated from Dylan 146's right hand toward the bar 113, and (4) the existence of a relationship between the bar 113 and the hand gesture may be determined by determining that the location of the bar 113 at the earlier time 101 corresponds to the depiction of the bar 113 included in the first image based on (a) information regarding the location of the bar 113 at the later time 201 and (b) information regarding the relative movement between the forward-facing camera 140 and the bar 113 between the earlier time 101 and the later time 201. Further, for example, the second image may include a depiction of Dylan 146 making a hand gesture with his left hand that is a thumbs-up gesture signifying that Dylan 146 has a positive opinion of the bar 113.

[0026] For example, with respect to the fourth vehicle 120, (1) the image may be a first image generated at an earlier time 101 by the forward-facing camera 153, including a depiction of the pedestrian 160; (2) the record may be a second image generated at a later time 201 by the cabin-view camera 155, including a depiction of the forest 159 gazing in the direction of the pedestrian 160; (3) a gaze vector 203 may be generated from the eyes of the forest 159 toward the pedestrian 160; and (4) the existence of a relationship between the pedestrian 160 and the gaze may be determined by determining that the location of the pedestrian 160 at the earlier time 101 corresponds to the depiction of the pedestrian 160 included in the first image, based on (a) information regarding the location of the pedestrian 160 at the later time 201, and (b) information regarding the relative movement between the forward-facing camera 153 and the pedestrian 160 between the earlier time 101 and the later time 201. Further, for example, the techniques of the present disclosure may employ emotion recognition technology to determine that a depiction in the second image of Forest 159 gazing in the direction of pedestrian 160 means that Forest 159 has a positive view of pedestrian 160.

[0027] For example, with respect to the second vehicle 118, (1) the image may be an image generated at a later time 201 by the rear-facing camera 128 that includes a depiction of the first vehicle 117, (2) the recording may be an audio recording generated at an earlier time 101 by the microphone 130 that includes a depiction of Bryce 133 making an audible comment that includes information implying the direction of the first vehicle 117 (e.g., "Look at that car to our left!"), and (3) the existence of a relationship between the first vehicle 117 and the audible comment may be determined by determining that the location of the first vehicle 117 at the later time 201 corresponds to the depiction of the first vehicle 117 included in the image based on (a) information regarding the location of the first vehicle 117 at the earlier time 101 and (b) information regarding the relative movement between the rear-facing camera 128 and the first vehicle 117 between the earlier time 101 and the later time 201. Further, for example, the techniques of the present disclosure may employ emotion recognition technology to determine that a depiction of Blythe 133 making an audible comment contained in an audio recording means that Blythe 133 has a negative opinion about the operator of the first vehicle 117 (e.g., "That guy is driving recklessly!").

[0028] 3 , for example, memory 304 may further store a database relationship establishment module 318. For example, database relationship establishment module 318 may include instructions that function to control processor 302 to store relationship information in a database. For example, the relationship information may include (1) a record that includes a depiction of a human-provided representation, (2) information regarding the existence of a relationship between an object and a human-provided representation, and (3) one or more of (a) an image that includes a depiction of the object or (b) another image that includes a depiction of the object.

[0029] For example, the instructions for storing the relevant information in a database may include instructions for storing the relevant information in a database. For example, the database may be stored in a data storage located in the vehicle. For example, the system 300 may further include a data storage 320. The data storage 320 may be communicatively connected to the processor 302. For example, the data in the database may be intended to be a private database for use by an occupant of the vehicle.

[0030] Alternatively or additionally, for example, the instructions for storing the relationship information in a database may include (1) instructions for transmitting the relationship information to a cloud computing platform and (2) instructions for storing the relationship information in the database. For example, the database may be stored in data storage located on the cloud computing platform. For example, the data in the database may be intended to be a public database for use by individuals authorized to access the cloud computing platform.

[0031] Further, for example, memory 304 may further store an image annotation module 322. For example, image annotation module 322 may include instructions operable to control processor 302 to annotate one or more of (1) an image including a depiction of an object, or (2) another image including a depiction of an object, with supplemental information. For example, the supplemental information may be based on one or more of (1) information implied by a human-provided representation, or (2) information generated contemporaneously with generation of the human-provided representation. For example, the supplemental information may include one or more of information regarding the identity of the object, information regarding the location of the object, information regarding features of the object, information regarding characteristics of features of the object, information regarding an opinion about the object, or the like.

[0032] 1 and 2 , for example, an image including a depiction of bar 113 with respect to third vehicle 119 may be annotated with supplemental information. For example, the supplemental information may be based on a hand gesture produced by Dylan 146's left hand, the hand gesture being a thumbs-up gesture signifying that Dylan 146 has a positive opinion of bar 113. Further, for example, the supplemental information may be based on an audio recording produced by microphone 143 simultaneously with the production of a hand gesture produced by Dylan 146's right hand, the hand gesture being a gesture pointing in the direction of bar 113, the hand gesture including a depiction of Dylan 146 making an audible comment including one or more of information regarding the location of bar 113, information regarding a feature of bar 113, information regarding a characteristic of a feature of bar 113, information regarding an opinion of bar 113, or the like (e.g., "That bar on the corner of 1st Street and Avenue A with the red double doors is great!").

[0033] For example, with respect to the fourth vehicle 120, an image including a depiction of pedestrian 160 may be annotated with supplemental information. For example, the supplemental information may be based on results generated by emotion recognition technology that indicated that Forrest 159 had a positive opinion of pedestrian 160. Further, for example, the supplemental information may be based on an audio recording generated by microphone 156 simultaneously with Forest 159 directing a gaze in the direction of pedestrian 160, including a depiction of Evan 158 making an audible comment (e.g., "Isn't that Grace Worthington?") that includes information regarding the identity of pedestrian 160. Further, for example, the technology of this disclosure may operate facial recognition technology to determine the identity of pedestrian 160 as Grace Worthington based on the depiction and supplemental information regarding pedestrian 160 in the image.

[0034] For example, with respect to the second vehicle 118, an image including a depiction of the first vehicle 117 may be annotated with supplemental information. For example, the supplemental information may be based on results generated by emotion recognition technology that meant Blythe 133 had a negative opinion about the operator of the first vehicle 117. Further, for example, the technology of the present disclosure may operate automatic number plate recognition (ANPR) technology that reads characters on the vehicle registration plate (i.e., license plate) of the first vehicle 117. Further, for example, the supplemental information may be based on results generated by the ANPR technology that include characters on the license plate of the first vehicle 117.

[0035] 3, alternatively or additionally, for example, memory 304 may further store an object recognition and classification module 324. For example, object recognition and classification module 324 may include instructions that function to control processor 302 to recognize and classify objects.

[0036] Alternatively or additionally, for example, the memory 304 may further store an image transformation module 326. For example, the image transformation module 326 may include instructions that control the processor 302 to generate another image including a representation of the object based on an image including a representation of the object. For example, the another image may be a transformation of the image. For example, (1) the representation of the object in the image may be associated with a first perspective of the object, and (2) the representation of the object in the other image may be associated with a second perspective of the object. For example, the second perspective may be the perspective of a human who generated the human-provided representation when the recording including the representation of the human-provided representation was generated. Alternatively, for example, the second perspective may be the perspective at which a measure of the recognizability of the object is greatest. Alternatively, for example, the second perspective may be the perspective of an image generated by another camera. For example, the other camera may be another camera on the (original) vehicle or a camera on another vehicle. For example, if the other camera is a camera on another vehicle, the image generated by the other camera may be communicated to the (original) vehicle.

[0037] Alternatively or additionally, for example, an image including a depiction of an object may be a member of a set of images including a depiction of the object. For example, the camera that generated the image may be configured to generate images at a particular generation rate. For example, the particular generation rate may be 10 Hz. For example, the memory 304 may further store an image quality measurement module 328 and an image designation module 330. For example, the image quality measurement module 328 may include instructions that control the processor 302 to determine an image from the set of images that has a maximum value for the image quality measurement of the object. For example, the image designation module 330 may include instructions that control the processor 302 to designate the image from the set of images that has a maximum value for the image quality measurement of the object as another image that includes a depiction of the object. For example, the database relationship establishment module 318 may include instructions that control the processor 302 to include another image in the relationship information stored in the database.

[0038] Alternatively, or in addition, for example, memory 304 may further store a map augmentation module 332. For example, map augmentation module 332 may include instructions that function to control processor 302 to include valid map information in a map related to the vicinity of an object's location. For example, the valid map information may include one or more of: (1) (a) an image including a depiction of the object or (b) another image including a depiction of the object; or (2) supplemental information. For example, the supplemental information may be based on one or more of: (a) information implied by a human-provided representation; or (b) information generated contemporaneously with the generation of the human-provided representation.

[0039] For example, the database query module 312 may include instructions that function to control the processor 302 to cause the database to generate information about objects based on a determination of the existence of a relationship in response to a query about the subject matter of a human-provided representation.

[0040] Additionally, for example, the memory 304 may further store a display presentation module 334. For example, the display presentation module 334 may include instructions that function to control the processor 302 to cause a display to present information about an object generated in response to a human-provided query about the subject of the display. For example, the display may be located in a vehicle.

[0041] 4A and 4B include a flow diagram illustrating an example method 400 associated with determining a relationship between an object and a human-provided representation associated with the object, in accordance with the techniques of this disclosure. Although method 400 is described in conjunction with system 300 shown in FIG. 3, in light of the description herein, it will be understood by those skilled in the art that method 400 is not limited to implementation by system 300 shown in FIG. 3. Rather, system 300 shown in FIG. 3 is an example of a system that may be used to implement method 400. Furthermore, while method 400 is depicted as a generally sequential process, various aspects of method 400 may be capable of being performed in parallel.

[0042] In FIG. 4A, in method 400, at operation 402, for example, the image receiving module 306 may receive an image that includes a depiction of an object but lacks a depiction of a human-provided indication associated with the object.

[0043] At operation 404, for example, the recording receiving module 308 may include instructions that operate to control the processor 302 to receive a recording that includes a depiction of a human-provided representation but lacks a depiction of an object.

[0044] For example, the image may be generated by a camera. For example, the recording may be an audio recording. Alternatively or additionally, the image may be a first image and the recording may be a second image. For example, the first image may have been generated by a first camera and the second image may have been generated by a second camera. For example, the first camera may be one or more of a forward-facing camera disposed on the vehicle or a rear-facing camera disposed on the vehicle. For example, the second camera may be a cabin-view camera disposed on the vehicle.

[0045] For example, the image may have been generated at a first time and the recording may have been generated at a second time. For example, the second time may be later than the first time. Alternatively, for example, the second time may be earlier than the first time.

[0046] For example, the human-provided indication may include one or more of a hand gesture, a gaze, an audible comment, or the like. For example, the hand gesture may be a gesture pointing in a particular direction, the gaze may be a particular direction, the audible comment may include information signifying a particular direction, or the like. Additionally or alternatively, the hand gesture may represent the perspective of the human who generated the hand gesture, the gaze may represent the perspective of the human who generated the gaze, the audible comment may represent the perspective of the human who generated the audible comment, or the like.

[0047] At operation 406, for example, the relationship determination module 310 may determine the existence of a relationship between the object and the human-provided representation. For example, (1) the image may have been generated by a camera at a first time, (2) the recording may have been generated at a second time, (3) the human-provided representation may refer to a particular direction, (4) the location of the object at the second time may be in a particular direction from the human who generated the human-provided representation, and (5) the instructions for determining the existence of a relationship may include instructions for determining that the location of the object at the first time corresponds to a representation of the object included in the image based on (a) information about the location of the object at the second time and (b) information about relative movement between the camera and the object between the first time and the second time. For example, the relative movement may include one or more of movement of a camera (e.g., included in a vehicle) or movement of the object.

[0048] In FIG. 4B, in method 400, at operation 408, for example, the database query module 312 may cause a database to generate information about the object based on a determination of the existence of a relationship in response to a query regarding the subject matter of the human-provided representation.

[0049] 4A , method 400 may further include, at operation 410, for example, the hand gesture module 314 may effect operation of a hand gesture technique in response to a human-provided indication including a hand gesture. For example, the hand gesture technique may include (1) operating gesture recognition technology to determine that the hand gesture is a gesture pointing in a particular direction, and (2) generating a hand gesture vector in the particular direction. For example, the origin of the hand gesture vector may be the hand positioned to generate the hand gesture.

[0050] Alternatively or additionally, at operation 412, for example, the gaze module 316 may effect manipulation of a gaze technique in response to a human-provided indication including a gaze. For example, the gaze technique may include (1) manipulating an eye gaze point tracking technique to determine that the gaze point of the eye is in a particular direction, and (2) generating a gaze vector in the particular direction. For example, the origin of the gaze vector may be the eye.

[0051] Further, in operation 414, for example, the database relationship establishment module 318 may store relationship information in the database. For example, the relationship information may include (1) a record including a depiction of the human-provided representation, (2) information regarding the existence of a relationship between the object and the human-provided representation, and (3) one or more of (a) an image including a depiction of the object or (b) another image including a depiction of the object.

[0052] For example, the database relationship establishment module 318 may store the relationship information in a database in operation 414. For example, the database may be stored in data storage located in the vehicle.

[0053] Alternatively, or in addition, at operation 414, for example, the database relationship establishment module 318 may (1) send the relationship information to a cloud computing platform and (2) store the relationship information in a database. For example, the database may be stored in data storage located on the cloud computing platform.

[0054] 4B , method 400 may further include, at operation 416, for example, causing image annotation module 322 to annotate one or more of (1) an image including a depiction of the object, or (2) another image including a depiction of the object, with supplemental information. For example, the supplemental information may be based on one or more of (1) information implied by a human-provided representation, or (2) information generated contemporaneously with the generation of the human-provided representation. For example, the supplemental information may include one or more of information regarding the identity of the object, information regarding the location of the object, information regarding features of the object, information regarding characteristics of features of the object, information regarding an opinion about the object, or the like.

[0055] Alternatively or additionally, at operation 418, for example, the object recognition and classification module 324 may recognize and classify the object.

[0056] Alternatively or additionally, in operation 420, for example, the image transformation module 326 may generate another image including a representation of the object based on an image including a representation of the object. For example, the another image may be a transformation of the image. For example, (1) the representation of the object in the image may be associated with a first perspective of the object, and (2) the representation of the object in the other image may be associated with a second perspective of the object. For example, the second perspective may be the perspective of a human who generated the human-provided representation when the record including the representation of the human-provided representation was generated. Alternatively, for example, the second perspective may be the perspective at which a measure of the recognizability of the object is greatest. Alternatively, for example, the second perspective may be the perspective of an image generated by another camera. For example, the other camera may be another camera on the (original) vehicle or a camera on another vehicle. For example, if the other camera is a camera on another vehicle, the image generated by the other camera may be communicated to the (original) vehicle.

[0057] Alternatively or additionally, for example, an image including a depiction of an object may be a member of a set of images including a depiction of the object. For example, the camera that generated the image may be configured to generate images at a particular generation rate. For example, the particular generation rate may be 10 Hertz. For example, at operation 422, the image quality measurement module 328 may determine the image from the set of images that has the greatest value for the measured image quality of the object. For example, at operation 424, the image designation module 330 may designate the image from the set of images that has the greatest value for the measured image quality of the object as another image that includes a depiction of the object. For example, the database relationship establishment module 318 may include the other image in the relationship information stored in the database.

[0058] Alternatively or additionally, at operation 426, for example, map augmentation module 332 may include valid map information in the map related to the vicinity of the object's location. For example, the valid map information may include one or more of: (1) (a) an image including a depiction of the object or (b) another image including a depiction of the object; or (2) supplemental information. For example, the supplemental information may be based on one or more of: (a) information implied by a human-provided representation; or (b) information generated contemporaneously with the generation of the human-provided representation.

[0059] Further, in operation 428, for example, the display presentation module 334 may cause a display to present information about the object generated in response to a human-provided query about the subject of the display. For example, the display may be located in a vehicle.

[0060] FIG. 5 includes a block diagram illustrating example elements located in a vehicle 500 in accordance with the techniques of this disclosure. As used herein, a “vehicle” may be any form of motorized transportation. In one or more implementations, the vehicle 500 may be an automobile. While the mechanisms described herein relate to automobiles, in light of the description herein, it will be understood by those skilled in the art that the embodiments are not limited to automobiles. For example, the functionality and / or operation of one or more of the second vehicle 118 (shown in FIGS. 1 and 2), the third vehicle 119 (shown in FIGS. 1 and 2), or the fourth vehicle 120 (shown in FIGS. 1 and 2) may be implemented by the vehicle 500.

[0061] In some embodiments, vehicle 500 may be configured to selectively switch between an automatic mode, one or more semi-automatic operating modes, and / or a manual mode. Such switching may be implemented in any suitable manner now known or later developed. As used herein, "manual mode" may refer to all or most of the navigation and / or operation of vehicle 500 being performed according to input received from a user (e.g., a human driver). In one or more arrangements, vehicle 500 may be a conventional vehicle configured to operate exclusively in manual mode.

[0062] In one or more embodiments, the vehicle 500 may be an automated vehicle. As used herein, "automated vehicle" may refer to a vehicle operating in an automated mode. As used herein, "automated mode" may refer to using one or more computing systems to control the vehicle 500 and navigate and / or operate the vehicle 500 along a travel route with minimal or no input from a human driver. In one or more embodiments, the vehicle 500 may be highly automated or fully automated. In one embodiment, the vehicle 500 may be configured with one or more semi-automated modes of operation in which one or more computing systems perform a portion of the navigation and / or operation of the vehicle along a travel route, and a vehicle operator (i.e., the driver) provides input to the vehicle 500 to perform a portion of the navigation and / or operation of the vehicle 500 along the travel route.

[0063] For example, the taxonomy and definitions of terms related to driving automation systems for road vehicles in standard J3016 202104, published by the Society of Automotive Engineers (SAE) International on January 16, 2014, and most recently revised on April 30, 2021, defines six levels of driving automation: (1) Level 0, no automation, where all aspects of dynamic driving tasks are performed by a human driver; (2) Level 1, driver assistance, where driver assistance systems, if selected, may perform either steering or acceleration / deceleration tasks using information about the driving environment, but all remaining dynamic driving tasks are performed by a human driver; and (3) Level 2, partial automation, where one or more driver assistance systems, if selected, may perform both steering and acceleration / deceleration tasks using information about the driving environment, but all remaining dynamic driving tasks are performed by a human driver. (4) Level 3, conditional automation, where the automated driving system may, if selected, perform all aspects of the DDT with the possibility that a human driver may respond appropriately to a request to intervene; (5) Level 4, high automation, where the automated driving system may, if selected, perform all aspects of the DDT even if the human driver does not respond appropriately to a request to intervene; and (6) Level 5, full automation, where the automated driving system may perform all aspects of the DDT under all roadway and environmental conditions that can be managed by a human driver.

[0064] Vehicle 500 may include various elements. Vehicle 500 may have any combination of the various elements shown in FIG. 5 . In various embodiments, vehicle 500 may not necessarily include all of the elements shown in FIG. 5 . Furthermore, vehicle 500 may have elements other than those shown in FIG. 5 . While various elements are shown in FIG. 5 as being located within vehicle 500, one or more of the elements may be located outside vehicle 500. Furthermore, the elements shown may be physically separated by large distances. For example, as described, one or more components of the system of the present disclosure may be implemented within vehicle 500, while other components of the system may be implemented within a cloud computing environment, as described below. For example, the elements may include one or more processors 510, one or more data stores 515, a sensor system 520, an input system 530, an output system 535, a vehicle system 540, one or more actuators 550, one or more autonomous driving modules 560, a communication system 570, and a system 300 for determining relationships between objects and human-provided representations associated with the objects.

[0065] In one or more arrangements, one or more processors 510 may be a main processor of vehicle 500. For example, one or more processors 510 may be an electronic control unit (ECU). For example, the functionality and / or operation of one or more of processor 123 (shown in FIGS. 1 and 2), processor 136 (shown in FIGS. 1 and 2), processor 149 (shown in FIGS. 1 and 2), or processor 302 (shown in FIG. 3) may be implemented by one or more processors 510.

[0066] The one or more data stores 515 may, for example, store one or more types of data. The one or more data stores 515 may include volatile memory and / or non-volatile memory. Examples of suitable memory for the one or more data stores 515 may include random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, any other suitable storage medium, or any combination thereof. The one or more data stores 515 may be components of the one or more processors 510. Additionally or alternatively, the one or more data stores 515 may be operably connected to the one or more processors 510 for use by the one or more processors 510. As used herein, "operably connected" includes direct or indirect connections and may include connections without direct physical contact. As used herein, the statement that a component may be "configured to" perform an operation may be understood to mean that the component does not require structural modification, but merely needs to be placed in an operational state to perform the operation (e.g., provided with power, having basic operating systems running, etc.). For example, the functionality and / or operation of one or more of memory 124 (shown in FIGS. 1 and 2), data storage 125 (shown in FIGS. 1 and 2), memory 137 (shown in FIGS. 1 and 2), data storage 138 (shown in FIGS. 1 and 2), memory 150 (shown in FIGS. 1 and 2), data storage 151 (shown in FIGS. 1 and 2), memory 304 (shown in FIG. 3), or data storage 320 (shown in FIG. 3) may be realized by one or more data stores 515.

[0067] In one or more arrangements, one or more data stores 515 may store map data 516. The map data 516 may include maps of one or more geographic areas. In some examples, the map data 516 may include information or data about roads, traffic control devices, road markings, structures, features, and / or landmarks within one or more geographic areas. The map data 516 may be in any suitable form. In some examples, the map data 516 may include aerial photographs of an area. In some examples, the map data 516 may include ground photographs of an area, including 360-degree ground photographs. The map data 516 may include measurements, dimensions, distances, and / or information about one or more items included in the map data 516 and / or relative to other items included in the map data 516. The map data 516 may include digital maps with information about road geometry. The map data 516 may be of high quality and / or high definition.

[0068] In one or more arrangements, the map data 516 may include one or more terrain maps 517. The one or more terrain maps 517 may include information about the ground, terrain, roads, terrain surfaces, and / or other features of one or more geographic areas. The one or more terrain maps 517 may include elevation data for one or more geographic areas. The map data 516 may be of high quality and / or high definition. The one or more terrain maps 517 may define one or more ground surfaces, which may include paved roads, unpaved roads, land, and other surfaces that define a ground surface.

[0069] In one or more arrangements, the map data 516 may include one or more stationary obstacle maps 518. The one or more stationary obstacle maps 518 may include information about one or more stationary obstacles located within one or more geographic areas. A "stationary obstacle" may be a physical object whose position does not change (or does not substantially change) over a period of time and / or whose size does not change (or does not substantially change) over a period of time. Examples of stationary obstacles may include trees, buildings, curbs, fences, railings, centerlines, utility poles, statues, monuments, signs, benches, furniture, mailboxes, large rocks, and hills. A stationary obstacle may be an object that extends above ground level. One or more stationary obstacles included in the one or more stationary obstacle maps 518 may have location data, size data, dimension data, material data, and / or other data associated therewith. The one or more stationary obstacle maps 518 may include measurements, dimensions, distances, and / or information about one or more stationary obstacles. The one or more static obstacle maps 518 may be of high quality and / or high definition. The one or more static obstacle maps 518 may be updated to reflect changes in the map area.

[0070] In one or more arrangements, one or more data stores 515 may store sensor data 519. As used herein, "sensor data" may refer to any information about sensors that the vehicle 500 may include, including capabilities and other information about the sensors. The sensor data 519 may relate to one or more sensors of the sensor system 520. For example, in one or more arrangements, the sensor data 519 may include information about one or more LiDAR sensors 524 of the sensor system 520.

[0071] In some arrangements, at least a portion of the map data 516 and / or sensor data 519 may be located in one or more data stores 515 located onboard the vehicle 500. Additionally or alternatively, at least a portion of the map data 516 and / or sensor data 519 may be located in one or more data stores 515 located remotely from the vehicle 500.

[0072] The sensor system 520 may include one or more sensors. As used herein, a "sensor" may refer to any device, component, and / or system that can detect and / or sense something. The one or more sensors may be configured to detect and / or sense in real time. As used herein, the term "real time" may refer to a level of processing responsiveness that allows a particular process or decision to be made to be recognized quickly enough by a user or system or to allow a processor to keep up with some external process.

[0073] In arrangements where the sensor system 520 includes multiple sensors, the sensors may function independently of one another. Alternatively, two or more of the sensors may function in combination with one another. In such cases, the two or more sensors may form a sensor network. The sensor system 520 and / or one or more sensors may be operatively connected to one or more processors 510, one or more data stores 515, and / or another element (including any of the elements shown in FIG. 5) of the vehicle 500. The sensor system 520 may acquire data regarding at least a portion of the vehicle 500's external environment (e.g., nearby vehicles). The sensor system 520 may include any suitable type of sensor. Various examples of different types of sensors are described herein. However, it will be understood by those skilled in the art that the embodiments are not limited to the specific sensors described herein.

[0074] The sensor system 520 may include one or more vehicle sensors 521. The one or more vehicle sensors 521 may detect, determine, and / or sense information about the vehicle 500 itself. In one or more arrangements, the one or more vehicle sensors 521 may be configured to detect and / or sense changes in the position and orientation of the vehicle 500, for example, based on inertial acceleration. In one or more arrangements, the one or more vehicle sensors 521 may include one or more accelerometers, one or more gyroscopes, an inertial measurement unit (IMU), a dead reckoning system, a global navigation satellite system (GNSS), a global positioning system (GPS), a navigation system 547, and / or other suitable sensors. The one or more vehicle sensors 521 may be configured to detect and / or sense one or more characteristics of the vehicle 500. In one or more arrangements, the one or more vehicle sensors 521 may include a speedometer to determine the current speed of the vehicle 500.

[0075] Additionally or alternatively, sensor system 520 may include one or more environmental sensors 522 configured to acquire and / or detect driving environment data. As used herein, "driving environment data" may include data or information regarding the external environment in which the vehicle is located, or one or more portions thereof. For example, one or more environmental sensors 522 may be configured to detect, quantify, and / or sense obstacles in at least a portion of the external environment of vehicle 500, and / or information / data regarding such obstacles. Such obstacles may be stationary objects and / or dynamic objects. One or more environmental sensors 522 may be configured to detect, measure, quantify, and / or sense other things in the external environment of vehicle 500, such as, for example, lane markers, signs, traffic lights, traffic signals, lanes, crosswalks, curbs near vehicle 500, off-road objects, etc. For example, the functionality and / or operation of sensor 230 (shown in FIG. 2 ) may be implemented by one or more environmental sensors 522.

[0076] Described herein are various examples of sensors for sensor system 520. Example sensors may be part of one or more vehicle sensors 521 and / or one or more environmental sensors 522. However, it will be understood by those skilled in the art that embodiments are not limited to the particular sensors described.

[0077] In one or more arrangements, the one or more environmental sensors 522 may include one or more radar sensors 523, one or more LiDAR sensors 524, one or more sonar sensors 525, and / or one or more cameras 526. In one or more arrangements, the one or more cameras 526 may be one or more high dynamic range (HDR) cameras or one or more infrared (IR) cameras. For example, the one or more cameras 526 may be used to record real-world conditions related to items of information that may appear on a digital map. For example, the functions and / or operations of one or more of forward-facing camera 127 (shown in Figures 1 and 2), rear-facing camera 128 (shown in Figures 1 and 2), cabin-view camera 129 (shown in Figures 1 and 2), forward-facing camera 140 (shown in Figures 1 and 2), rear-facing camera 141 (shown in Figures 1 and 2), cabin-view camera 142 (shown in Figures 1 and 2), forward-facing camera 153 (shown in Figures 1 and 2), rear-facing camera 154 (shown in Figures 1 and 2), or cabin-view camera 155 (shown in Figures 1 and 2) may be performed by one or more cameras 526.

[0078] Input system 530 may include any device, component, system, element, mechanism, or group thereof that allows information / data to be input to a machine. Input system 530 may receive input from a vehicle occupant (e.g., the driver or passenger). Output system 535 may include any device, component, system, element, mechanism, or group thereof that allows information / data to be presented to a vehicle occupant (e.g., the driver or passenger). For example, the functions and / or operations of one or more of microphone 130 (shown in FIGS. 1 and 2), microphone 143 (shown in FIGS. 1 and 2), or microphone 156 (shown in FIGS. 1 and 2) may be realized by input system 530. For example, the functions and / or operations of one or more of display 131 (shown in FIGS. 1 and 2), display 144 (shown in FIGS. 1 and 2), or display 157 (shown in FIGS. 1 and 2) may be realized by output system 535.

[0079] Various examples of one or more vehicle systems 540 are shown in FIG. 5 . However, it will be understood by those skilled in the art that the vehicle 500 may include more, fewer, or different vehicle systems. While certain vehicle systems may be defined separately, each or any of the systems, or portions thereof, may be otherwise combined or separated within the vehicle 500 via hardware and / or software. For example, the one or more vehicle systems 540 may include a propulsion system 541, a braking system 542, a steering system 543, a throttle system 544, a transmission system 545, a signaling system 546, and / or a navigation system 547. Each of these systems may include one or more devices, components, and / or combinations thereof, now known or later developed.

[0080] Navigation system 547 may include one or more devices, applications, and / or combinations thereof, now known or later developed, configured to determine the geographic location of vehicle 500 and / or determine travel routes for vehicle 500. Navigation system 547 may include one or more map applications that determine travel routes for vehicle 500. Navigation system 547 may include a global positioning system, a local positioning system, a geolocation system, and / or combinations thereof.

[0081] The one or more actuators 550 may be any element or combination of elements operable to modify, adjust, and / or change one or more of the vehicle systems 540 or its components in response to receiving signals or other inputs from the one or more processors 510 and / or one or more autonomous driving modules 560. Any suitable actuator may be used. For example, the one or more actuators 550 may include motors, pneumatic actuators, hydraulic pistons, relays, solenoids, and / or piezoelectric actuators.

[0082] The one or more processors 510 and / or one or more autonomous driving modules 560 may be operatively connected to communicate with various vehicle systems 540 and / or their individual components. For example, the one or more processors 510 and / or one or more autonomous driving modules 560 may communicate to send and / or receive information from the various vehicle systems 540 to control the movement, speed, operation, heading, direction, etc. of the vehicle 500. The one or more processors 510 and / or one or more autonomous driving modules 560 may control some or all of these vehicle systems 540 and, therefore, may be partially or fully autonomous.

[0083] The one or more processors 510 and / or the one or more autonomous driving modules 560 may be operable to control the navigation and / or operation of the vehicle 500 by controlling the vehicle systems 540 and / or one or more of its components. For example, when operating in an autonomous mode, the one or more processors 510 and / or the one or more autonomous driving modules 560 may control the direction and / or speed of the vehicle 500. The one or more processors 510 and / or the one or more autonomous driving modules 560 may accelerate the vehicle 500 (e.g., by increasing the supply of fuel provided to the engine), slow the vehicle 500 (e.g., by decreasing the supply of fuel to the engine and / or by applying the brakes), and / or change direction of the vehicle 500 (e.g., by turning the front two wheels). As used herein, "cause" or "causing" may mean, either directly or indirectly, to make, force, compel, direct, command, order, and / or enable an event or action to occur, or to make, force, compel, direct, command, order, and / or at least enable the event or action to be in a state in which such event or action can occur.

[0084] The communication system 570 may include one or more receivers 571 and / or one or more transmitters 572. The communication system 570 may receive and transmit one or more messages over one or more wireless communication channels. For example, the one or more wireless communication channels may comply with the Institute of Electrical and Electronics Engineers (IEEE) 802.11p standard for adding wireless access in vehicular environments (WAVE) (based on Dedicated Short-Range Communications (DSRC)), the 3rd Generation Partnership Project (3GPP®) Long Term Evolution (LTE®) Vehicle-to-Everything (V2X) (LTE®-V2X) standard (including the LTE® Uu interface between mobile communication devices and Evolved Node Bs of the Universal Mobile Telecommunications System), the 3GPP® Fifth Generation (5G) New Radio (NR) Vehicle-to-Everything (V2X) standard (including the 5G NR Uu interface), or the like. For example, the communication system 570 may include “connected vehicle” technology. “Connected vehicle” technology may include, for example, devices that exchange communications between the vehicle and other devices over a packet-switched network. The other devices may include, for example, another vehicle (e.g., “vehicle-to-vehicle” (V2V) technology), roadside infrastructure (e.g., “vehicle-to-infrastructure” (V2I) technology), cloud platforms (e.g., “vehicle-to-cloud” (V2C) technology), pedestrians (e.g., “vehicle-to-pedestrian” (V2P) technology), or networks, e.g., “vehicle-to-network” (V2N) technology. “Vehicle-to-everything” (V2X) technology may integrate aspects of these individual communication technologies. For example, the functionality and / or operation of one or more of communication device 126 (shown in FIGS. 1 and 2 ), communication device 139 (shown in FIGS. 1 and 2 ), or communication device 152 (shown in FIGS. 1 and 2 ) may be implemented by communication system 570.

[0085] Additionally, the one or more processors 510, the one or more data stores 515, and the communication system 570 may be configured to one or more of: form a micro-cloud, participate as a member of a micro-cloud, or act as a leader of a micro-cloud. A micro-cloud may be characterized by the distribution of one or more computational resources or one or more data storage resources among the members of the micro-cloud for collaboration in performing operations. The members may include at least connected vehicles.

[0086] Vehicle 500 may include one or more modules, at least some of which are described herein. The modules may be implemented as computer-readable program code that, when executed by one or more processors 510, implements one or more of the various processes described herein. One or more of the modules may be components of one or more processors 510. Additionally or alternatively, one or more of the modules may be executed on and / or distributed among other processing systems to which one or more processors 510 may be operatively connected. The modules may include instructions (e.g., program logic) executable by one or more processors 510. Additionally or alternatively, one or more data stores 515 may include such instructions.

[0087] In one or more arrangements, one or more of the modules described herein may include artificial intelligence or computational intelligence elements, such as neural networks, fuzzy logic, or other machine learning algorithms. Further, in one or more arrangements, one or more of the modules may be distributed among multiple modules described herein. In one or more arrangements, two or more of the modules described herein may be combined into a single module.

[0088] The vehicle 500 may include one or more autonomous driving modules 560. The one or more autonomous driving modules 560 may be configured to receive data from the sensor system 520 and / or from any other type of system capable of capturing information about the vehicle 500 and / or the environment external to the vehicle 500. In one or more arrangements, the one or more autonomous driving modules 560 may use the data to generate one or more driving scene models. The one or more autonomous driving modules 560 may determine the position and speed of the vehicle 500. The one or more autonomous driving modules 560 may determine the location of obstacles, obstructions, or other environmental features, including traffic signs, trees, shrubs, nearby vehicles, pedestrians, etc.

[0089] One or more autonomous driving modules 560 may be configured to receive and / or determine location information about obstacles in the external environment of the vehicle 500 for use by one or more processors 510 and / or one or more of the modules described herein, and to estimate the position and orientation of the vehicle 500, the vehicle's position in global coordinates, based on signals from multiple satellites or any other data and / or signals that may be used to determine the current state of the vehicle 500 or the position of the vehicle 500 relative to the vehicle's 500's environment used in creating a map or determining the position of the vehicle 500 relative to the map data.

[0090] The one or more autonomous driving modules 560 may be configured to determine one or more travel paths, a current autonomous driving maneuver for the vehicle 500, a future autonomous driving maneuver, and / or modifications to the current autonomous driving maneuver based on data from any other suitable sources, such as data acquired by the sensor system 520, a driving scene model, and / or determinations from the sensor data 519. As used herein, a "driving maneuver" may refer to one or more actions that affect the movement of the vehicle. Examples of driving maneuvers include accelerating, decelerating, braking, turning, moving the vehicle 500 laterally, changing lanes of travel, merging into lanes of travel, and / or reversing, just to name a few possibilities. The one or more autonomous driving modules 560 may be configured to implement the determined driving maneuvers. The one or more autonomous driving modules 560 may directly or indirectly implement such autonomous driving maneuvers. As used herein, "cause" or "causing" means, either directly or indirectly, to make, command, command to occur, and / or enable an event or action to occur, or to make, command, command, and / or at least enable an event or action to be in a state in which such event or action can occur. One or more autonomous driving modules 560 may be configured to perform various vehicle functions and / or send data to, receive data from, interact with, and / or control vehicle 500 or one or more of its systems (e.g., one or more of vehicle systems 540). For example, the functions and / or operations of an automobile navigation system may be implemented by one or more autonomous driving modules 560.

[0091] Detailed embodiments are disclosed herein. However, in light of the description herein, those skilled in the art will understand that the embodiments of the present disclosure are intended merely as examples. Therefore, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to variously employ the aspects of the present disclosure in substantially any suitable detailed configuration. Furthermore, the terms and phrases used herein are not intended to be limiting, but rather to provide an understandable description of possible implementations. While various embodiments are shown in FIGS. 1-3, 4A, 4B, and 5, the embodiments are not limited to the illustrated structures or applications.

[0092] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, comprising one or more executable instructions that implement the specified logical function(s). In light of the description herein, those skilled in the art will appreciate that in some alternative implementations, the functions noted in the blocks may occur out of the order depicted by the figures. For example, two blocks depicted in succession may in fact be executed substantially simultaneously, or the blocks may be executed in the reverse order, depending on the functionality involved.

[0093] The above-described systems, components, and / or processes can be implemented in hardware or a combination of hardware and software, either centralized within one processing system or distributed with various elements spread across several interconnected processing systems. Any type of processing system or other apparatus configured to perform the methods described herein is suitable. A typical combination of hardware and software can be a processing system having computer-readable program code that, when loaded and executed, controls the processing system such that the processing system performs the methods described herein. The systems, components, and / or processes can also be embodied in a computer-readable storage, such as a computer program product or other data program storage device, tangibly embodying a program of instructions executable by the machine to perform the methods and processes described herein. These elements can also be embodied in an application product that comprises all the features that enable implementation of the methods described herein and that, when loaded on a processing system, can execute the methods.

[0094] Furthermore, the mechanisms described herein may take the form of a computer program product in which computer-readable program code is embodied, e.g., stored, in one or more computer-readable media. Any combination of one or more computer-readable media may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. As used herein, the phrase "computer-readable storage medium" refers to a non-transitory storage medium. A computer-readable storage medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of computer-readable storage media include, in a non-exhaustive list, the following: a portable computer diskette, a hard disk drive (HDD), a solid-state drive (SSD), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), an optical storage device, a magnetic storage device, or any suitable combination thereof. As used herein, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0095] Generally, as used herein, a module includes a routine, program, object, component, data structure, etc. that performs a particular task or implements a particular data type. In a further aspect, a memory generally stores such modules. The memory associated with a module may be a buffer or may be a cache, random access memory (RAM), ROM, flash memory, or another suitable electronic storage medium incorporated within a processor. In still a further aspect, a module as used herein may be implemented as an application-specific integrated circuit (ASIC), as a hardware component of a system-on-chip (SoC), as a programmable logic array (PLA), or as another suitable hardware component (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a field-programmable gate array (FPGA), or the like) incorporating a defined set of configurations (e.g., instructions) to perform the functions of the present disclosure.

[0096] Program code embodied on a computer-readable medium may be transmitted using any appropriate medium, including, but not limited to, wireless, wired, fiber optic, cable, radio frequency (RF), etc., or any suitable combination of the above. Computer program code for performing operations for aspects of the technology of this disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, Smalltalk, C++, or the like, and traditional procedural programming languages ​​such as the “C” programming language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection to the external computer may be made (e.g., through the Internet using an Internet Service Provider).

[0097] The terms "a" and "an," as used herein, are defined as one or more than one. The term "plurality," as used herein, is defined as two or more than two. The term "another," as used herein, is defined as at least a second or more. The terms "including" and / or "having," as used herein, are defined as comprising (i.e., open language). The phrase "at least one of ... or ..." as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. For example, the phrase "at least one of A, B, or C" includes A only, B only, C only, or any combination thereof (e.g., AB, AC, BC, or ABC).

[0098] Aspects of the present specification may be embodied in other forms without departing from the spirit or essential attributes thereof, and reference should accordingly be made to the following claims, rather than the foregoing specification, as indicating the scope herein.

Claims

1. a processor; Memory and The system comprises: an image receiving module comprising instructions that, when executed by the processor, cause the processor to receive an image that includes a representation of an object but lacks a representation of a human-provided indication associated with the object; a record receiving module comprising instructions that, when executed by the processor, cause the processor to receive a record that includes the description of the human-provided representation but lacks the description of the object; a relationship determination module comprising instructions that, when executed by the processor, cause the processor to determine the existence of a relationship between the object and the human-provided representation; a database query module comprising instructions that, when executed by the processor, cause the processor to generate information about the object in a database based on the presence determination in response to a query regarding the subject matter of the human-provided representation; The system remembers this.

2. the recording is an audio recording; or The system of claim 1 , wherein the image is at least one of a first image and the recording is a second image.

3. the first image is generated by a first camera; The system of claim 2 , wherein the second image is generated by a second camera.

4. The first camera a forward-facing camera located in the vehicle, or At least one of the rear-facing cameras located on the vehicle, The system of claim 3 , wherein the second camera is a cabin view camera located in the vehicle.

5. The human-provided representation may include: Hand gestures or Staring, or The system of claim 1 , further comprising at least one of an audible commentary.

6. The hand gesture is a gesture of pointing in a specific direction, the gaze is in the particular direction; or The system of claim 5 , wherein the audible comment comprises at least one of: information that signifies the particular direction;

7. The hand gesture represents the view of a person who generated the hand gesture; said gaze refers to the view of the person who generated said gaze; or The system of claim 5 , wherein the audible comment is at least one of: representing the opinion of a person who generated the audible comment.

8. the image was generated by a camera at a first time; the record was generated at a second time; the human-provided indicia signifying a particular direction; the location of the object at the second time is in the particular direction from the human who generated the human-provided representation; The instructions for determining the existence of the relationship include: information about the location of the object at the second time; and based on information about the relative movement between the camera and the object between the first time and the second time; The system of claim 1 , further comprising instructions for determining that the location of the object at the first time corresponds to the depiction of the object contained in the image.

9. The memory includes: a hand gesture module comprising instructions that, when executed by the processor, cause the processor to perform a hand gesture technique in response to the human-provided indication including a hand gesture, the hand gesture technique comprising: operating gesture recognition technology to determine that the hand gesture is a gesture pointing in the particular direction; generating a hand gesture vector in the particular direction, the origin of the hand gesture vector being the hand positioned to generate the hand gesture; a hand gesture module including: a gaze module comprising instructions that, when executed by the processor, cause the processor to operate a gaze technique in response to the human-provided indication including a gaze, the gaze technique comprising: operating an eye gaze point tracking technique to determine that the eye gaze point is in said particular direction; generating a gaze vector in the particular direction, the gaze vector having an origin at the eye; The system of claim 8 , further storing at least one of the gaze modules comprising:

10. The memory further stores a database relationship establishment module comprising instructions that, when executed by the processor, cause the processor to store relationship information in the database, the relationship information comprising: the record including the description of the human-provided representation; information regarding the existence of the relationship between the object and the human-provided representation; and The system of claim 1 , further comprising at least one of the image including the representation of the object or another image including the representation of the object.

11. 11. The system of claim 10, wherein the memory further stores an image annotation module comprising instructions that, when executed by the processor, cause the processor to annotate the at least one of the image including the depiction of the object or the other image including the depiction of the object with supplemental information, the supplemental information being based on at least one of information implied by the human-provided representation or information generated contemporaneously with generation of the human-provided representation.

12. The system of claim 10 , wherein the memory further stores an object recognition and classification module comprising instructions that, when executed by the processor, cause the processor to recognize and classify the object.

13. 11. The system of claim 10, wherein the memory further stores an image transformation module comprising instructions that, when executed by the processor, cause the processor to generate, based on the image including the representation of the object, another image including the representation of the object, wherein the another image is a transformation of the image.

14. The image containing the depiction of the object is a member of a set of images containing the depiction of the object, The memory includes: an image quality measurement module comprising instructions that, when executed by the processor, cause the processor to determine an image from the set of images that has the greatest image quality measurement value for the object; and 11. The system of claim 10, further storing an image designation module including instructions that, when executed by the processor, cause the processor to designate the image of the set of images in which the measurement value of the image quality of the object is the maximum value as the other image containing the depiction of the object.

15. The memory further stores a map augmentation module comprising instructions that, when executed by the processor, cause the processor to include valid map information in a map relating to a neighborhood of the object's location, the valid map information comprising: the at least one of the image including the depiction of the object or the other image including the depiction of the object; or 11. The system of claim 10, further comprising at least one of supplemental information based on at least one of information implied by the human-provided representation or information generated contemporaneously with generation of the human-provided representation.

16. receiving, by a processor, an image including a representation of an object but lacking a representation of a human-provided representation associated with the object; receiving, by a processor, a record including the description of the human-provided representation but lacking the description of the object; determining, by the processor, the existence of a relationship between the object and the human-provided representation; generating, by the processor, information about the object in a database based on the presence determination in response to a query regarding the subject of the human-provided representation; A method comprising:

17. the image was generated at a first time, The method of claim 16 , wherein the record was generated at a second time.

18. 18. The method of claim 17, wherein the second time is later than the first time.

19. 18. The method of claim 17, wherein the second time is earlier than the first time.

20. 1. A non-transitory computer-readable medium for determining relationships between objects and human-provided representations associated with the objects, the non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to: receiving an image including a representation of the object but lacking a representation of the human-provided representation associated with the object; receiving a record including the description of the human-provided representation but lacking the description of the object; determining the existence of the relationship between the object and the human-provided representation; A non-transitory computer-readable medium that causes a database to generate information about the object based on the determination of the existence of the relationship in response to a query related to the subject matter of the human-provided representation.