Vehicle data association device and method therefor

JP2022151619A5Active Publication Date: 2025-12-22INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022017464
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-03-25
Filing Date
2022-02-07
Publication Date
2025-12-22
Estimated Expiration
2042-02-07

AI Technical Summary

Technical Problem

Autonomous vehicles face challenges in distinguishing between relevant and irrelevant sensor data, requiring significant training for artificial neural networks (ANNs) to accurately process and analyze large volumes of data from multiple sensors.

Method used

A system that utilizes human speech and gaze/gesture recognition to label relevant sensor data by cross-referencing instructor inputs with sensor data, employing machine learning algorithms to enhance ANN training and improve data association in autonomous vehicles.

Benefits of technology

Enhances the ability of ANNs to identify and label relevant sensor data, improving the accuracy of driving decisions by leveraging human instructor feedback, thereby refining the training process and enhancing the vehicle's perception of hazardous situations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a vehicle data association device and a method capable of determining related sensor data and non-related sensor data.SOLUTION: A vehicle data association device comprises an internal audio / image data analysis device, an external image analysis device, and an object data generation device. The internal audio / image data analysis device receives at least one of sounds or images in a vehicle (first data) and identifies in the first data the second data representing an audio indicator (a human voice associated with importance of an object outside the vehicle) or an image indicator (a human being in the vehicle associated with the importance of the object outside the vehicle). The external image analysis device receives an image (third data) outside the vehicle, identifies an object corresponding to at least one of the audio indicator and a video indicator in the third data, and the object data generation device generates data corresponding to the object.SELECTED DRAWING: Figure 13
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Various aspects of the present disclosure relate to speech recognition and object recognition from image data based on speech.

Background Art

[0002] Autonomous vehicles and semi-autonomous vehicles typically rely on multiple sensors to detect information about the vehicle's surroundings and make driving decisions based on this information. Such sensors may include, for example, multiple cameras, one or more light detection and ranging (LIDAR) systems, one or more radio detection and ranging (RADAR) systems, microphones, accelerometers, and / or position sensors. Since these sensors generate a significant amount of data, an autonomous vehicle may need to analyze this large amount of data for driving operations.

[0003] One particular challenge in processing this data is the ability to distinguish relevant sensor data from irrelevant sensor data. The use of artificial neural networks (ANNs) in the processing of sensor data and driving decision-making is increasing. Artificial neural networks may be particularly well-suited to this task since they can be configured to receive large amounts of data and analyze it quickly.

[0004] However, a significant amount of training is required to successfully implement an ANN for the analysis of such sensor data. One particularly difficult task is teaching the ANN to distinguish relevant sensor data from irrelevant sensor data. Put another way, while a human driver may be able to relatively easily distinguish relevant visual or auditory information, an ANN may not be able to distinguish it without further training.

Brief Description of the Drawings

[0005] In drawings, similar reference numerals generally refer to the same part across different drawings. Drawings are not necessarily to scale and are generally focused on illustrating exemplary principles of the disclosure. The following description will describe various exemplary aspects of the disclosure with reference to the following drawings.

[0006] [Figure 1] This document illustrates various aspects of autonomous vehicles relating to this disclosure. [Figure 2] This document illustrates various exemplary electronic components of a vehicle safety system relating to various aspects of this disclosure. [Figure 3] This shows an example vehicle equipped with multiple sensors. [Figure 4] The vehicle interior 400 according to one aspect of this disclosure is shown. [Figure 5] This document presents an object labeling algorithm based on human speech. [Figure 6] This shows an example of gaze being used to identify an object. [Figure 7] This document shows a gaze-tracking detector relating to one aspect of the present disclosure. [Figure 8] This document shows a calculation of mirror gaze relating to a certain aspect of this disclosure. [Figure 9] This shows a hand gesture detector that can be configured to detect one or more hand gestures or hand positions. [Figure 10] This disclosure shows a data synthesis apparatus and labeler relating to certain aspects of this disclosure. [Figure 11] This document describes a data storage device relating to one aspect of this disclosure. [Figure 12] This document shows a vehicle data association device relating to one aspect of this disclosure. [Figure 13] This document describes how to associate vehicle data. [Modes for carrying out the invention]

[0007] The following detailed description will refer to the accompanying drawings illustrating exemplary details and embodiments of the embodiments of this disclosure.

[0008] In this specification, the term “exemplary” is used to mean “serving as an example, illustration, or reference.” Any aspect or design described herein as “exemplary” should not necessarily be construed as being preferable or advantageous to any other aspect or design.

[0009] Unless otherwise specified, the same reference numerals are used throughout the drawings to indicate the same or similar elements, features, and structures.

[0010] The phrases "at least one" and "one or more" may be understood to include one or more quantities (e.g., one, two, three, four, [...], etc.). In this specification, the phrase "at least one of" relating to a group of elements may be used to mean at least one element from the group consisting of those elements. For example, in this specification, the phrase "at least one of" relating to a group of elements may be used to mean selecting one of the enumerated elements, one of several enumerated elements, several individual enumerated elements, or several individual enumerated elements.

[0011] The terms “plural” and “multiple” in the specification and claims clearly refer to more than one quantity. Therefore, any phrase explicitly suggesting the aforementioned terms referring to a quantity of elements (e.g., “plural, multiple [elements]”) clearly refers to more than one of those elements. For example, the phrase “plural” may be understood to include two or more quantities (e.g., two, three, four, five, [...], etc.).

[0012] In the specification and claims, phrases such as “a group,” “a set,” “a collection,” “a sequence,” “a series,” and “a bundle” refer to one or more quantities, if any. The terms “proper subset,” “reduced subset,” and “smaller subset” refer to a subset of a set that is not equal to a set, and by example, a subset of a set that contains fewer elements than a set.

[0013] As used herein, the term “data” may be understood to include any suitable analog or digital information provided, for example, as a file, part of a file, a set of files, a signal or stream, part of a signal or stream, and a set of signals or streams. Furthermore, the term “data” may be used to mean, for example, a reference to information in pointer form. However, the term “data” is not limited to the examples given herein and may represent any information in various forms as understood in the art.

[0014] For example, the terms “processor” or “controller” as used herein may be understood as any kind of technical entity that enables the processing of data. The data may be processed according to one or more specific functions performed by the processor or controller. Furthermore, as used herein, the processor or controller may be understood as any kind of circuit, for example, any kind of analog or digital circuit. Thus, the processor or controller may be an analog circuit, a digital circuit, a mixed-signal circuit, a logic circuit, a processor, a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field-programmable gate array (FPGA), an integrated circuit, an application-specific integrated circuit (ASIC), or any combination thereof, and may include these. Any other kind of implementation of each of the functions described in more detail below may be understood as a processor, a controller, or a logic circuit. It will be understood that any two (or more) of the processors, controllers, or logic circuits detailed herein may be realized as a single entity having equivalent functions, etc., and conversely, any single processor, controller, or logic circuit detailed herein may be realized as two (or more) separate entities having equivalent functions, etc.

[0015] As used herein, “memory” is understood to mean a computer-readable medium (e.g., non-temporary computer-readable medium) capable of storing data or information for retrieval. Therefore, references to “memory” included herein should be understood to mean volatile or non-volatile memory, including, among other things, random-access memory (RAM), read-only memory (ROM), flash memory, solid-state memory, magnetic tape, hard disk drives, optical drives, 3D XPoint®, or any combination thereof. In this specification, the term “memory” also includes, among other things, registers, shift registers, processor registers, and data buffers. The term “software” refers to any type of executable instruction, including firmware.

[0016] Unless otherwise explicitly specified, the term “transmit” encompasses both direct (point-to-point) transmission and indirect transmission (through one or more intermediate points). Similarly, the term “receive” encompasses both direct and indirect reception. Furthermore, the terms “transmit,” “receive,” “communicate,” and other similar terms encompass both physical transmission (e.g., transmission of radio signals) and logical transmission (e.g., transmission of digital data through logical software-level connections). For example, a processor or controller may transmit or receive data in the form of radio signals through a software-level connection with another processor or controller, where physical transmission and reception are handled by radio layer components such as RF transceivers and antennas, and logical transmission and reception through software-level connections are performed by the processor or controller. The term “communicate” encompasses one-way or two-way communication in either or both the receiving and receiving directions, i.e., in either or both the receiving and sending directions. The term “calculate” encompasses both “direct” calculations using mathematical formulas / relationships and “indirect” calculations through lookup tables or hash tables and other array indexing or array lookup operations.

[0017] "Vehicle" may be understood to include any type of driven object. For example, a vehicle may be a driven object having a combustion engine, a reaction engine, an electric drive, a hybrid drive, or a combination thereof. A vehicle may, among other things, be an automobile, bus, minibus, van, truck, mobile home, vehicle trailer, motorcycle, bicycle, tricycle, locomotive, freight train, mobile robot, personal mobility, boat, ship, submarine, submersible, drone, aircraft, or rocket.

[0018] The term "self-driving vehicle" may refer to a vehicle that can implement at least one navigation change without driver input. The navigation change may represent or include one or more changes in vehicle steering, braking, or acceleration / deceleration. Even if the vehicle is not fully autonomous (e.g., fully operational with or without a driver present), the vehicle may be described as self-driving. Self-driving vehicles may include vehicles that can operate under driver control for a specific period and can operate without driver control during other periods. Self-driving vehicles may include vehicles that control only some aspects of vehicle navigation, such as steering (e.g., to maintain the vehicle path between lane limits) or some steering operations in certain situations (but not all situations), while other aspects of vehicle navigation (e.g., braking, or braking in certain situations) may be left to the driver. Self-driving vehicles may include vehicles that share control of one or more aspects of vehicle navigation in certain situations (e.g., hands-on, such as in response to driver input) and vehicles that control one or more aspects of vehicle navigation in certain situations (e.g., hands-off, such as not depending on driver input). Self-driving vehicles may include vehicles that control one or more aspects of vehicle navigation under certain situations, such as under certain environmental conditions (e.g., spatial area, lane conditions). In some embodiments, self-driving vehicles may handle some or all aspects of vehicle braking, speed control, velocity control, and / or steering. Self-driving vehicles may include vehicles that can operate without a driver present. The level of self-driving of a vehicle may be described or determined by the Society of Automotive Engineers (SAE) level of the vehicle (e.g., as defined by SAE, e.g., in SAE J3016 2018: Taxonomy and Definitions for Terms Related to On-Road Motor Vehicle Automated Driving Systems) or by other relevant professional organizations. The SAE level may have values ranging from a minimum level, e.g., level 0 (as an example, substantially no driving automation) to a maximum level, e.g., level 5 (as an example, full driving automation).

[0019] In the context of the present disclosure, "vehicle operation data" may be understood to represent any type of feature related to the operation of a vehicle. By way of example, "vehicle operation data" may represent the state of a vehicle, such as the type of tires of the vehicle, the type of vehicle, and / or the time of manufacture of the vehicle. More generally, "vehicle operation data" may represent or include static features or static vehicle operation data (as an example, features or data that do not change over time). As another example, additionally or alternatively, "vehicle operation data" may represent or include features that change during the operation of the vehicle, such as environmental conditions, such as weather conditions or road conditions during the operation of the vehicle, fuel level, liquid level, operating parameters of the vehicle's drive source, and the like. More generally, "vehicle operation data" may represent or include changing features or changing vehicle operation data (as an example, features or data that change over time).

[0020] Various aspects of the disclosure herein may utilize one or more machine learning models to perform or control the functions of a vehicle (or other functions described herein). For example, the term "model" as used herein may be understood as any type of algorithm that provides output data from input data (e.g., any type of algorithm that generates or calculates output data from input data). A machine learning model may be executed for a computing system to gradually improve the execution of a particular task. In some aspects, the parameters of a machine learning model may be adjusted in a training phase based on training data. A trained machine learning model may be used in an inference phase to make predictions or decisions based on input data. In some aspects, a trained machine learning model may be used to generate additional training data. A further machine learning model may be adjusted in a second training phase based on the generated additional training data. A trained further machine learning model may be used in an inference phase to make predictions or decisions based on input data.

[0021] The machine learning models described herein may take any appropriate form and utilize any appropriate method (for example, for training purposes). For example, any machine learning model may utilize supervised learning, semi-supervised learning, unsupervised learning, or reinforcement learning methods.

[0022] In supervised learning, a model may be built using a training set of data that includes both inputs and corresponding desired outputs (for example, each input may be associated with a desired or expected output for that input). Each training instance may include one or more inputs and desired outputs. Training may involve repeating training instances and teaching the model to predict outputs for new inputs (for example, inputs not included in the training set) using an objective function. In semi-supervised learning, some inputs in the training set may lack their respective desired outputs (for example, one or more inputs may not be associated with any desired or expected output).

[0023] In unsupervised learning, a model may be constructed from a training set of data that contains only the inputs but not the desired outputs. Unsupervised models may be used, for example, to find structures in data by discovering patterns in the data (e.g., grouping or clustering data points). Techniques that may be implemented in unsupervised learning models may include, for example, self-organizing maps, nearest neighbor mapping, K-means clustering, and singular value decomposition.

[0024] Reinforcement learning models may include positive or negative feedback to improve accuracy. Reinforcement learning models may attempt to maximize one or more objectives / rewards. Techniques that may be implemented in reinforcement learning models may include, for example, Q-learning, temporal difference (TD), and deep adversarial networks.

[0025] Various embodiments described herein may utilize one or more classification models. In some classification models, the output may be limited to a limited set of values ​​(e.g., one or more classes). A classification model may output a class for an input set of one or more input values. The input set may include sensor data such as image data, radar data, and lidar data. Classification models described herein may classify specific driving and / or environmental conditions, such as weather conditions and road conditions. References to classification models herein may refer to models that implement any one or more of the following techniques, for example, linear classifiers (e.g., logistic regression or naive Bayesian classifiers), support vector machines, decision trees, boosting trees, random forests, neural networks, or nearest neighbors.

[0026] Various embodiments described herein may utilize one or more regression models. A regression model may output a number from a contiguous range based on (for example, starting with or using a set of one or more input values). References to regression models herein may refer to models that implement any one or more of the following techniques (or other suitable techniques), such as linear regression, decision trees, random forests, or neural networks.

[0027] The machine learning models described herein may be or may include an ANN. The ANN may be any type of neural network, such as a convolutional neural network, an autoencoder network, a variational autoencoder network, a sparse autoencoder network, a recurrent neural network, a deconvolutional network, a generative adversarial network, an advanced neural network, and a sum-product neural network. The ANN may contain any number of layers. Training of the ANN (e.g., fitting the layers of the neural network) may use or be based on any type of training principle, such as backpropagation (e.g., using a backpropagation algorithm).

[0028] Throughout this disclosure, terms such as driving parameter set, driving model parameter set, safety layer parameter set, driver assistance, autonomous driving model parameter set, and / or similar (e.g., driving safety parameter set) are used as synonyms.

[0029] Furthermore, throughout this disclosure, terms such as driving parameters, driving model parameters, safety layer parameters, driver assistance, and / or automated driving model parameters, and / or similar (e.g., driving safety parameters) are used as synonyms.

[0030] Figure 1 shows an exemplary vehicle according to various embodiments of the present disclosure, namely vehicle 100. In some embodiments, vehicle 100 may include one or more processors 102, one or more image acquisition devices 104, one or more position sensors 106, one or more speed sensors 108, one or more radar sensors 110, and / or one or more lidar sensors 112.

[0031] In some embodiments, the vehicle 100 may include a safety system 200 (as described below with respect to Figure 2). Since the vehicle 100 and safety system 200 are inherently illustrative, they may be simplified for illustrative purposes. The locations and correlated distances of the elements (as mentioned above, the figures are not to scale) are provided as examples and are not limited thereto. The safety system 200 may include various components depending on the requirements of a particular implementation.

[0032] Figure 2 shows various exemplary electronic components of a vehicle according to various embodiments of the present disclosure, i.e., a safety system 200. In some embodiments, the safety system 200 may include one or more processors 102, one or more image acquisition devices 104 (e.g., one or more cameras), one or more position sensors 106 (e.g., in particular, a Global Navigation Satellite System (GNSS), a Global Positioning System (GPS)), one or more speed sensors 108, one or more radar sensors 110, and / or one or more lidar sensors 112. According to at least one embodiment, the safety system 200 may further include one or more memories 202, one or more map databases 204, one or more user interfaces 206 (e.g., a display, a touchscreen, a microphone, a loudspeaker, one or more buttons and / or switches, etc.), and / or one or more radio transceivers 208, 210, 212. In some embodiments, the radio transceivers 208, 210, and 212 may be configured according to the same, different, or any combination thereof, radio communication protocols or standards. For example, a radio transceiver (e.g., the first radio transceiver 208) may be configured according to a short-range mobile radio communication standard (e.g., Bluetooth®, Zigbee®, among others). In another example, a radio transceiver (e.g., the second radio transceiver 210) may be configured according to a medium-range or wide-area mobile radio communication standard (e.g., among others, 3G (e.g., Universal Mobile Telecommunications System), 4G (e.g., Long Term Evolution, and / or 5G mobile radio communication standards) compliant with the corresponding 3rd Generation Partnership Project (3GPP) standards.As a further example, a wireless transceiver (e.g., a third wireless transceiver 212) may be configured according to a wireless local area network communication protocol or standard (e.g., IEEE 802.11, 802.11a, 802.11b, 802.11g, 802.11n, 802.11p, 802.11-12, 802.11ac, 802.11ad, 802.11ah). One or more wireless transceivers 208, 210, 212 may be configured to transmit signals via an antenna system through an air interface.

[0033] In some embodiments, one or more processors 102 may include an application processor 214, an image processor 216, a communication processor 218, and / or any other suitable processing devices. The image acquisition device 104 may include any number of image acquisition devices and components depending on the requirements of the particular application. The image acquisition device 104 may include one or more image acquisition devices (e.g., a camera, a charge-coupled device (CCD), or any other type of image sensor).

[0034] In at least one embodiment, the safety system 200 may include a data interface that communicates one or more processors 102 to one or more image acquisition devices 104. For example, the first data interface may include an optional wired and / or wireless first link 220 configured to transmit image data acquired by one or more image acquisition devices 104 to one or more processors 102 (e.g., an image processor 216).

[0035] In some embodiments, the wireless transceivers 208, 210, 212 may be coupled to one or more processors 102 (e.g., a communication processor 218) via, for example, a second data interface. The second data interface may include an optional wired and / or wireless second link 222 configured to transmit wireless transmission data acquired by the wireless transceivers 208, 210, 212 to one or more processors 102, e.g., a communication processor 218.

[0036] In some embodiments, the memory 202 and one or more user interfaces 206 may be coupled to each of one or more processors 102, for example, via a third data interface. The third data interface may include an optional wired and / or wireless third link 224. Furthermore, the position sensor 106 may be coupled to each of one or more processors 102, for example, via the third data interface.

[0037] Such transmission may include (e.g., one-way or two-way) communication between vehicle 100 and one or more other (target) vehicles in the environment of vehicle 100 (e.g., to facilitate the adjustment of vehicle 100's navigation in consideration of, or in conjunction with, other (target) vehicles in the environment of vehicle 100), or further, broadcast transmission to unspecified recipients in the vicinity of the transmitting vehicle 100.

[0038] One or more of the transceivers 208, 210, and 212 may be configured to implement one or more vehicle-to-vehicle / vehicle-to-infrastructure (V2X) communication protocols. V2X communication protocols may include vehicle-to-vehicle (V2V), vehicle-to-infrastructure (V2I), vehicle-to-network (V2N), vehicle-to-pedestrian (V2P), vehicle-to-device (V2D), vehicle-to-grid (V2G), and other protocols.

[0039] Each of the one or more processors 102, each of the processors 214, 216, 218, may include various types of hardware-based processing devices. For example, each of the processors 214, 216, 218 may include a microprocessor, a preprocessor (e.g., an image preprocessor), a graphics processor, a central processing unit (CPU), support circuitry, a digital signal processor, an integrated circuit, memory, or any other type of device suitable for application execution and image processing and analysis. In some embodiments, each of the processors 214, 216, 218 may include any type of single-core or multi-core processor, a mobile device microcontroller, a central processing unit, etc. Each of these processor types may include multiple processing units having local memory and an instruction set. Such processors may include a video input device for receiving image data from multiple image sensors and may also include a video output function.

[0040] The processors 214, 216, and 218 disclosed herein may be configured to perform specific functions according to program instructions that can be stored in one of the one or more memories 202. In other words, one of the one or more memories 202 may store software that, when executed by a processor (e.g., one or more processors 102), controls the operation of a system, such as a safety system. One of the one or more memories 202 may store, for example, one or more databases, and image processing software, as well as trained systems such as neural networks or deep neural networks. One or more memories 202 may include any number of random access memories, read-only memories, flash memories, disk drives, optical storage devices, tape storage devices, removable storage devices, and other types of storage devices.

[0041] In some embodiments, the safety system 200 may further include components such as a speed sensor 108 (e.g., a speedometer) for measuring the speed of the vehicle 100. The safety system may also include one or more accelerometers (single-axis or multi-axis) (not shown) for measuring the acceleration of the vehicle 100 along one or more axes. The safety system 200 may further include additional sensors or different sensor types, such as an ultrasonic sensor (which may be integrated into the headlights of the vehicle 100), a thermal sensor, one or more radar sensors 110, and one or more lidar sensors 112. The radar sensors 110 and / or lidar sensors 112 may be configured to provide pre-processed sensor data, such as a radar target list or a lidar target list. A third data interface may connect the speed sensor 108, one or more radar sensors 110, and one or more lidar sensors 112 to at least one of one or more processors 102.

[0042] One or more memories 202 may store, for example, data indicating the locations of known landmarks, for example, in a database or in any different format. One or more processors 102 may process sensory information of the vehicle 100's environment (e.g., images, radar signals, or depth information from lidar or stereo processing of two or more images) together with positional information such as GPS coordinates, vehicle ego-motion, etc. to determine the vehicle 100's current location relative to known landmarks and to refine the determination of the vehicle's location. Certain aspects of this technique may be included in localization techniques such as mapping and routing models.

[0043] The map database 204 may include any type of database for storing (digital) map data for the vehicle 100, for example, the safety system 200. The map database 204 may include data relating to the locations of various items in a reference coordinate system, including roads, buildings using water, geographical features, businesses, points of interest, restaurants, gas stations, etc. The map database 204 may store not only the locations of such items but also descriptors associated with these items, including, for example, names associated with any of the stored features. In this embodiment, one of the one or more processors 102 may download information from the map database 204 via a wired or wireless data connection to a communication network (for example, via a cellular network and / or the Internet, etc.). In some cases, the map database 204 may store a sparse data model including a polynomial representation of specific road features (e.g., lane markings) or a target trajectory for the vehicle 100. The map database 204 may include stored representations of various recognized landmarks that can be provided to determine or update the known position of the vehicle 100 relative to the target trajectory. The marker representation may include data fields such as the marker type and marker location, among many other potential identifiers.

[0044] Furthermore, the safety system 200 may include, for example, a driving model implemented in an advanced driver-assistance system (ADAS) and / or a driver assistance and automated driving system. For example, the safety system 200 may include a computer implementation of a formal model, such as a safe driving model (for example, as part of a driving model). The safe driving model may be, or include, a mathematical model that formalizes interpretations of applicable laws, standards, policies, etc., that can be applied to an automated (ground) vehicle. The safe driving model may be designed to achieve, for example, the following three objectives: First, the interpretation of the law must be sound in the sense that it conforms to the way humans interpret the law. Second, the interpretation must result in useful driving policies, that is, sensible driving policies rather than overly defensive driving that inevitably confuses other human drivers, disrupts traffic, and consequently limits the scalability of system deployment. Third, the interpretation must be efficiently verifiable in the sense that it can be rigorously proven that an automated (autonomous) vehicle correctly implements the interpretation of the law. A safe driving model may, for example, be a mathematical safety model that enables the identification and execution of appropriate responses to dangerous situations in which self-inflicted accidents can be avoided, or may include such a mathematical model.

[0045] As described above, the vehicle 100 may include a safety system 200, similarly described with reference to Figure 2.

[0046] The vehicle 100 may include, for example, one or more processors 102 that are integrated with or separate from the engine control unit (ECU) of the vehicle 100.

[0047] The safety system 200 may generally directly or indirectly control the operation of the vehicle 100 by generating data to control the vehicle's ECU and / or other components, or to assist in such control.

[0048] One of the challenges in implementing an ANN for autonomous driving is training the ANN to distinguish between relevant sensor data and irrelevant (e.g., irrelevant, low direct relevance) sensor data. This can be analogous to an inexperienced driver. An inexperienced driver may have developed excellent vision and hearing and may be able to accurately perceive the surroundings of the vehicle, but they may struggle to distinguish highly important information from less important information. For example, an inexperienced driver may overemphasize the importance of an ambulance on the other side of a divided highway, or fail to understand the importance of a small child on a bicycle just a few meters from the road. As an inexperienced driver develops their driving skills, they learn to identify hazards in sensory information and to assign appropriate weights (e.g., important or unimportant, relevant or irrelevant) to individual aspects of sensory information.

[0049] Allan Networks (ANNs) must be similarly trained to distinguish between relevant and irrelevant information. That is, an ANN must receive sensor information (e.g., individual sensor data streams or multiple sensor data streams) and be trained to identify aspects or portions of sensor data that are particularly relevant for driving decisions. Conversely, an ANN may be configured to identify less relevant or irrelevant sensor data.

[0050] As described above, training an ANN may modify one or more weights associated with its nodes or layers, and / or one or more functions (e.g., one or more activation functions). Since there are various types and implementations of ANNs that can affect the details of ANN training, assuming that those skilled in the art are expected to understand the details of ANN training at the node and / or function level, this disclosure primarily describes higher-level training functions in which one or more streams of sensor data are analyzed for relevance and simultaneously cross-referenced. According to one aspect of this disclosure, an ANN may use these techniques to identify and label relevant sensor data for processing driving decisions. According to another aspect of this disclosure, an ANN may use these techniques to distinguish relevant information from irrelevant information.

[0051] In known efforts to train ANNs (Autonomous Navigation Engines) for purposes such as evaluating sensor data in the context of autonomous driving, vendors may collect hundreds or even thousands of hours of sensor data that demonstrate human driver behavior in relation to corresponding recording data around the vehicle, including road boundaries and drivable surfaces, traffic signs, static objects, movable objects, and location data. The ANN can use this data to learn to match specific external sensor inputs with specific driver behavior. Such conventional methods make it difficult to teach the ANN specific hazards in specific situations because there are often many noticeable driver responses or because accidents do not actually occur. In other words, the ANN's ability to learn from these situations may be limited because experienced human drivers (e.g., model drivers in training programs) can avoid or prevent hazards through normal operation without any noticeable signals of avoidable hazards or actions taken to avoid them. In contrast, human instructors can leverage their experience to recognize potential hazards and provide guidance. In some cases, an instructor may comment on or instruct about potential hazards even before a collision is imminent or even before a collision may occur. The ability to perceive driving instructions represents an efficient strategy for augmenting or replacing existing methods of training ANNs for driving operations.

[0052] This disclosure describes, in particular, strategies for incorporating input or instructions from a human driving instructor and learning this information by cross-referencing it with one or more additional sources, such as the behavior of an inexperienced driver in the vehicle and / or additional sensor information, so that the ANN can learn to label dangerous driving situations and / or map these dangerous situations to examples of correct and incorrect human driver behavior. In this way, human driving behavior may be recorded together with a sensor-based environmental model, and potentially dangerous driving situations may be pointed out, and correct responses to perceived dangers may be recorded. Since potential dangers may not materialize as actual dangers during vehicle operation, conventional approaches may fail to recognize their importance without explicit commentary on such potential dangers.

[0053] Driving instructors are accustomed to recognizing and explaining dangerous situations to inexperienced drivers (e.g., their students). Inexperienced drivers will exhibit both correct and incorrect behaviors through driving instructions. The primary task of the driving instructor is to comment on or "label" these behaviors. This interaction between instructor and student can be collected to improve automated vehicle assessment of dangerous situations (e.g., to train ANNs to better evaluate sensor data).

[0054] As mentioned above regarding the similarities with inexperienced drivers, driver training programs (e.g., classes, instructional sessions, hands-on sessions) are often conducted in the context of instructors and students being together in a vehicle or vehicle simulator, or in another vehicle where a passenger provides instructions to the driver. Throughout this disclosure, the term “vehicle” is used in the context of driver training, but it is made clear that the term “vehicle” may also refer to a vehicle simulator. Typically, the student operates the vehicle (e.g., sitting in the driver’s seat), and the instructor provides driving inputs, such as verbal instructions, physical cues (e.g., gestures, pointing, other body language), or, in some cases, further, physical operation of the steering wheel or brakes. These driving inputs may be used for ANN training. In other words, these driving inputs may be analyzed and cross-referenced with sensor data to distinguish between highly relevant and less relevant sensor data.

[0055] Autonomous vehicles, semi-autonomous vehicles (e.g., vehicles that perform one or more autonomous driving operations but cannot operate continuously and completely independently of human control), and even primarily non-autonomous vehicles are often equipped with multiple sensors capable of detecting information about areas inside or outside the vehicle (e.g., inside the vehicle, near the outside of the vehicle, etc.) and generating corresponding sensor data. One or more sensors may be connected to one or more processors (e.g., electrically connected) or configured to wirelessly transmit sensor data to one or more processors (e.g., via one or more transmitters and receivers). Figure 3 shows an exemplary vehicle with multiple sensors. In this figure, the vehicle includes multiple external sensors (e.g., outward-facing sensors) configured to detect the environment around the vehicle, including, but not limited to, drivable paths, lane markings, static movable objects (e.g., other vehicles), and traffic signs. The sensors may be positioned to cover 360° around the vehicle so that both forward and reverse operations can be detected. The sensor may include one or more image sensors (e.g., cameras, video cameras, depth cameras, etc.) 302a and one or more distance sensors (e.g., light detection and ranging (LIDAR), radio wave detection and ranging (RADAR)) 302b. The external sensor may optionally include one or more microphones (not shown) that can be configured to detect external and / or ambient noise.

[0056] The vehicle may be equipped with one or more inward-facing visual sensors 304a that can be configured to detect information inside the vehicle. These may include one or more image sensors (e.g., mono or stereo cameras) and / or one or more novel sensors, such as event cameras that can detect gaze and / or gestures from a driving instructor and / or the driver. The vehicle may be equipped with one or more inward-facing microphones 304b that can be configured to record audio from inside the vehicle by a driving instructor and / or the driver. The vehicle may be equipped with vehicle operation sensors 306 that can be configured to detect speed, acceleration, steering, and / or braking. The vehicle may include one or more position sensors 308 (e.g., one or more Global Navigation Satellite System (GNSS) sensors) that can be configured to detect position, location, and / or direction of travel data. The vehicle may include a data storage system 308 (e.g., memory, hard drive, solid-state drive, optical drive, etc.) in which some or all of the above sensor data may be stored (e.g., over the duration of driving instruction, over a predetermined duration, indefinitely, etc.). The vehicle may include one or more timekeepers (e.g., clock, processor clock, GNSS time management device, etc.) (not shown) so that the device may create synchronized timestamps across all sensor data streams. Such synchronization may be useful or even necessary for subsequent comparison of the data streams. The vehicle may include a high-bandwidth data connection that the vehicle can use to upload collected sensor data to an external device (e.g., server, data repository, etc.). Alternatively or additionally, the vehicle may include one or more data transfer devices (not shown) which may include one or more buses, one or more ports, one or more data transfer cables, etc. Such uploading of sensor data may occur continuously or after the training has been completed.

[0057] The vehicle may include one or more processors and one or more non-temporary computer-readable media containing instructions. When executed, the instructions cause one or more processors to interpret the gaze, gestures, or voice of the driver and / or driving instructor and to create labels that can associate dangerous situations with external sensor data. These instructions may include one or more machine learning algorithms for interpreting external and internal scenes.

[0058] In some cases, data relating to the student (inexperienced driver) may be useful, but data relating to the driving instructor's gaze, gestures, voice, or any of these may be more useful or consistently useful than data relating to the student. In some cases, student response and related steering, acceleration, and / or braking data may be useful for data referencing and / or label generation as described herein. If the vehicle includes one or more processors and non-temporary computer-readable media as described herein, the sensor data referencing and label generation functions as described herein may be performed within the vehicle. Alternatively or additionally, these functions may be performed outside the vehicle (e.g., on a device to which relevant data is uploaded, whether simultaneously with / in parallel with or after driving instructions). For example, a central memory processing facility may exist for all vehicles of a participating driving school. In this case, all relevant / eligible vehicles may upload their data to the central memory processing facility.

[0059] Sensor data retrieval and label generation may be performed in real time with vehicle operation (e.g., online) or after vehicle operation (e.g., offline). Whether online or offline, this configuration may be chosen based on the desired implementation. Since sensor data retrieval and label generation are performed from sensor data that can be fully recorded, no degradation in analytical quality is expected from an offline configuration.

[0060] The device may label external sensor data based on input from inside the vehicle. Figure 4 shows an interior (e.g., a stylized diagram of the interior through a windshield) 400 according to one aspect of the present disclosure. In this diagram, the windshield is shown to include an instructor section 402 (e.g., generally corresponding to a passenger seat) and a driver section 410 (generally corresponding to a driver's seat). As described above, the interior may include one or more microphones 404 that can be configured to detect and record voice or human sounds (e.g., non-verbal utterances) inside the vehicle. The vehicle may include one or more cameras (not shown) that can be configured to detect information and generate corresponding image data of the interior of the vehicle. Any aspect of the interior of the vehicle can be utilized in the principles and methods disclosed herein, but one or more processors may be configured to detect specific markers in the image data that can provide relevant information. These markers may include the gaze of instructor 406, the facial expression of instructor 407, the gesture of instructor 408 (e.g., a hand gesture), the gaze of the inexperienced driver 412, or any of these.

[0061] According to one aspect of this disclosure, one or more processors may label external sensor data based on one or more comments from a driving instructor. Human speech can convey varying degrees of information content. Therefore, in order to process and understand human speech relevant to sensor data, it may be useful, firstly, to analyze the human speech for the type of content it contains. Such types of content may include, for example, the location of a hazard (e.g., "front right"), the identification of objects involved in the hazard (e.g., "a group of elderly women"), statements of possible outcomes (e.g., "they might step into the road without looking"), statements of correct actions (e.g., "look," "get ready to brake," "slow down," "turn the steering wheel to the right," "honk the horn"), or any of these.

[0062] According to certain aspects of this disclosure, one or more processors may be configured to label the irrelevance of objects based on at least human speech. To use human speech (for example, to identify danger or relevance based on human speech), one or more processors may be configured to recognize patterns or keywords in human speech. The following describes keyword analysis for labeling relevant objects.

[0063] One or more processors may be configured to use voice data as a primary mode to trigger the attribution of increasing relevance to sensor data. One or more processors may use object delivery keywords to classify specific objects or groups of objects in relation to human voice. One or more processors may be configured to apply object detection algorithms in parallel with external sensor inputs (e.g., using object detection network analysis). For example, if the voice recognition class is bicycles, all bicycles in the image sensor data are marked (e.g., labeled) as hazardous or potentially hazardous. If several objects of the same type are detected (e.g., if there is more than one bicycle in the image sensor data), one or more processors may be configured to rank the multiple objects by relevance. Relevance may be determined based on any desirable factor. According to one aspect of this disclosure, one or more processors may be configured to rank multiple objects corresponding to the same keyword in terms of relevance based on distance from their own vehicle.

[0064] The received human voice can be assumed to contain multiple keywords, each potentially providing information about the location or relevance of a danger or object. While the keyword structure can be implemented in various ways, an exemplary keyword implementation is described below. In this implementation, keywords may be defined by category as follows: a) Warning Keywords: Warning keywords may include any keyword that suggests an increased need for vigilance or the possibility of imminent danger. Such keywords may include, but are not limited to, “caution,” “be careful,” “cautious,” “look,” or “be vigilant.” Alternatively or additionally, such warning keywords may correspond to utterances that are associated with the need for vigilance or the possibility of imminent danger, such as “oh,” “wow,” or “um,” even though these may not directly correspond to the original words. b) Directional keywords: Directional keywords may include any keyword that indicates direction, such as direction to the speaker, direction to the driver, direction to the vehicle, or other directions. Such directional keywords may include, but are not limited to, “forward,” “left,” “right,” “up,” “down,” “backward,” or others. c) Qualitative location keywords: Qualitative location keywords may include any keyword that suggests a location relative to a reference point. The reference point may be anything without limitation, including but not limited to, people, roadways, buildings, vehicles, animals, and other objects. Examples of qualitative location keywords may include, but not limited to, "at the next intersection," "behind the bus stop," "on the sidewalk," "next to the white car," "above the roadway," "below the bridge," "in front of the dumpster," "next to the sign," or others. d) Object description keywords: Object description keywords may include any keywords that describe one or more objects. In this case, objects may include living things, inanimate things, animals, people, or any combination thereof. Object description keywords may be used in relation to objects that pose a risk to the driver or vehicle. Object description keywords may include, but are not limited to, "a group of children," "two bicycles," "trailer," "elderly woman," "cyclist," "jogger," etc. e) Action Keywords: Action keywords may include any keywords that suggest instructions regarding actions to be taken by the driver or vehicle. In the context of driving instructions, a driving instructor may need to provide instructions in the form of commands in relation to perceived hazards. Action keywords may be closely associated with commands so that the presence of one or more action keywords may enable the underlying system to identify the instruction and associate it with a driving command. Examples of action keywords include, but are not limited to, “slowly,” “faster,” “stop,” “pull over to the side,” “cross quickly,” “overtake quickly,” “change direction,” “turn the steering wheel,” “apply the brakes,” and “signal.”

[0065] Keyword recognition in human speech requires at least functional level speech recognition. That is, one or more processors must analyze microphone data representing human speech, and from this human speech, one or more processors must identify words, phrases, sentences, or any of these. Computerized speech recognition is known and will not be described in detail herein. Rather, it is assumed that those skilled in the art will understand how to implement one or more speech recognition programs for detecting human speech in audio data. Such speech recognition programs are typically configured to receive an audio file and output text corresponding to the recognized human speech. In the above description relating to keywords, it is assumed that such recognized material can be used for keyword recognition in text or in any other suitable format.

[0066] In one aspect of this disclosure, the device may be programmed to perform an inference phase in which human interaction is detected. In the inference phase, voice recognition may be the primary mode in a machine learning toolchain. Upon recognition of a warning keyword, a new warning label may be initiated (e.g., generated, implemented, executed). This keyword may then be associated with external sensor data. External sensor data occurring simultaneously with the warning keyword may be considered most relevant. Alternatively or additionally, external sensor data slightly preceding the warning keyword may be particularly relevant. This is best explained by the fact that the driving instructor's cognitive process requires time to perceive external conditions of the vehicle (e.g., objects and hazards, etc.) and then to formulate and express verbal statements or responses in relation to the perceived conditions. Thus, any spoken words may be attributable to sensor data occurring immediately before the words are spoken. Alternatively, due to the relative persistence of objects and specific hazardous situations, a hazard instantaneously perceived before a keyword about a hazard is likely to still exist when the keyword is spoken. Therefore, one or more processors may be configured to attribute relevance to portions of sensor data that occur simultaneously with spoken keywords.

[0067] Depending on the richness of information contained in the instructor's instructions, it may be possible to map the entire scene recorded by the sensor to the instructions, or rather, to map only specific areas or specific objects to the instructions for labeling purposes.

[0068] In other words, after identifying objects related to a keyword, one or more processors may be configured to further narrow the focus area within external sensor data (e.g., image sensor data, radar, lidar, etc.) by applying qualitative keyword locations. More specifically, a speaker may utter a qualitative keyword location corresponding to a particular object among a group of objects. More specifically, continuing with the hypothetical bicycle example, the instructor may say, "Watch out! Bicycle. Near the sign." Thus, "Watch out" is a attention keyword, which causes one or more processors to label the incoming sensor data as highly relevant for a given duration. The keyword "bicycle" represents an object in the sensor data, and one or more processors may be configured to find one or more bicycles in the sensor data. For example, a forward-facing camera may supply sensor data representing the vicinity of a vehicle where three bicycles are present. One or more processors may be configured to label each of the three bicycles as particularly relevant. Next, the keyword "near the sign" represents the location of the most relevant bicycle. In this way, one or more processors may be configured to locate a sign and determine the proximity of that sign to various bicycles. The bicycle closest to the sign may be identified as a bicycle by the sign that the rider should pay attention to. One or more processors may be configured to label this bicycle as being of higher importance than other identified bicycles.

[0069] Figure 5 illustrates an object labeling algorithm based on human speech. In this figure, four road scenes and three children are shown as 502, 504, 506, and 508. These scenes are shown chronologically, illustrating the progression of object identification based on human speech. In the first scene, 502, the instructor uses the alert keyword "attention." As described above, this alert keyword may indicate that one or more objects in the field of view necessitate further caution. Since no identifying information regarding the source of the hazard is provided, one or more processors may label the entire field of view as particularly relevant. This is illustrated in the fully shaded figure of 502. That is, one or more processors may label all image data occurring simultaneously with, or essentially simultaneously with, the alert keyword as highly relevant. One or more processors may be configured to derive further keywords from subsequent human speech to further identify hazards.

[0070] In the following scene 504, the instructor supplements with the object description keyword “child.” One or more processors may be configured to recognize this keyword and search for image data that has already been marked as highly relevant to objects associated with the keyword “child.” One or more processors may utilize any known object detection algorithm for this purpose. When one or more processors detect one or more “children” in the image sensor data, they may be configured to restrict the highly relevant areas to areas that generally correspond to “children.” In this way, one or more processors can label smaller areas as highly relevant, thereby further distinguishing between relevant and irrelevant materials.

[0071] In the following scene 506, the instructor uses the qualitative location keyword "on the road." Assuming this qualitative location keyword is spoken in a close temporal relationship with the alert keyword, one or more processors may be configured to further narrow down the highly relevant area to the area where the child is on the road. Needless to say, this requires the further context of "road," the determination of where the road is located, and the context of "on." One or more processors may use the phrasing of the contexts "on" and "on the road" to search for one of the identified children currently on the road. In this case, there is currently only one child on the road, and therefore, one or more processors may restrict the highly relevant area to the area surrounding that single child on the road.

[0072] In the following scene 508, further certainty may be provided regarding the determined highly relevant areas. The decisions made in scenes 502, 504, and 506 may be made in relation to confidence levels; that is, one or more additional verbal keywords or gestures may be used to increase or decrease the confidence level regarding the locations shown in 506.

[0073] The scene in Figure 5 is shown as being implemented in the order of alert keywords, object description keywords, and qualitative location keywords, but it is specifically stated that alert keywords, direction keywords, qualitative location keywords, object description keywords, behavior keywords, or any of these categories or any order may be used in a similar procedure to that shown in Figure 5. In other words, Figure 5 is an illustrative depiction of keyword usage for identifying relevant objects, but the keywords and / or the order of keywords may differ from that in Figure 5. The order of keywords may depend heavily on the order of words in the instructor's voice, and therefore, the underlying system should ideally be able to process keywords in various orders.

[0074] Alternatively or additionally, one or more processors may be configured to identify objects using the gaze of an instructor or passenger, and / or one or more detected human gestures. Since instructors are likely to provide verbal or auditory instructions, the gaze and human gestures may function primarily to complement verbal instructions, allowing one or more processors to further refine or improve the reliability of object identification based on gaze or human gestures. Nevertheless, it is conceivable that an instructor may provide human gestures to identify hazards without also providing verbal or auditory instructions. It is therefore expressly stated that the principles and methods disclosed herein with respect to object identification and / or determination of certainty using gaze and human gestures may be used in conjunction with or independently of verbal instructions.

[0075] In one aspect of this disclosure, one or more processors may be configured to recognize gaze direction. That is, one or more processors may be configured to determine the direction of human gaze (e.g., focus of attention, direction of eyes) from image sensor data. The underlying assumption that enables gaze detection is that people frequently look at the object they are describing. From the perspective of this assumption, one or more processors may determine the direction of human gaze and associate this direction with an object near the vehicle. Once the object is determined, it may further be possible to associate the object with simultaneously spoken text or keywords. For example, an instructor who sees a cyclist moving towards the vehicle's path may gaze in the cyclist's direction while loudly speaking one or more keywords such as "Watch out!". By detecting the instructor's gaze, the gaze direction may be associated with the cyclist, and then the warning keyword "Watch out" may allow one or more processors to associate a higher level of importance or relevance with the cyclist.

[0076] During gaze direction recognition, one or more processors may be configured to determine the direction of human gaze (e.g., focus of attention, direction of eyes) from image sensor data. The underlying assumption enabling gaze detection is that people frequently look at the object they are describing. From this assumption, one or more processors may determine the direction of human gaze and associate this direction with an object near the vehicle. Once the object is determined, it may further be possible to associate the object with simultaneously spoken text or keywords. For example, an instructor who sees a cyclist moving towards the vehicle's path may gaze in the cyclist's direction while loudly speaking one or more keywords such as "Watch out!". By detecting the instructor's gaze, the direction of gaze may be associated with the cyclist, and then the warning keyword "Watch out" may allow one or more processors to associate a higher level of importance or relevance with the cyclist.

[0077] In another aspect of this disclosure, the labeling of data based on human voice may be supported by gaze (e.g., the direction the instructor is looking, an external object corresponding to the gaze) and / or gestures (e.g., the instructor pointing in a particular direction). Optionally, the inexperienced driver's response may be recorded and mapped to the instructor's comments, such mapping may occur over a limited (e.g., predetermined) time after the comments, because the relevance of such response is in most situations closely related to the stimulus (e.g., the relevance of the response may be higher if the response occurs in close temporal proximity to the stimulus (instruction), but lower if there is a temporal gap from the stimulus).

[0078] Estimating gaze direction is crucial for many human-machine interaction applications. Knowledge of gaze direction provides information about the user's focus of attention. In a real-time framework for classifying gaze direction, one or more processors may initially be configured to implement a face detector. This face detector may be any known system for face detection. According to one aspect of this disclosure, the face detector may include the Viola-Jones algorithm, which is a known framework for object detection that has been successfully implemented for face detection.

[0079] Figure 6 shows an example of gaze used for object identification. In this image, a bicycle 602 and two pedestrians 604 are located near the vehicle. Without further information, it may not be immediately clear whether the bicycle 602 or the pedestrians 604 pose a danger to the vehicle, or if they are more relevant to the vehicle. In this example, the instructor provides the verbal statement “Watch out!” and gazes at the bicycle. One or more processors may be configured to interpret the instructor’s voice and detect the words “Watch out!”. The keyword “Watch out!” alone makes it clear that there is something near the vehicle that needs to be alerted, but it is not immediately clear what the object is (e.g., is it the bicycle or the pedestrians). However, if the instructor gazes at the bicycle while saying “Watch out!”, one or more processors may link the instructor’s gaze to the bicycle, rather than considering the entire vicinity irrelevant, and thus narrow the relevant area 606 to the area around the bicycle. Similarly, if an instructor shouts "Watch out!" while pointing to a bicycle or otherwise gesturing to the bicycle, one or more processors may link the point or gesture to the bicycle, and thereby narrow down the relevant area 606 to the area around the bicycle.

[0080] Alternatively or additionally, one or more processors may be configured to utilize one or more directional keywords in conjunction with gaze and / or hand gestures to further enhance confidence in identified objects within a location, where gaze direction may be of particular importance. One or more processors may utilize gaze direction to identify particularly important objects in sensor data, and / or particularly important objects, if they are among multiple objects. Continuing with the bicycle example, assuming one or more processors detect three bicycles in the sensor data, the instructor is likely to be gazing at the bicycles considered particularly relevant or risky. One or more processors may be configured to determine a gaze direction and associate that gaze direction with external sensor data so that it can be determined that the instructor is gazing at a particular bicycle. In this way, the gazed-on bicycle may be labeled as more important than the other identified bicycles.

[0081] One or more processors may be configured to perform rough eye region detection on the detected faces after applying face detection (for example, after faces have been detected). Eye region detection may utilize any known eye region detection algorithm or procedure. Known procedures for eye region detection may rely on geometric relationships and facial markers to locate and identify eyes. Once eye regions are detected, one or more processors may be configured to classify the gaze direction. One or more processors may be configured to implement a convolutional neural network (CNN) for gaze detection. In this way, one or more processors may determine the gaze direction.

[0082] The direction of gaze may depend on the direction of the instructor's eyes. According to one aspect of this disclosure, the CNN may be configured to independently determine the gaze direction of each eye (e.g., determining the direction of the left eye, then the direction of the right eye). One or more processors may use these determined directions to calculate a fused score and classify the gaze. That is, one or more processors may be configured to determine the average or center point of two determined gaze directions such that the determined left eye direction and the determined right eye direction can be harmonized. This harmonized or fused direction score may be classified as a gaze.

[0083] Several datasets are available for gaze classification. For example, Eye Chimera is a known database that enables gaze detection. According to one aspect of this disclosure, one or more processors and / or CNNs may use Eye Chimera to detect the gaze of an instructor or other person inside a vehicle.

[0084] Figure 7 shows a gaze-tracking detector according to one aspect of the present disclosure. The gaze-tracking detector may include a face detector 702, an eye-region localizer 704, and a gaze-direction classifier 706. These components may be implemented as a single component (e.g., on a processor, a group of processors, an integrated circuit, a system-on-a-chip, etc.) or as multiple components. These components may be implemented as software to be executed by one or more processors. These components may be implemented as an ANN. The face detector 702 may receive image sensor data from inside a vehicle and may execute one or more face detection algorithms on the image sensor data. The specific face detection algorithm to be employed may be selected for a given implementation. According to one aspect of the present disclosure, the Viola-Jones algorithm may be used, but the implementation is not limited to the use of the Viola-Jones algorithm. The face detector 702 may output a label or other identifier for the portion of the image corresponding to the image sensor data in which a face was detected. The eye region localizer 704 may receive image sensor data, image sensor data corresponding to a detected face, and an identifier corresponding to an area of ​​the detected face, or any of these, and may perform eye region localization on the data corresponding to the area of ​​the detected face. During eye region localization, the eye region localizer 704 may implement one or more algorithms to determine the presence of eyes in the area corresponding to the detected face. The eye region localizer 704 may implement any known eye region localization algorithm to locate the eyes. These may include, but are not limited to, shape-based models (e.g., algorithms for detecting eyes based on a semi-elliptical head model, or algorithms for locating eye regions based on a generalized head transform), feature-based shape methods, appearance-based methods, or any combination thereof. Once the eye region localizer 704 has located the eye region, it may output an identifier for the location in the image represented by the sensor data corresponding to the eye region. The gaze direction classifier 706 may receive this identifier and determine the gaze direction.Various strategies are known for determining gaze direction from image data, and an appropriate gaze direction procedure may be selected for a given implementation. In some situations, gaze may be determined solely from eye position. In other situations, gaze may be determined from eye position relative to head position. In other situations, gaze may be determined from eye position relative to a fixed reference point, such as a part inside a vehicle. The gaze direction classifier 706 may output a gaze direction identifier that can represent the direction of gaze. The identifier may be a facial feature, a body part, a part or axis of the head, a reference point inside a vehicle, a reference point outside a vehicle, or something else.

[0085] In some aspects of this disclosure, the gaze detector may be implemented within an ANN (including, but not limited to, a CNN). The ANN may be particularly well suited for rapid evaluation of image sensor data for determining gaze direction. The ANN may be trained to detect gaze direction using annotated data 708. In general principle, it may be inferred that the instructor is gazing at an object that is the subject of the instructor's keywords. Therefore, if keywords are detected, and if one or more keywords are used to identify a highly relevant object as described herein, annotated data identifying the highly relevant object (e.g., its location, its type or identity, or otherwise) may be sent to the gaze direction classifier 706. The gaze direction classifier 706 may use this annotated data to compare the location of an object near the vehicle with the detected gaze direction. The gaze detector may determine the accuracy of the gaze direction classification by mapping the gaze direction to external image sensor data. That is, assuming that the object described by the keywords is the subject of the instructor's gaze, the gaze direction must correspond to the location of the detected object. The difference between the gaze direction and the object's location can both be used in the training phase to improve the gaze detector's results.

[0086] Of particular note is that the gaze direction may include a straight line of sight gaze direction and / or a mirror gaze direction. In the line of sight gaze direction, the vehicle may determine the direction of the instructor's gaze from internal vehicle sensor data (e.g., camera data or others). One or more processors also have sensor data corresponding to the vicinity outside the vehicle (image sensor data, lidar, radar, etc.). One or more processors may determine the speaker's gaze direction, such as gaze direction toward the speaker, gaze direction toward the vehicle, gaze direction toward a fixed point inside the vehicle, or others. Each of the internal sensor data (e.g., internal microphone, internal camera) and external sensor data (e.g., image sensor data, lidar, radar) may be time-stamped, so that the internal sensor data can be compared with the external sensor data that occurs simultaneously. One or more processors may be configured to analyze, upon detection of a gaze direction, simultaneously occurring external sensor data (e.g., external data with the same, similar, simultaneous, or overlapping timestamps as the detected gaze direction) and tag or label objects in the sensor data corresponding to the gaze as highly relevant.

[0087] One or more processors may be configured to distinguish between line-of-sight gaze on an object and mirror gaze (referred to herein as “mirror gaze”). Figure 8 illustrates the calculation of mirror gaze according to one aspect of the present disclosure. The figure shows a representation of a vehicle having a front windshield 802 and a rear windshield 804. A driver or passenger 806 is shown inside the vehicle and gazing toward the rearview mirror 808. If the gaze detector and / or one or more processors simply detect the direction (e.g., angle) of gaze and associate that direction with an object outside the vehicle, the gaze detector and / or one or more processors may overlook the rearview mirror 808 and instead assume that the gaze continues along a straight path beyond the front windshield 802, as shown in 810. If an object is located along this path, as shown in 812, the gaze detector and / or one or more processors may incorrectly associate the gaze of the driver or passenger 806 with the object 812. Instead, the gaze detector and / or one or more processors must consider the presence of the rearview mirror 808 and its effect on the gaze of the driver or passenger 806. Specifically, the driver or passenger 806 is looking at the reflection of the obstacle 814 along the reflection path 816.

[0088] To achieve this, one or more processors may be configured to determine whether the driver or passenger is looking at a mirror (e.g., an inspection mirror or other mirror) and to determine the angle of reflection. The law of reflection states that the angle of reflection is equal to the angle of incidence. That is,

number

[0089] As described above, according to another aspect of this disclosure, one or more processors may be configured to recognize one or more human gestures in image sensor data (e.g., data relating to the driver and / or passengers from an inward-facing camera). Such human gestures may include at least the following: a) Pointing in a specific direction: Instructors may give commands while pointing in a specific direction. Such pointing is often performed in conjunction with verbal references. That is, instructors may name an object (e.g., using object description keywords) and point to the object. Pointing to an object may be performed in conjunction with a statement about the object's location (e.g., using directional keywords or qualitative location keywords). In this case, pointing helps to reinforce the identification of the object through the verbal command, or otherwise simplify the identification of the object. Alternatively, pointing may be performed instead of a statement about the object's location, such as simply pointing to the object and saying its name (e.g., "cyclist"). b) Attention signs (e.g., raising the index finger): Certain gestures may be associated with an increased need for attention. These gestures may be culturally specific, and consequently, there may be no universally applicable gesture. Rather, specific gestures associated with attention may be selected to meet the specific implementation needs. For example, in some countries, raising the index finger without any further verbal or nonverbal communication may indicate an increased need for attention. c) Negative signs: Certain gestures may be associated with negation that suggests the current sequence of actions is wrong and should be stopped, or that previous instructions should be ignored. Such negative gestures may also be culturally specific and may be selected for a given implementation. Examples of such negative gestures may include, but are not limited to, shaking the head from side to side, shaking one or both hands from side to side, or others. d) Stop gestures: Certain gestures may be associated with the need to stop. Such stop gestures may also be culturally specific and may be selected for a given implementation. Known gestures associated with the need to stop include, but are not limited to, extending one arm forward with the wrist bent, or extending both arms forward with the wrists bent and nearly parallel to each other.

[0090] According to one aspect of this disclosure, one or more processors and / or ANNs may be configured to detect one or more hand gestures and / or to detect the orientation of one or more hand gestures. Gesture recognition has been used in many systems for a variety of purposes. Gesture recognition in image databases is known, and any suitable method for determining gestures in image data may be used.

[0091] Figure 9 shows a hand gesture detector that may be configured to detect one or more hand gestures or hand positions as described above. The hand gesture detector may include a hand detector 902 that can receive image sensor data (e.g., image sensor data from an inward-facing camera / image sensor data inside a vehicle) and may employ one or more hand detection algorithms to detect a hand in the image sensor data. Various known hand detection algorithms and procedures can be used. For example, known hand detection algorithms may use skin color, hand shape, protrusions on the outer surface of the hand, finger shape, or any of these to detect a hand. Whatever the selected implementation, the hand detector detects a hand in the image data. The handle localizer 904 may then locate the position of the detected hand and output the hand position relative to the image represented by the image sensor data. Once a hand is detected and its position is determined, the hand gesture recognition device 908 may identify a hand gesture.

[0092] According to certain aspects of this disclosure, the hand gesture recognition device 908 may be configured as an ANN (including, for example, a CNN, but not limited to, the following). For example, known implementations of CNNs as hand gesture recognition devices have been shown to be able to identify subtle hand gestures such as left / right swipes, up / down flicks, taps, none, and other actions from image sensor data. Such CNNs may be further trained to identify other hand gestures such as pointing, attention signs, stop gestures, or any other desirable hand gestures.

[0093] The next step for the hand gesture detector depends on the decision of the hand gesture recognition device 908. If the recognized hand gesture in 908 is not a pointing gesture (e.g., a warning sign, a stop gesture, etc.), the hand gesture recognition device 908 may output an identification of the determined hand gesture. The determined hand gesture may be associated with one or more keywords or actions. For example, a warning sign may be treated similarly to a warning keyword, such as initially labeling the entire vicinity of the vehicle as highly relevant. Based on the warning sign, one or more processors may use gaze, other keywords, subsequent pointing, or any of these to further refine or more carefully identify objects in the image sensor data that are highly relevant.

[0094] If the hand gesture recognized by the hand gesture recognition device 908 is a pointing gesture, the hand gesture recognition device 908 may output an identifier corresponding to the pointing gesture to the hand direction classifier 912. The hand direction classifier 912 may be implemented as an ANN (including, but not limited to, a CNN). The hand direction classifier may be configured to determine the direction in which the hand is pointing.

[0095] Similar to gaze classification, hand direction classification may be trained with annotated data. As a general principle, if an instructor is pointing, it may be assumed that the instructor is pointing to the object that is the subject of the instructor's keyword. Therefore, if a keyword is detected, and if one or more keywords are used to identify a highly relevant object as described herein, annotated data identifying the highly relevant object (e.g., its location, type or identity, or otherwise) may be sent to the hand direction classifier 912. The hand direction classifier 912 may use this annotated data to compare the location of an object near the vehicle with the detected hand direction. The hand detector may determine the accuracy of the hand direction classification by mapping the hand direction to external image sensor data. That is, assuming that the object described by the keyword is the subject of the instructor's pointing, the hand direction must correspond to the location of the detected object. Any difference between the hand direction and the object location may be used in the training phase to improve the hand detector's results.

[0096] In another aspect of this disclosure, one or more processors may increase or decrease the certainty that a detected object is the subject of an instructor's attention based on a pointing gesture. Humans often point to objects of their attention. One or more processors may be configured to determine a pointing gesture from image sensor data (e.g., one or more cameras near the vehicle, one or more cameras directed at the occupants of the vehicle). One or more processors may associate a pointing gesture with direction to the same extent that gaze is associated with direction. The direction of a pointing gesture may then be determined in relation to external sensor data (e.g., cameras, lidars, radars) configured to detect information about the outside or near the vehicle, and the direction of a pointing gesture may then be associated with one or more objects in the external image sensor data. If the objects in the external sensor data have already been identified (e.g., from verbal keywords), the pointing gesture may be used to increase or decrease the certainty of the detected object. In other words, if the object identified by the verbal keyword corresponds to the object in the direction of the pointing gesture, the certainty that the instructor's attention is directed towards that object may increase. Conversely, if the object identified by the verbal keyword corresponds to a different object than the one that appears to be in the direction of the pointing gesture, the certainty that the instructor's attention is directed towards that object may decrease.

[0097] In one aspect of this disclosure, one or more processors may be configured to attribute an increased risk level to all external sensor data during a predetermined duration associated with a warning keyword. As described above, the predetermined duration may begin immediately before, during, or immediately after speaking the keyword. The predetermined duration may be configurable. In this way, the duration during which the importance of sensor data associated with the keyword is increased may be configured for implementations based on any desired factors. These factors may include, but are not limited to, the instructor's personal attributes, a particular type of keyword, a particular type of hazard, or regional or cultural differences.

[0098] During a predetermined duration in which the relevance of sensor data increases, one or more processors may be configured to further specify hazardous areas by utilizing voice, gaze, gestures, or any of these as additional context. This will be described in more detail here.

[0099] Figure 10 shows a data synthesis apparatus and labeler according to one embodiment of the present disclosure. In this figure, microphone data from inside the vehicle is shown at 1002, and image sensor data corresponding to the vicinity of the vehicle is shown at 1004. At 1003, one or more processors detect voice data keywords corresponding to the timestamp 10:42:22:06 (the sample timestamp is provided for illustrative purposes only and is not intended to limit). In this case, the voice data keywords include "Watch out! Child! On the road!". After identifying the voice keywords, one or more processors may use the timestamp of the image sensor data 1004 to identify the corresponding section of the image sensor data. One or more processors may be configured to find a specific portion of the image sensor data corresponding to the timestamp at 1003. Alternatively or additionally, one or more processors may be configured to find portions of the image sensor data corresponding to a time slightly before and / or slightly after the timestamp at 1003. For example, given the understanding that human voices indicating danger are generally uttered shortly after the danger is first detected, one or more processors may be configured to consider image sensor data corresponding to 10:42:21:06 to 10:42:24:06 (e.g., one second before and two seconds after the keyword). Needless to say, the duration before and after the keyword considered by one or more processors is a matter of preference and implementation, and should not be understood as limiting.

[0100] One or more processors may, after identifying a corresponding section 1005 of image sensor data, examine the image sensor data to identify one or more objects corresponding to keyword 1003 of microphone data. Continuing the above example, one or more processors may identify image sensor data 1005 corresponding to a child on the road. If an object corresponding to one or more of the keywords in the microphone data is located in the corresponding section of the image sensor data, one or more processors may generate a label corresponding to the detected object. Details of label generation may depend on a given implementation. According to one aspect of this disclosure, the label may include an image data identifier representing the portion of image data corresponding to the detected keyword, an object identifier representing the identity of the object corresponding to the detected keyword, an object label representing the name or type of the identified object, a priority label representing the relevance of the detected object, or any of these. The label may be part of the external image sensor data (e.g., labeled data) or it may be independent of the external image sensor data (e.g., stored separately from the external image sensor data).

[0101] In some implementations, it may be desirable for one or more processors within the vehicle to perform the keyword and image sensor matching procedure described herein. In such configurations, each vehicle may include one or more processors configured to identify keywords and microphone data and correlate the identified keywords with objects in the image sensor data, as described herein. This may be performed in real time or at any given latency.

[0102] According to another aspect of this disclosure, it may be desirable for one or more central databases to perform the keyword and image sensor matching procedures described herein. Figure 11 shows a data storage device according to one aspect of this disclosure, where the data storage device is configured to receive microphone data and external image sensor data for label generation. In this implementation, one or more vehicles may be equipped with one or more data storage modules that can be configured to receive and store at least microphone data and image sensor data. The microphone data and image sensor data may be time-stamped to enable temporal comparison of the data streams. The actual type of data storage module is generally not important and may include, but is not limited to, one or more hard drives, one or more solid-state drives, one or more optical drives, or others. The data stored in the data storage module may be transferred from time to time to one or more central databases. This transfer may be performed using any data transfer method without limitation. The transfer may be performed as a wired or wireless transfer. Alternatively or additionally, one or more elements of the data storage module may be physically removed from the vehicle and connected directly to one or more servers for uploading. This configuration may be used in any given implementation. One example of such an implementation may be in the context of a driving school, where the school uses multiple vehicles for driving instructions. Each of the vehicles may record and store its own microphone data and external image sensor data, which may then be uploaded to a central database for processing, sometimes or periodically.

[0103] Figure 12 shows a vehicle data association device 1200 according to one embodiment of the present disclosure. The vehicle data association device may include an internal audio / image data analyzer 1202, an external image analyzer 1204, and an object data generation device 1206, wherein the internal audio / image data analyzer 1202 is configured to identify second data representing an audio indicator or an image indicator within first data representing at least one of audio from inside the vehicle or images from inside the vehicle, where the audio indicator is human voice and the image indicator represents human behavior inside the vehicle; the external image analyzer 1204 is configured to identify an object corresponding to at least one of the audio indicator or a video indicator within third data representing images near the outside of the vehicle; and the external image analyzer 1206 is configured to generate object data and classify the third data.

[0104] Figure 13 shows a method for associating vehicle data. This method comprises the steps of: identifying a second data representing an audio indicator or an image indicator within a first data representing at least one of audio from inside the vehicle or an image from inside the vehicle, wherein the audio indicator is human voice and the image indicator represents human behavior inside the vehicle; identifying an object corresponding to at least one of the audio indicator or video indicator within a third data representing an image near the outside of the vehicle; and generating object data to classify the third data.

[0105] In some aspects of this disclosure, one or more methods may be optionally employed to cross-validate the concept of hazard. For example, one or more responsibility-based safety algorithms or other risk assessment procedures may be used, for example, to determine a measure of discrepancy between data and estimation models.

[0106] In one aspect of this disclosure, confidence may be associated with an object identified due to the direction of gaze. That is, the detection of gaze and the relationship between the detected gaze and an object in the external sensor data may depend on several variables, each having a specific range of error. Thus, confidence may be assigned to an object considered to be associated with the detected gaze, where confidence indicates the likelihood or certainty that the labeled object corresponds to the gaze (e.g., representing an object that is being focused on by the instructor). Various further relationships (e.g., keywords and / or gestures) may be used to increase certainty.

[0107] In some aspects of this disclosure, direction keywords may be used to increase certainty regarding the direction of gaze and / or an object that is thought to correspond to the direction of gaze. While instructors may not always include direction keywords, if they do, one or more processors may use those direction keywords to identify an object in external sensor data that is thought to correspond to the instructor's attention. If one or more processors have already linked an object to the instructor's gaze, the confidence associated with the object identification may be increased by adding a direction keyword corresponding to the same object. Conversely, if an object keyword suggests an object other than one previously identified as likely to be associated with the instructor's gaze, the confidence associated with the previously identified object may be decreased.

[0108] One or more processors may label hazardous conditions using any of several levels of detail by employing the principles and methods described herein. That is, one or more processors may determine a general concept of the hazard (for example, generally, by identifying a hazard type or object type), or they may determine a specific object associated with the hazard, or even a type of hazard associated with that specific object.

[0109] In some aspects of this disclosure, one or more processors may be configured to terminate a labeling session after a predetermined duration following a keyword and / or gesture. As described above, identified hazards are most relevant when they are temporally close to the spoken keyword or gesture, and consequently, the relevance of identifying a corresponding object or labeling a corresponding hazard decreases as the temporal distance from the keyword and / or gesture increases. One way to address this is to define a predetermined duration from the keyword or gesture at which the labeling procedure terminates. For example, the labeling procedure may terminate 0.5 seconds, 1 second, 2 seconds, 5 seconds, or 10 seconds after the keyword or gesture. In some aspects of this disclosure, this predetermined duration may be configurable for a specific instruction, a specific context, a specific culture or country, or other reasons. Alternatively or additionally, the relevance of the identified object or hazard may be inversely proportional to the time from the keyword or gesture. In this way, one or more processors may be configured to assign associations to objects or dangers (for example, in labels associated with objects), where the association may increase as it approaches a keyword or gesture in time, and decrease as it approaches a keyword or gesture in time.

[0110] In some aspects of this disclosure, the subject matter of this disclosure may enable observation of human interaction during driving instructions, for example, rather than merely observing human driving, as is done in conventional autonomous vehicle training. This allows the underlying system to acquire further information about dangerous situations that might otherwise be missed in conventional training, such as when only other traffic participants are being observed. As disclosed herein, the vehicle learns not by copying human driving behavior, but rather by learning how to focus on potentially dangerous situations, which may not necessarily result in a direct response from a human driver, but can rather be used to focus on the most relevant parts of image sensor data. The labeling of such training data may be performed semi-automatically; that is, the training data may be labeled without further involvement from human participants, even if human behavior is used for labeling.

[0111] According to certain aspects of this disclosure, labeled data may include correct responses to specific scenarios. That is, one or more processors may use instructions such as those described herein to search for relevant or hazardous situations in image sensor data. In addition to simply identifying these relevant or hazardous situations, one or more processors may record or acquire desired responses to the scenarios in the form of driver reactions.

[0112] Further aspects of this disclosure are disclosed as an example.

[0113] Example 1: A vehicle data association device comprising an internal audio / image data analyzer, an external image analyzer, and an object data generation device, wherein the internal audio / image data analyzer is configured to identify, in first data representing at least one of audio from inside the vehicle or images from inside the vehicle, a second data representing an audio indicator or an image indicator, wherein the audio indicator is a human voice and the image indicator represents human behavior inside the vehicle; the external image analyzer is configured to receive a third data representing an image near the outside of the vehicle and to identify, in the third data representing an image near the outside of the vehicle, an object data generation device is configured to generate object data and classify the third data.

[0114] Example 2: The vehicle data association device according to Example 1, wherein the identity of the object includes one or more coordinates that define the boundaries of the object.

[0115] Example 3: The vehicle data association device according to Example 1 or 2, wherein identifying the second data includes the internal audio / image data analyzer identifying one or more keywords in the audio.

[0116] Example 4: The vehicle data association device according to Example 3, wherein the one or more keywords include at least one of the following: alert keywords, direction keywords, qualitative location keywords, object description keywords, behavior keywords, or any combination thereof.

[0117] Example 5: The vehicle data association device according to Example 3 or 4, wherein the internal audio / image data analyzer is configured to transmit data representing one or more keywords to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the one or more keywords.

[0118] Example 6: The vehicle data association device according to any one of Examples 3 to 5, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords.

[0119] Example 7: The vehicle data association device according to any one of Examples 1 to 6, wherein identifying the second data includes the internal audio / image data analysis device identifying human gestures in the video.

[0120] Example 8: The vehicle data association device according to Example 7, wherein the human gesture is at least one of pointing in a direction, making an attention gesture, making a denial gesture, making a stop gesture, or any combination thereof.

[0121] Example 9: The vehicle data association device according to Example 7 or 8, wherein the internal audio / image data analyzer is configured to transmit data representing the gesture to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the gesture.

[0122] Example 10: The vehicle data association device according to any one of Examples 7 to 9, wherein the internal audio / image data analyzer is configured to transmit data representing the gesture to the external image analyzer, and the external image analyzer is configured to map the pointing action to the third data and to identify the object based on the mapped relationship between the pointing action and the object.

[0123] Example 11: The vehicle data association device according to any one of Examples 1 to 10, wherein identifying the second data includes the internal audio / image data analysis device identifying the direction of human gaze in the video.

[0124] Example 12: The vehicle data association device described in Example 11, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the Viola-Jones algorithm.

[0125] Example 13: The vehicle data association device described in Example 11, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the Viola-Jones algorithm or a similar algorithm.

[0126] Example 14: A vehicle data association device according to any one of Examples 11 to 13, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the eye-chimera (EYE part from the Cognitive process Inference by the mutual use of the eye and expRession Analysis).

[0127] Example 15: A vehicle data association device according to any one of Examples 12 to 14, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using an Eye-Chimera or similar algorithm for cognitive process inference through the interuse of eye and expression analysis.

[0128] Example 16: The vehicle data association device according to any one of Examples 13 to 15, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the gaze direction.

[0129] Example 17: A vehicle data association device according to any one of Examples 13 to 16, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to map the gaze direction to the third data and to identify the object based on the mapped relationship between the gaze direction and the object.

[0130] Example 18: The vehicle data association device according to any one of Examples 1 to 17, wherein the image from inside the vehicle includes an image of the driver or passenger of the vehicle.

[0131] Example 19: The vehicle data association device according to any one of Examples 1 to 18, wherein the external image analysis device is configured to identify the object based on at least two of one or more keywords, one or more gestures, or gaze directions.

[0132] Example 20: The first data includes image sensor data and / or microphone data, as described in any one of Examples 1 to 19, of the vehicle data association device.

[0133] Example 21: A vehicle data association device according to any one of Examples 1 to 20, wherein the third data includes image sensor data, light detection and ranging (LIDAR) sensor data, radio wave detection and ranging (RADAR) sensor data, or any combination thereof.

[0134] Example 22: The vehicle data association device according to any one of Examples 1 to 21, further comprising a vehicle sensor data analyzer configured to receive vehicle sensor data from a vehicle sensor, and the external image analyzer further configured to generate object data based on the vehicle sensor data.

[0135] Example 23: The vehicle data association device described in Example 22, wherein the vehicle sensors include a steering sensor, an accelerometer, a braking sensor, a speedometer, or any combination thereof.

[0136] Example 24: The vehicle data association device according to any one of Examples 1 to 23, further comprising a vehicle actuator data analyzer configured to receive actuator data, and the external image analyzer further configured to generate object data based on the vehicle actuator data.

[0137] Example 25: The vehicle data association device described in Example 24, wherein the vehicle actuator data includes data representing steering wheel position, brake position, brake pedal depression, braking force, speed, velocity, acceleration, or any combination thereof.

[0138] Example 26: A vehicle data association device according to any one of Examples 22 to 25, wherein the sensor data analyzer and / or the vehicle actuator data analyzer is configured to determine the behavior of the vehicle related to an object represented by the object data from the sensor data and / or the vehicle actuator data.

[0139] Example 27: The vehicle data association device according to any one of Examples 1 to 26, further comprising an artificial neural network, wherein the internal audio / image data analyzer, the external image analyzer, or the object data generation device is implemented as the artificial neural network.

[0140] Example 28: The vehicle data association device according to any one of Examples 1 to 27, further comprising a memory configured to store the first data, the second data, the third data, the object data, or any combination thereof.

[0141] Example 29: A vehicle data association device according to any one of Examples 1 to 28, wherein at least two of the first data, the second data, the third data, the object data, or any combination thereof each include a timestamp, and multiple data sources are synchronized via the timestamp.

[0142] Example 30: A vehicle data association device according to any one of Examples 1 to 29, wherein the object data generation device is configured to associate an object with the second data for a predetermined period of time following the audio indicator or the image indicator, and the object data generation device is configured not to associate the object with the second data after the predetermined period of time following the audio indicator or the image indicator has ended.

[0143] Example 31: A vehicle data association device according to any one of Examples 1 to 30, wherein the object data generation device is configured to determine the priority of the object based on the audio indicator or the image indicator.

[0144] Example 32: The vehicle data association device as described in Example 31, wherein the priority is based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated amount of damage associated with a collision with the object, or any combination thereof.

[0145] Example 33: The vehicle data association device according to any one of Examples 1 to 32, wherein the object data includes labels for one or more detected objects.

[0146] Example 34: The vehicle data association device according to any one of Examples 1 to 33, wherein the object data generation device is further configured to generate sensor data labels, the sensor data labels being labels representing at least one of the identity of the object, the behavior of the object, or the priority of the object.

[0147] Example 35: The vehicle data association device according to Example 34, wherein the object data generation device is further configured to output sensor data including the sensor data label.

[0148] Example 36: The vehicle data association device according to Example 35 or 34, wherein the output sensor data includes the third data and the sensor data label.

[0149] Example 37: A non-temporary computer-readable medium comprising instructions, the instructions, when executed, cause one or more processors to identify second data representing an audio indicator or an image indicator in first data representing at least one of audio from inside the vehicle or images from inside the vehicle, wherein the audio indicator is human voice and the image indicator represents human behavior inside the vehicle; identify an object corresponding to at least one of the audio indicator or the video indicator in third data representing images near the outside of the vehicle; and generate object data to classify the third data.

[0150] Example 38: The non-temporary computer-readable medium according to Example 37, wherein the identity of the object includes one or more coordinates defining the boundaries of the object.

[0151] Example 39: A non-temporary computer-readable medium according to Example 37 or 38, wherein identifying the second data includes identifying one or more keywords in the audio.

[0152] Example 40: The non-temporary computer-readable medium according to Example 39, wherein one or more keywords include at least one of the following: alert keywords, directional keywords, qualitative location keywords, object description keywords, behavioral keywords, or any combination thereof.

[0153] Example 41: A non-temporary computer-readable medium according to Example 39 or 40, wherein the instruction is further configured to cause one or more processors to transmit data representing one or more keywords to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and one or more keywords.

[0154] Example 42: The non-temporary computer-readable medium according to any one of Examples 39 to 41, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords.

[0155] Example 43: A non-temporary computer-readable medium according to any one of Examples 37 to 42, wherein identifying the second data includes identifying human gestures in the video.

[0156] Example 44: The non-temporary computer-readable medium according to Example 43, wherein the human gesture is at least one of pointing in a direction, making an attention gesture, making a denial gesture, making a stop gesture, or any combination thereof.

[0157] Example 45: A non-temporary computer-readable medium according to any one of Examples 37 to 44, wherein identifying the second data includes identifying the direction of human gaze in the video.

[0158] Example 46: The non-temporary computer-readable medium according to Example 45, wherein the instruction is further configured to cause one or more processors to identify the direction of human gaze by using the Viola-Jones algorithm.

[0159] Example 47: A non-temporary computer-readable medium according to either Example 45 or 46, wherein the instruction is further configured to cause one or more processors to identify the direction of human gaze by using the eye portion (Eye-Chimera) from cognitive process inference by the interuse of eye and expression analysis.

[0160] Example 48: A non-temporary computer-readable medium according to any one of Examples 45 to 47, wherein the instruction is further configured to cause one or more processors to identify the direction of human gaze by using an Eye-Chimera or similar algorithm for cognitive process inference through the interuse of eye and expression analysis.

[0161] Example 49: The image from inside the vehicle is a non-temporary computer-readable medium according to any one of Examples 37 to 48, including an image of the driver or passenger of the vehicle.

[0162] Example 50: A non-temporary computer-readable medium according to any one of Examples 37 to 49, wherein the instruction is further configured to cause one or more processors to identify the object based on at least two of one or more keywords, one or more gestures, or gaze directions.

[0163] Example 51: The first data is a non-temporary computer-readable medium as described in any one of Examples 37 to 50, including image sensor data and / or microphone data.

[0164] Example 52: The third data is a non-temporary computer-readable medium according to any one of Examples 37 to 51, comprising image sensor data, light detection and ranging (LIDAR) sensor data, radio wave detection and ranging (RADAR) sensor data, or any combination thereof.

[0165] Example 53: A non-temporary computer-readable medium according to any one of Examples 37 to 52, wherein the instruction is further configured to cause one or more processors to receive vehicle sensor data from a vehicle sensor and generate object data based on the vehicle sensor data.

[0166] Example 54: The vehicle sensor is a non-temporary computer-readable medium as described in Example 53, including a steering sensor, an accelerometer, a braking sensor, a speedometer, or any combination thereof.

[0167] Example 55: A non-temporary computer-readable medium according to any one of Examples 37 to 54, wherein the instruction is further configured to cause one or more processors to receive actuator data, and the external image analysis device is further configured to generate object data based on the vehicle actuator data.

[0168] Example 56: The vehicle actuator data includes data representing steering wheel position, brake position, brake pedal depression, braking force, speed, velocity, acceleration, or any combination thereof, in the non-temporary computer-readable medium described in Example 55.

[0169] Example 57: The non-temporary computer-readable medium according to any one of Examples 53 to 56, wherein the instruction is further configured to cause one or more processors to determine from the sensor data and / or the vehicle actuator data the behavior of the vehicle relating to an object represented by the object data.

[0170] Example 58: The instruction is a non-temporary computer-readable medium as described in any one of Examples 37 to 57, implemented within an artificial neural network.

[0171] Example 59: A non-temporary computer-readable medium according to any one of Examples 37 to 58, wherein at least two of the first data, the second data, the third data, the object data, or any combination thereof each include a timestamp, and multiple data sources are synchronized via the timestamp.

[0172] Example 60: A non-temporary computer-readable medium according to any one of Examples 37 to 59, wherein the instruction is further configured to cause one or more processors to associate an object with the second data for a predetermined period of time following the audio indicator or the image indicator, and the object is not associated with the second data after the predetermined period of time following the audio indicator or the image indicator ends.

[0173] Example 61: A non-temporary computer-readable medium according to any one of Examples 37 to 60, wherein the instruction is further configured to cause one or more processors to determine the priority of the object based on the audio indicator or the image indicator.

[0174] Example 62: The non-temporary computer-readable medium described in Example 61, wherein the priority is based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated amount of damage associated with a collision with the object, or any combination thereof.

[0175] Example 63: The object data is a non-temporary computer-readable medium according to any one of Examples 37 to 62, comprising labels for one or more detected objects.

[0176] Example 64: The non-temporary computer-readable medium according to any one of Examples 37 to 63, wherein the instruction is further configured to cause one or more processors to generate a sensor data label, the sensor data label being a label representing at least one of the identity of the object, the behavior of the object, or the priority of the object.

[0177] Example 65: The non-temporary computer-readable medium according to Example 64, wherein the instruction is further configured to cause one or more processors to output sensor data including the sensor data label.

[0178] Example 66: The output sensor data is a non-temporary computer-readable medium as described in Example 65 or 64, including the third data and the sensor data label.

[0179] Example 67: A means for associating vehicle data, the means for associating vehicle data comprising an internal audio / image data analyzer, an external image analyzer, and an object data generation device, wherein the internal audio / image data analyzer is configured to identify a second data representing an audio indicator or an image indicator in first data representing at least one of audio from inside the vehicle or images from inside the vehicle, wherein the audio indicator is a human voice and the image indicator represents human behavior inside the vehicle; the external image analyzer is configured to identify an object corresponding to at least one of the audio indicator or the video indicator in third data representing an image near the outside of the vehicle; and the external image analyzer is configured to generate object data and classify the third data.

[0180] Example 68: The vehicle data association means according to Example 67, wherein the identity of the object includes one or more coordinates that define the boundaries of the object.

[0181] Example 69: The means for associating vehicle data according to Example 67 or 68, wherein identifying the second data includes the internal audio / image data analyzer identifying one or more keywords in the audio.

[0182] Example 70: The vehicle data association means according to Example 69, wherein the one or more keywords include at least one of a warning keyword, a direction keyword, a qualitative location keyword, an object description keyword, a behavior keyword, or any combination thereof.

[0183] Example 71: The vehicle data association means according to Example 69 or 70, wherein the internal audio / image data analyzer is configured to transmit data representing one or more keywords to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the one or more keywords.

[0184] Example 72: The vehicle data association means according to any one of Examples 69 to 71, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords.

[0185] Example 73: The means of vehicle data association according to any one of Examples 67 to 72, wherein identifying the second data includes the internal audio / image data analyzer identifying human gestures in the video.

[0186] Example 74: The means for associating vehicle data according to Example 73, wherein the human gesture is at least one of pointing in a direction, making an attention gesture, making a denial gesture, making a stop gesture, or any combination thereof.

[0187] Example 75: The vehicle data association means according to Example 73 or 74, wherein the internal audio / image data analyzer is configured to transmit data representing the gesture to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the gesture.

[0188] Example 76: The vehicle data association means according to any one of Examples 73 to 75, wherein the internal audio / image data analyzer is configured to transmit data representing the gesture to the external image analyzer, and the external image analyzer is configured to map the pointing action to the third data and to identify the object based on the mapped relationship between the pointing action and the object.

[0189] Example 77: The means of vehicle data association according to any one of Examples 67 to 76, wherein identifying the second data includes the internal audio / image data analyzer identifying the direction of human gaze in the video.

[0190] Example 78: The vehicle data association means described in Example 77, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the Viola-Jones algorithm.

[0191] Example 79: The vehicle data association means described in Example 77, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the Viola-Jones algorithm or a similar algorithm.

[0192] Example 80: A means of vehicle data association according to any one of Examples 77 to 79, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using the eye portion (Eye-Chimera) from cognitive process inference through the interuse of eye and expression analysis.

[0193] Example 81: A means of vehicle data association according to any one of Examples 78 to 80, wherein the internal audio / image data analysis device is configured to identify the direction of human gaze by using an Eye-Chimera or similar algorithm for cognitive process inference through the interuse of eye and expression analysis.

[0194] Example 82: The vehicle data association means according to any one of Examples 79 to 81, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the gaze direction.

[0195] Example 83: The vehicle data association means according to any one of Examples 79 to 82, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to map the gaze direction to the third data and to identify the object based on the mapped relationship between the gaze direction and the object.

[0196] Example 84: The vehicle data association means according to any one of Examples 67 to 83, wherein the image from inside the vehicle includes an image of the driver or passenger of the vehicle.

[0197] Example 85: The vehicle data association means according to any one of Examples 67 to 84, wherein the external image analysis device is configured to identify the object based on at least two of one or more keywords, one or more gestures, or gaze directions.

[0198] Example 86: The means for associating vehicle data according to any one of Examples 67 to 85, wherein the first data includes image sensor data and / or microphone data.

[0199] Example 87: The vehicle data association means according to any one of Examples 67 to 86, wherein the third data includes image sensor data, light detection and ranging (LIDAR) sensor data, radio wave detection and ranging (RADAR) sensor data, or any combination thereof.

[0200] Example 88: The vehicle data association means according to any one of Examples 67 to 87, wherein the vehicle data association device further comprises a vehicle sensor data analyzer configured to receive vehicle sensor data from a vehicle sensor, and the external image analyzer is further configured to generate object data based on the vehicle sensor data.

[0201] Example 89: The means for associating vehicle data as described in Example 88, wherein the vehicle sensors include a steering sensor, an accelerometer, a braking sensor, a speedometer, or any combination thereof.

[0202] Example 90: The vehicle data association means according to any one of Examples 67 to 89, wherein the vehicle data association device further comprises a vehicle actuator data analyzer configured to receive actuator data, and the external image analyzer further comprises an external image analyzer configured to generate object data based on the vehicle actuator data.

[0203] Example 91: The vehicle data association means described in Example 90, wherein the vehicle actuator data includes data representing steering wheel position, brake position, brake pedal depression, braking force, speed, velocity, acceleration, or any combination thereof.

[0204] Example 92: The vehicle data association means according to any one of Examples 88 to 91, wherein the sensor data analyzer and / or the vehicle actuator data analyzer is configured to determine the behavior of the vehicle related to an object represented by the object data from the sensor data and / or the vehicle actuator data.

[0205] Example 93: The vehicle data association means according to any one of Examples 67 to 92, wherein the vehicle data association device further comprises an artificial neural network, and the internal audio / image data analysis device, the external image analysis device, or the object data generation device is implemented as the artificial neural network.

[0206] Example 94: A vehicle data association means according to any one of Examples 67 to 93, further comprising a memory configured to store the first data, the second data, the third data, the object data, or any combination thereof.

[0207] Example 95: A means for associating vehicle data according to any one of Examples 67 to 94, wherein at least two of the first data, the second data, the third data, the object data, or any combination thereof each include a timestamp, and the multiple data sources are synchronized via the timestamp.

[0208] Example 96: The vehicle data association means according to any one of Examples 67 to 95, wherein the object data generation device is configured to associate an object with the second data for a predetermined period of time following the audio indicator or the image indicator, and the object data generation device is configured not to associate the object with the second data after the predetermined period of time following the audio indicator or the image indicator has ended.

[0209] Example 97: The vehicle data association means according to any one of Examples 67 to 96, wherein the object data generation device is configured to determine the priority of the object based on the audio indicator or the image indicator.

[0210] Example 98: The vehicle data association means described in Example 97, wherein the priority is based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated amount of damage associated with a collision with the object, or any combination thereof.

[0211] Example 99: The vehicle data association means according to any one of Examples 67 to 98, wherein the object data includes labels for one or more detected objects.

[0212] Example 100: The vehicle data association means according to any one of Examples 67 to 99, wherein the object data generation device is further configured to generate sensor data labels, the sensor data labels being labels representing at least one of the identity of the object, the behavior of the object, or the priority of the object.

[0213] Example 101: The vehicle data association means according to Example 100, wherein the object data generation device is further configured to output sensor data including the sensor data label.

[0214] Example 102: The output sensor data includes the third data and the sensor data label, as per the vehicle data association means described in Example 101 or 100.

[0215] Example 103: A method for associating vehicle data, comprising the steps of: identifying a second data representing an audio indicator or an image indicator in first data representing at least one of audio from inside the vehicle or an image from inside the vehicle, wherein the audio indicator is a human voice and the image indicator represents human behavior inside the vehicle; identifying an object corresponding to at least one of the audio indicator or the video indicator in third data representing an image of the vicinity of the outside of the vehicle; and generating object data to classify the third data.

[0216] Example 104: The vehicle data association method according to Example 103, wherein the identity of the object includes one or more coordinates that define the boundaries of the object.

[0217] Example 105: The vehicle data association method according to Example 103 or 104, wherein the step of identifying second data includes the step of identifying one or more keywords in the audio.

[0218] Example 106: The vehicle data association method according to Example 105, wherein the one or more keywords include at least one of a warning keyword, a direction keyword, a qualitative location keyword, an object description keyword, a behavior keyword, or any combination thereof.

[0219] Example 107: The vehicle data association method according to Example 105 or 106, further comprising the step of transmitting data representing one or more keywords to an external image analyzer, the external image analyzer being configured to identify the object based on the relationship between the object and the one or more keywords.

[0220] Example 108: The vehicle data association method according to any one of Examples 105 to 107, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords.

[0221] Example 109: A vehicle data association method according to any one of Examples 103 to 108, wherein the step of identifying second data includes the step of identifying human gestures in the video.

[0222] Example 110: The vehicle data association method according to Example 109, wherein the human gesture is at least one of pointing in a direction, making an attention gesture, making a denial gesture, making a stop gesture, or any combination thereof.

[0223] Example 111: The vehicle data association method according to Example 109 or 110, further comprising the step of identifying the object based on the relationship between the object and the gesture.

[0224] Example 112: The vehicle data association method according to any one of Examples 109 to 111, further comprising the steps of mapping a pointing action to the third data and identifying the object based on the mapped relationship between the pointing action and the object.

[0225] Example 113: A method of vehicle data association according to any one of Examples 103 to 112, wherein identifying the second data includes the internal audio / image data analyzer identifying the direction of human gaze in the video.

[0226] Example 114: The vehicle data association method according to Example 113, further comprising the step of identifying the direction of human gaze by using the Viola-Jones algorithm.

[0227] Example 115: The vehicle data association method according to any one of Examples 112 to 114, further comprising a step of identifying the direction of human gaze by using the eye portion (Eye-Chimera) from cognitive process inference through the interuse of eye and expression analysis.

[0228] Example 116: The vehicle data association method according to any one of Examples 113 to 115, further comprising a step of identifying the direction of human gaze by using an Eye-Chimera or similar algorithm for cognitive process inference through the interuse of eye- and expression analysis.

[0229] Example 117: The vehicle data association method according to any one of Examples 114 to 116, further comprising the step of identifying the object based on the relationship between the object and the gaze direction.

[0230] Example 118: The vehicle data association method according to any one of Examples 114 to 117, further comprising the steps of mapping the gaze direction to the third data and identifying the object based on the mapped relationship between the gaze direction and the object.

[0231] Example 119: The vehicle data association method according to any one of Examples 103 to 118, wherein the image from inside the vehicle includes an image of the driver or passenger of the vehicle.

[0232] Example 120: The vehicle data association method according to any one of Examples 103 to 119, further comprising the step of identifying the object based on at least two of one or more keywords, one or more gestures, or gaze direction.

[0233] Example 121: The vehicle data association method described in any one of Examples 103 to 120, wherein the first data includes image sensor data and / or microphone data.

[0234] Example 122: The vehicle data association method according to any one of Examples 103 to 121, wherein the third data includes image sensor data, light detection and ranging (LIDAR) sensor data, radio detection and ranging (RADAR) sensor data, or any combination thereof.

[0235] Example 123: The vehicle data association method according to any one of Examples 103 to 122, further comprising the steps of receiving vehicle sensor data from a vehicle sensor and generating object data based on the vehicle sensor data.

[0236] Example 124: The vehicle data association method described in Example 123, wherein the vehicle sensors include a steering sensor, an accelerometer, a braking sensor, a speedometer, or any combination thereof.

[0237] Example 125: The vehicle data association method according to any one of Examples 103 to 124, further comprising the steps of receiving actuator data and generating object data based on the vehicle actuator data.

[0238] Example 126: The vehicle data association method according to Example 125, wherein the vehicle actuator data includes data representing steering wheel position, brake position, brake pedal depression, braking force, speed, velocity, acceleration, or any combination thereof.

[0239] Example 127: The vehicle data association method according to any one of Examples 123 to 126, further comprising the step of determining the behavior of the vehicle related to an object represented by the object data from the sensor data and / or the vehicle actuator data.

[0240] Example 128: The vehicle data association method according to any one of Examples 103 to 127, further comprising the step of implementing any one or more elements described in Examples 103 to 127 into an artificial neural network.

[0241] Example 129: The vehicle data association method according to any one of Examples 103 to 128, further comprising the step of storing the first data, the second data, the third data, the object data, or any combination thereof in memory.

[0242] Example 130: A vehicle data association method according to any one of Examples 103 to 129, wherein at least two of the first data, second data, third data, object data, or any combination thereof each include a timestamp, and multiple data sources are synchronized via the timestamp.

[0243] Example 131: The vehicle data association method according to any one of Examples 103 to 130, further comprising the step of associating an object with the second data for a predetermined period of time following the audio indicator or the image indicator, and not associating the object with the second data after the end of the predetermined period of time following the audio indicator or the image indicator.

[0244] Example 132: The vehicle data association method according to any one of Examples 103 to 131, further comprising the step of determining the priority of the object based on the audio indicator or the image indicator.

[0245] Example 133: The vehicle data association method described in Example 132, wherein the priority is based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated amount of damage associated with a collision with the object, or any combination thereof.

[0246] Example 134: The vehicle data association method according to any one of Examples 103 to 133, wherein the object data includes labels for one or more detected objects.

[0247] Example 135: The vehicle data association method according to any one of Examples 103 to 134, further comprising the step of generating a sensor data label, wherein the sensor data label is a label representing at least one of the identity of the object, the behavior of the object, or the priority of the object.

[0248] Example 136: The vehicle data association method according to Example 135, further comprising the step of outputting sensor data including the sensor data label.

[0249] Example 137: The vehicle data association method according to Example 136 or 135, wherein the output sensor data includes the third data and the sensor data label.

[0250] While this disclosure is illustrated and described with reference to particular embodiments, those skilled in the art will understand that various modifications of form and detail can be made therein without departing from the spirit and scope of this disclosure as defined in the appended claims. Therefore, the scope of this disclosure is indicated in the appended claims and is intended to encompass all modifications within the meaning and equivalence of the claims. [Other adjacent items] (Item 1) A vehicle data association device, wherein the vehicle data association device is Internal audio / image data analysis device, External image analysis device, Object data generation device and Equipped with, The aforementioned internal audio / image data analysis device is Identifying a second data representing an audio indicator or an image indicator within a first data representing at least one of audio or images from inside the vehicle, wherein the audio indicator is a human voice and the image indicator is configured to identify a human action inside the vehicle. The external image analysis device is configured to identify an object corresponding to at least one of the audio indicator or the image indicator within third data representing an image of the vicinity of the outside of the vehicle. The object data generation device is configured to generate object data and classify the third data. Vehicle data association device. (Item 2) The vehicle data association device according to item 1, wherein the object data generation device is configured to classify the third data based on the relevance of the third data in training a trainable model. (Item 3) The vehicle data association device according to item 1, wherein the object data includes at least one of the object's identity, the object's behavior, or the object's priority. (Item 4) The vehicle data association device according to item 1, wherein the internal audio / image data analyzer identifies the second data, which includes the internal audio / image data analyzer identifying one or more keywords in the audio. (Item 5) The vehicle data association device according to item 4, wherein the internal audio / image data analysis device is configured to transmit data representing one or more keywords to the external image analysis device, and the external image analysis device is configured to identify the object based on the relationship between the object and the one or more keywords. (Item 6) The vehicle data association device according to item 4, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords. (Item 7) The vehicle data association device according to item 1, wherein the internal audio / image data analyzer identifies the second data, which includes the internal audio / image data analyzer identifying a human gesture in the image, the human gesture being at least one of pointing in a direction, making an attention gesture, making a denial gesture, making a stop gesture, or any combination thereof. (Item 8) The vehicle data association device according to item 7, wherein the internal audio / image data analysis device is configured to transmit data representing the human gesture to the external image analysis device, and the external image analysis device is configured to map a pointing action to the third data and to identify the object based on the mapped relationship between the pointing action and the object. (Item 9) The vehicle data association device according to item 1, wherein the internal audio / image data analysis device identifies the second data, which includes identifying the direction of human gaze in the image. (Item 10) The vehicle data association device according to item 9, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to identify the object based at least partially on the relationship between the object and the gaze direction. (Item 11) The vehicle data association device according to item 9, wherein the internal audio / image data analyzer is configured to transmit data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to map the gaze direction to the third data and to identify the object based on the mapped relationship between the gaze direction and the object. (Item 12) The vehicle data association device according to item 1, wherein the external image analysis device is configured to identify the object based on at least two of one or more keywords, one or more gestures, or gaze direction. (Item 13) The vehicle data association device further comprises a vehicle sensor data analyzer configured to receive vehicle sensor data from vehicle sensors, and the external image analyzer further comprises an external image analyzer configured to generate object data based on the vehicle sensor data, wherein the vehicle sensors include steering sensors, accelerometers, braking sensors, speedometers, or any combination thereof, as described in item 1. (Item 14) The vehicle data association device further comprises a vehicle actuator data analyzer configured to receive actuator data, and the external image analyzer further comprises an external image analyzer configured to generate object data based on the vehicle actuator data, wherein the vehicle actuator data includes data representing steering wheel position, brake position, brake depression, braking force, speed, velocity, acceleration, or any combination thereof, as described in item 1. (Item 15) The vehicle data association device according to item 14, wherein the sensor data analysis device and / or the vehicle actuator data analysis device is configured to determine the behavior of the vehicle related to an object represented by the object data from the sensor data and / or the vehicle actuator data. (Item 16) The vehicle data association device described in item 1, wherein the object data generation device is configured to determine the priority of the object based on the audio indicator or the image indicator, the priority being based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated amount of damage associated with a collision with the object, or any combination thereof. (Item 17) The object data includes labels for one or more detected objects, as described in item 1 of the vehicle data association device. (Item 18) The vehicle data association device according to item 1, wherein the object data generation device is further configured to generate sensor data labels, the sensor data labels being labels representing at least one of the identity of the object, the behavior of the object, or the priority of the object. (Item 19) A non-temporary computer-readable medium comprising instructions, wherein, when executed, the instructions are directed to one or more processors. Identifying, within first data representing at least one of audio from inside the vehicle or images from inside the vehicle, second data representing an audio indicator or an image indicator, wherein the audio indicator is a human voice and the image indicator represents human behavior inside the vehicle. In the third data representing an image of the vicinity of the exterior of the vehicle, an object corresponding to at least one of the audio indicator or the video indicator is identified. To generate object data and classify the aforementioned third data. A non-temporary computer-readable medium that enables the operation of [the process]. (Item 20) The identity of the object is a non-temporary computer-readable medium as described in item 19, which includes one or more coordinates defining the boundaries of the object. (Item 21) Identifying the second data includes identifying one or more keywords in the audio, as described in item 19 of the non-temporary computer-readable medium. (Item 22) The instruction is further configured to cause one or more processors to transmit data representing one or more keywords to the external image analyzer, and the external image analyzer is configured to identify the object based on the relationship between the object and the one or more keywords, as described in item 21, for a non-temporary computer-readable medium. (Item 23) A means for associating vehicle data, wherein the means for associating vehicle data is Internal audio / image data analysis device, External image analysis device, Object data generation device and Equipped with, The internal audio / image data analysis device is configured to identify, within first data representing at least one of audio from inside the vehicle or images from inside the vehicle, second data representing an audio indicator or an image indicator, wherein the audio indicator is human voice and the image indicator represents human behavior inside the vehicle. The external image analysis device is configured to identify an object corresponding to at least one of the audio indicator or the video indicator within third data representing an image of the vicinity of the outside of the vehicle. The external image analysis device is configured to generate object data and classify the third data. A method for associating vehicle data.

Claims

1. 1. A vehicle data association device, comprising: an internal audio / image data analyzer; an external image analyzer; an object data generating device; Equipped with said internal audio / image data analyzer identifying second data representing an audio indicator or a visual indicator within first data representing at least one of audio or visuals from within a vehicle, the audio indicator being a human voice and the visual indicator being representative of a human activity within the vehicle; the external image analyzer is configured to identify an object corresponding to at least one of the audio indicator or the visual indicator within third data representing an image near the exterior of the vehicle; the object data generation device is configured to generate object data based on the identified object and classify the third data for training a machine learning model implemented in a driver assistance system or an autonomous driving system that performs or controls a function of a vehicle based on an image of an exterior vicinity of the vehicle. Vehicle data association device.

2. The vehicle data association device of claim 1 , wherein the object data generator is configured to classify the third data based on the relevance of the third data in training a trainable model.

3. The vehicle data association device of claim 1 or 2, wherein the object data includes at least one of the object's identity, the object's behavior, or the object's priority.

4. 4. The vehicle data association device of claim 1, wherein the internal audio / image data analyzer identifying second data comprises the internal audio / image data analyzer identifying one or more keywords within the audio.

5. 5. The vehicle data association device of claim 4, wherein the internal audio / image data analyzer is configured to send data representing the one or more keywords to the external image analyzer, and the external image analyzer is configured to identify the object based on a relationship between the object and the one or more keywords.

6. 6. The vehicle data association device of claim 4 or 5, wherein the external image analysis device is configured to repeatedly identify the object using at least two keywords.

7. 7. The vehicle data association device of claim 1, wherein the internal audio / image data analyzer identifying the second data includes the internal audio / image data analyzer identifying a human gesture in the image, the human gesture being at least one of pointing in a direction, making an attention gesture, making a negation gesture, making a stop gesture, or any combination thereof.

8. 8. The vehicle data association device of claim 7, wherein the internal audio / image data analyzer is configured to send data representing the human gesture to the external image analyzer, the external image analyzer being configured to map a pointing behavior to the third data and identify the object based on the mapped relationship between the pointing behavior and the object.

9. 9. The vehicle data association device of claim 1, wherein the internal audio / image data analyzer identifying second data includes the internal audio / image data analyzer identifying a gaze direction of a person within the image.

10. 10. The vehicle data association device of claim 9, wherein the internal audio / image data analyzer is configured to send data representing the gaze direction to the external image analyzer, and the external image analyzer is configured to identify the object based at least in part on a relationship between the object and the gaze direction.

11. 11. The vehicle data association device of claim 9 or 10, wherein the internal audio / image data analyzer is configured to send data representing the gaze direction to the external image analyzer, the external image analyzer being configured to map the gaze direction to the third data and identify the object based on the mapped relationship between the gaze direction and the object.

12. 12. The vehicle data association device of claim 1, wherein the external image analyzer is configured to identify the object based on at least two of one or more keywords, one or more gestures, or gaze direction.

13. 13. The vehicle data association device of claim 1, further comprising a vehicle sensor data analyzer configured to receive vehicle sensor data from vehicle sensors, the external image analyzer further configured to generate object data based on the vehicle sensor data, the vehicle sensors including a steering sensor, an accelerometer, a braking sensor, a speedometer, or any combination thereof.

14. 14. The vehicle data association device of claim 13, further comprising a vehicle actuator data analyzer configured to receive actuator data, and wherein the external image analyzer is further configured to generate object data based on the vehicle actuator data, the vehicle actuator data including data representing steering wheel position, brake position, brake application, braking force, speed, velocity, acceleration, or any combination thereof.

15. 15. The vehicle data association device of claim 14, wherein the vehicle sensor data analyzer and / or the vehicle actuator data analyzer are configured to determine a behavior of the vehicle in relation to an object represented by the object data from the vehicle sensor data and / or the vehicle actuator data.

16. 16. The vehicle data association device of claim 1, wherein the object data generation unit is configured to determine a priority of the object based on the audio indicator or the image indicator, the priority being based on the importance of avoiding a collision with the object, the risk of a collision with the object, the estimated damage associated with a collision with the object, or any combination thereof.

17. 17. A vehicle data association device according to any preceding claim, wherein the object data comprises labels of one or more detected objects.

18. 18. The vehicle data association device of claim 1, wherein the object data generation unit is further configured to generate a sensor data label, the sensor data label being a label representing at least one of an identity of the object, a behavior of the object, or a priority of the object.

19. one or more processors, identifying, within first data representing at least one of audio from within a vehicle or an image from within the vehicle, second data representing an audio indicator or an image indicator, wherein the audio indicator is a human voice and the image indicator is representative of a human activity within the vehicle; identifying an object corresponding to at least one of the audio indicator or the visual indicator within third data representing an image near the exterior of the vehicle; generating object data based on the identified objects and classifying the third data for training a machine learning model implemented in a driver assistance or autonomous driving system that performs or controls a vehicle function based on images of the vehicle's exterior and vicinity; A program to execute.

20. 20. The program of claim 19, wherein the object identity includes one or more coordinates that define a boundary of the object.

21. 21. The program of claim 19 or 20, wherein identifying second data includes identifying one or more keywords within the audio.

22. 22. The program of claim 21, further causing the one or more processors to send data representing the one or more keywords to an external image analysis device, the external image analysis device configured to identify the object based on a relationship between the object and the one or more keywords.

23. A vehicle data association means, the vehicle data association means comprising: an internal audio / image data analyzer; an external image analyzer; an object data generating device; Equipped with the internal audio / image data analyzer is configured to identify, within first data representing at least one of audio from within a vehicle or images from within the vehicle, second data representing an audio indicator or an image indicator, the audio indicator being a human voice and the image indicator being indicative of a human activity within the vehicle; the external image analyzer is configured to identify an object corresponding to at least one of the audio indicator or the visual indicator within third data representing an image near the exterior of the vehicle; the object data generation device is configured to generate object data based on the identified object and classify the third data for training a machine learning model implemented in a driver assistance system or an autonomous driving system that performs or controls a function of a vehicle based on an image of an exterior vicinity of the vehicle. A means of vehicle data association.

24. A non-transitory computer-readable storage medium storing the program according to any one of claims 19 to 22.