Federated learning for connected camera applications in vehicles
By integrating sensors and federated learning technology into vehicles, external events can be automatically detected and reported, solving the problem of reliance on manual detection and reporting in existing technologies, and achieving automated event detection with real-time event processing and data privacy protection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-13
- Publication Date
- 2026-03-27
AI Technical Summary
Existing vehicle incident detection and reporting methods mainly rely on manual detection and reporting, which cannot automate the handling of external vehicle events in real time, especially when passengers cannot see the event and therefore cannot respond in a timely manner.
Sensors (such as cameras, radar, lidar, and microphones) are used to detect external events to the vehicle, and machine learning and federated learning techniques are combined to automatically record and report events, including the uploading of video and audio data and vehicle control actions.
It enables real-time automated detection and reporting of incidents, provides real-time video/audio evidence, assists vehicle operators and emergency services, and improves the accuracy of data processing and privacy protection through federated learning.
Smart Images

Figure CN116416788B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The technical field generally relates to vehicles, and more specifically to automated event detection for vehicles based on captured images using federated learning. BACKGROUND
[0002] Modern vehicles (e.g., cars, motorcycles, boats, or any other type of automobile) can be equipped with vehicle communication systems that facilitate different types of communication between the vehicle and other entities. For example, the vehicle communication systems can provide vehicle-to-infrastructure (V2I), vehicle-to-vehicle (V2V), vehicle-to-pedestrian (V2P), and / or vehicle-to-grid (V2G) communication. Collectively, these can be referred to as vehicle-to-anything (V2X) communication, which enables communication of information between a vehicle and any other suitable entity. Various applications (e.g., V2X applications) can use V2X communication to transmit and / or receive safety messages, maintenance messages, vehicle status messages, etc.
[0003] Modern vehicles can also include one or more cameras that provide backup assistance, take images of the vehicle’s driver to determine driver drowsiness or attentiveness, provide road images while the vehicle is in motion for collision avoidance, provide structure recognition, such as road signs, etc. For example, a vehicle can be equipped with multiple cameras, and can use images from multiple cameras (referred to as “wrap-around view cameras”) to create a “wrap-around” or “bird’s eye” view of the vehicle. Some cameras (referred to as “long-range cameras”) can be used to capture long-range images (e.g., for object detection for collision avoidance, structure recognition, etc.).
[0004] Such vehicles can also be equipped with sensors for performing target tracking, such as radar device(s), LiDAR (light detection and ranging) device(s), etc. Target tracking includes identifying a target object and tracking the target object over time as the target object moves relative to the vehicle observing the target object. Images from one or more cameras of the vehicle can also be used to perform target tracking.
[0005] These communication protocols, cameras, and / or sensors can be used to monitor the vehicle and the environment around the vehicle.
[0006] In practice, vehicles operate in different environments that exhibit regional variations, resulting in data diversity that can complicate labeling and reduce the rate of convergence of machine learning models. Accordingly, it is desirable to improve model performance in a manner that accounts for regional variations. Other desirable features and characteristics will become apparent from the subsequent detailed description and the appended claims, taken in conjunction with the accompanying drawings and the foregoing technical field and background. SUMMARY
[0007] A method for assigning an object type to a detected object using a location-considering local classification model is provided. The method involves obtaining, at a vehicle, sensor data for a detected object external to the vehicle from a sensor of the vehicle, obtaining, at the vehicle, location data associated with the detected object, obtaining, at the vehicle, a local classification model associated with an object type, assigning, at the vehicle, the object type to the detected object using the local classification model based on an output of the local classification model from the sensor data and the location data, and initiating, at the vehicle, an action responsive to assigning the object type to the detected object.
[0008] An apparatus for a vehicle is provided, the apparatus comprising a sensor to provide sensor data for a detected object external to the vehicle, a navigation system to provide location data for the vehicle contemporaneously with the sensor data, a memory comprising computer readable instructions and a local classification model associated with an object type, and a processing device to execute the computer readable instructions. The computer readable instructions control the processing device to perform operations involving assigning the object type to the detected object using the local classification model based on an output of the local classification model from the sensor data and the location data, and initiating an action responsive to assigning the object type to the detected object.
[0009] A vehicle system is provided, comprising a remote server and a vehicle. The remote server is configurable to provide a classification model for an object type. The vehicle is coupled to the remote server over a network to obtain the classification model from the remote server. The vehicle comprises a sensor to provide sensor data for a detected object external to the vehicle, a navigation system to provide location data for the vehicle contemporaneously with the sensor data, a memory comprising computer readable instructions and a processing device to execute the computer readable instructions. The computer readable instructions control the processing device to perform operations involving assigning the object type to the detected object using the first classification model based on an output of the classification model from the sensor data, and responsive to assigning the object type to the detected object, determining a local classification model for the object type using the sensor data and the location data associated with the detected object. BRIEF DESCRIPTION OF DRAWINGS
[0010] In the following exemplary embodiments will be described with reference to the following drawings, wherein like elements are referred to with like reference numerals, and wherein:
[0011] Figure 1 A vehicle comprising a sensor and a processing system according to one or more embodiments described herein is depicted;
[0012] Figure 2a block diagram depicting a system for automated event detection for a vehicle according to one or more embodiments described herein;
[0013] Figure 3 a flow diagram depicting a method for processing environmental data for a vehicle according to one or more embodiments described herein;
[0014] Figure 4 a block diagram depicting a federated learning system according to one or more embodiments described herein;
[0015] Figure 5 a flow diagram depicting a federated learning process adapted to be implemented in conjunction with one or more vehicles in a federated learning system according to one or more embodiments described herein; and
[0016] Figure 6 a flow diagram depicting an enhanced labeling process adapted to be implemented at one or more vehicles in a federated learning system according to one or more embodiments described herein. DETAILED DESCRIPTION
[0017] The following detailed description is merely exemplary in nature and is not intended to limit the application and use. Furthermore, there is no intention to be bound by any expressed or implied theory presented in the preceding technical field, background, brief summary or the following detailed description. As used herein, the term module refers to an application specific integrated circuit (ASIC), an electronic circuit, a processor (shared, dedicated, or group) and memory that execute one or more software or firmware programs, a combinational logic circuit, and / or other suitable components that provide the described functionality.
[0018] The subject matter described herein relates generally to automated event detection for a vehicle, e.g., for automated traffic event recording and reporting. One or more exemplary embodiments described herein provide for recording an event outside of a vehicle, e.g., a traffic event (e.g., a traffic stop by law enforcement, an accident, etc.), or any other event outside of the vehicle, which serves as a trigger and then takes action, e.g., controls the vehicle and / or reports the event to an emergency dispatcher, another vehicle, etc.
[0019] Conventional approaches to event detection and reporting for vehicles are inadequate. For example, event detection and reporting is largely a manual process requiring human detection and triggering of reporting. Consider the example of a law enforcement officer driving past a target vehicle for a traffic stop. In this case, the passenger of the target vehicle would have to manually detect that the target vehicle is being driven past and then manually initiate a recording of the traffic stop, such as on a mobile phone (e.g., a smartphone) or camera system within the target vehicle. If reporting of the event to, such as a family member, emergency response agency, etc., is desired, such reporting is typically performed manually, such as through a phone call. Moreover, if an event occurs with another vehicle other than the target vehicle, the passenger of the target vehicle can not be aware of the event (e.g., the passenger cannot see the event).
[0020] One or more embodiments described herein address these and other drawbacks of the prior art by detecting an event, initiating a recording using one or more sensors (e.g., a camera, a microphone, etc.), and reporting the event. As one example, a method in accordance with one or more embodiments can include detecting an event (e.g., detecting an emergency vehicle, a law enforcement vehicle, etc.), initiating a recording of audio and / or video, superimposing data (e.g., speed, location, timestamp, etc.) on the video, uploading the audio and / or video recording to a remote processing system (e.g., a cloud computing node of a cloud computing environment), and issuing an alert (also referred to as a “notification”). In some examples, the audio and / or video recording can be used to reconstruct the scene or event. In some examples, one or more vehicles involved in the event (e.g., the target vehicle, the law enforcement vehicle, etc.) can be controlled, such as by causing a window to roll down, causing a light to turn on, causing an alarm within one or more of the vehicles to be sounded, etc.
[0021] One or more embodiments described herein provide advantages and improvements over the prior art. For example, the described technical solution can provide video / audio of what is happening when an event occurs (e.g., when a vehicle operator is pulled over by law enforcement) and also provide evidence of the event in real-time, including alerting third parties of the event. Moreover, the described technical solution can provide real-time assistance to a vehicle operator during a traffic stop or accident by providing real-time video and / or audio. Other advantages of the technology can include behavioral adjustments from parties involved in the event to improve outcomes. Additional advantages include using data about a detected event to control a vehicle, such as causing the vehicle to steer away from an approaching emergency vehicle.
[0022] Figure 1 A vehicle 100 including a sensor and processing system 110 in accordance with one or more embodiments described herein is depicted. In the depicted example, the vehicle 100 includes a camera 102, a microphone 104, a processing system 110, and a communication system 112. The camera 102 can be configured to capture video of a scene in front of the vehicle 100. The microphone 104 can be configured to capture audio of the scene in front of the vehicle 100. The processing system 110 can be configured to detect an event (e.g., an emergency vehicle, a law enforcement vehicle, etc.), initiate a recording of audio and / or video, superimpose data (e.g., speed, location, timestamp, etc.) on the video, upload the audio and / or video recording to a remote processing system (e.g., a cloud computing node of a cloud computing environment), and issue an alert (also referred to as a “notification”). The communication system 112 can be configured to communicate with the remote processing system. Figure 1In the example of FIG. 1, vehicle 100 includes processing system 110, cameras 120, 121, 122, 123, cameras 130, 131, 132, 133, radar sensor 140, light detection and ranging (LiDAR) sensor 141, and microphone 142. Vehicle 100 can be a car, truck, van, bus, motorcycle, boat, airplane, or another suitable vehicle 100.
[0023] Cameras 120-123 are surround view cameras that capture images of the exterior and vicinity of vehicle 100. Together, the images captured by cameras 120-123 form a surround view (sometimes referred to as a “top-down view” or “bird’s eye view”) of vehicle 100. These images can be used to operate the vehicle (e.g., to park, back up, etc.). These images can also be used to capture events, such as a traffic stop, accident, etc. Cameras 130-133 are long-range cameras that capture images of the exterior of the vehicle and further away from vehicle 100 than cameras 120-123. For example, these images can be used for object detection and avoidance. These images can also be used to capture events, such as a traffic stop, accident, etc. It should be understood that although eight cameras 120-123 and 130-133 are shown, more or fewer cameras can be implemented in various embodiments.
[0024] The captured images can be displayed on a display (not shown) to provide an exterior view of vehicle 100 to a driver / operator of vehicle 100. The captured images can be displayed as live images, still images, or some combination thereof. In some examples, the images can be combined to form a composite view, such as a surround view. In some examples, the images captured by cameras 120, 123 and 130, 133 can be stored to data store 111 of processing system 110 and / or to remote data store 151 associated with remote processing system 150.
[0025] Radar sensor 140 measures distance to a target object by emitting electromagnetic waves and measuring the reflected waves with a sensor. This information is useful to determine the distance / location of the target object relative to vehicle 100. It should be understood that radar sensor 140 can represent multiple radar sensors.
[0026] Light detection and ranging (LiDAR) sensor 141 measures distance to a target object (e.g., other vehicle 154) by illuminating the target with a pulsed or continuous wave laser and measuring the reflected pulses or continuous wave with a detector sensor. This information is useful to determine the distance / location of the target object relative to vehicle 100. It should be understood that LiDAR sensor 141 can represent multiple LiDAR sensors.
[0027] Microphone 142 can record sound waves (e.g., sound or audio). This information can be useful to record sound information about the vehicle 100 and / or the environment near the vehicle 100. It should be understood that microphone 142 can represent multiple microphones and / or a microphone array, which can be disposed in or on the vehicle such that microphone 142 can record sound waves inside the vehicle (e.g., the passenger cabin) and / or outside the vehicle.
[0028] Data generated from cameras 120-123, 130-133, radar sensor 140, lidar sensor 141, and / or microphone 142 can be used to detect and / or track target objects relative to vehicle 100 to detect events, etc. Examples of target objects include other vehicles (e.g., other vehicle 154), emergency vehicles, vulnerable road users (VRUs) such as pedestrians, bicycles, animals, potholes, oil on the road surface, debris on the road surface, fog, flooding, etc.
[0029] Processing system 110 includes a data / communication engine 112, a decision engine 114 for detection and classification, a control engine 116, a data store 111, and a machine learning (ML) model 118. Data / communication engine 112 receives / collection data from sensors associated with vehicle 100 (e.g., one or more of cameras 120, 123, 130, 133, radar sensor 140, lidar sensor 141, microphone 142, etc.), and / or from other sources such as remote processing system 150 and / or other vehicle 154. Decision engine 114 processes the data to detect and classify events. Decision engine 114 can utilize ML model 118 in accordance with one or more embodiments described herein. Examples of how decision engine 114 processes data are shown in particular in Figure 2 and are further described herein. Control engine 116 controls vehicle 100 in order to plan a route and perform driving maneuvers (e.g., change lanes, change speed, etc.), initiate recording from sensors of vehicle 100, store recorded data in data store 111 of remote processing system 150 and / or data store 151, and perform other suitable actions. Processing system 110 can include other components, engines, modules, etc., such as processors (e.g., central processing units, graphics processing units, microprocessors, etc.), memory (e.g., random access memory, read-only memory, etc.), data storage (e.g., solid state drives, hard disk drives, mass storage, etc.), etc.
[0030] The processing system 110 can be communicatively coupled to a remote processing system 150, which can be an edge processing node that is part of an edge processing environment, a cloud processing node that is part of a cloud processing environment, and the like. The processing system 110 can also be communicatively coupled to one or more other vehicles (e.g., other vehicles 154). In some examples, the processing system 110 is communicatively coupled to the processing system 150 and / or another vehicle 154 directly (e.g., using V2V communication), while in other examples, the processing system 110 is communicatively coupled to the processing system 150 and / or another vehicle 154 indirectly, such as through a network 152. For example, the processing system 110 can include a network adapter that enables the processing system 110 to send data to and / or receive data from other sources, such as other processing systems including the remote processing system 150 and other vehicles 154, data repositories, and the like. As an example, the processing system 110 can send data to and / or receive data from the remote processing system 150 directly and / or via the network 152.
[0031] The network 152 represents any one or combination of different types of suitable communication networks, such as, for example, cable networks, public networks (e.g., the Internet), private networks, wireless networks, cellular networks, or any other suitable private and / or public networks. Moreover, the network 152 can have any suitable communication range associated therewith and can include, for example, global networks (e.g., the Internet), metropolitan area networks (MANs), wide area networks (WANs), local area networks (LANs), personal area networks (PANs). Additionally, the network 152 can include any type of media that can carry a network traffic, including, but not limited to, coaxial cable, twisted-pair wire, optical fiber, hybrid fiber coaxial (HFC) media, microwave terrestrial transceivers, radio frequency communication media, satellite communication media, or any combination thereof. In accordance with one or more embodiments described herein, the remote processing system 150, the other vehicles 154, and the processing system 110 communicate via vehicle-to-infrastructure (V2I), vehicle-to-vehicle (V2V), vehicle-to-pedestrian (V2P), and / or vehicle-to-grid (V2G) communication.
[0032] Features and functionality of the components of the processing system 110 are now described in more detail. The processing system 110 of the vehicle 100 assists in automated event detection of the vehicle.
[0033] According to one or more embodiments described herein, the processing system 110 combines sensor inputs and driving behavior with artificial intelligence (AI) and machine learning (ML) (e.g., federated learning) to determine when the vehicle 100 is involved in an incident (e.g., a traffic stop) and to automatically take actions such as using sensors associated with the vehicle to record data, connect to third parties, control the vehicle (e.g., roll down windows, turn on hazard lights and / or interior lights, etc.), add overlay information to recorded data (e.g., speed, GPS, and time stamps to recorded video), and the like. The processing system 110 can also issue notifications / alerts such as providing a message on a display of the vehicle 100 to the operator / passenger communicating that an incident has occurred, notifying emergency contacts and / or emergency dispatchers of the incident.
[0034] According to one or more embodiments described herein, the processing system 110 can perform automatic AI / ML triggering of features based on fusion of sensor data (e.g., data from cameras, microphones, etc.) with driving behavior observed by external vehicle sensors (e.g., one or more of the cameras 120, 123, 130, 133, microphones 142, etc.). The processing system 110 can trigger AI / ML (e.g., through enhanced federated learning) in conjunction with triggers to initiate recording such as emergency lights, sirens, speed, and vehicle harsh maneuvers. In some examples, the processing system 110 can cause local and / or remote data capture / recording to ensure data ownership / security. For example, raw data is saved locally in the vehicle 100 (e.g., in the data storage 111) and in a mobile device (not shown) of the operator / passenger of the vehicle 100. In addition, federated learning data and 3D reconstruction primitives can be uploaded to third parties such as the remote processing system 150 and / or another vehicle 154.
[0035] In some examples, processing system 110 can also implement third-party data collection and / or notification, and can provide multi-vehicle observation / processing. For example, processing system 110 can send an alert to an emergency dispatch service (e.g., remote processing system 150) to initiate emergency operations. In some examples, this can include sending data collected by vehicle sensors to remote processing system 150. Processing system 110 can also send an alert to another vehicle 154 to cause the other vehicle (e.g., a third-party witness) to collect data using one or more sensors (not shown) associated with the other vehicle 154. In some examples, processing system 110 can access data from another vehicle 154 (which can represent one or more vehicles) to perform 3D scene reconstruction through federated learning (or other suitable machine learning techniques). For example, multi-view / camera scene reconstruction techniques can be implemented using video collected from vehicle 100 (as well as one or more neighboring vehicles (e.g., another vehicle 154)) for 3D scene modeling, and audio associated with the video can be preserved or enhanced through noise cancellation techniques. When data from neighboring vehicles (e.g., another vehicle 154) near vehicle 100 is processed for 3D scene reconstruction in vehicle 100, an edge processing node, or a cloud computing node (e.g., using motion stereo or shape from motion), vehicle 100 sends relevant data / models using machine learning or federated learning methods to protect data privacy. That is, data that is not deemed irrelevant (e.g., data collected from before the event, data collected from the passenger cabin of neighboring vehicles, etc.) is not sent for scene reconstruction to provide data privacy.
[0036] According to one or more embodiments described herein, processing system 110 can prepare vehicle 100, such as by rolling up / down the windows, turning on hazard lights, turning on interior lights, providing a message on a display of the vehicle (e.g., “Emergency vehicle behind”), etc., upon detecting a law enforcement vehicle.
[0037] Turning now to Figure 2 , a block diagram of a system 200 for automated event detection for a vehicle is provided according to one or more embodiments described herein. In this example, system 200 includes vehicle 100, remote processing system 150, and another vehicle 154 communicatively coupled over network 152. It should be appreciated that another vehicle 154 can be similarly configured as vehicle 100 as shown and described herein. In some examples, additional other vehicles can be implemented. Figure 1
[0038] Figure 1 As shown, the vehicle 100 includes a sensor 202 (e.g., one or more of the cameras 120, 123, 130, 133, radar sensors 140, lidar sensors 141, microphones 142, etc.). At block 204, the data / communication engine 112 receives data from the sensor 202. This can include actively collecting data (e.g., causing the sensor 202 to collect data) or passively receiving data from the sensor 202.
[0039] The decision engine 114 processes the data collected by the data / communication engine 112 at block 204. In particular, at block 206, the decision engine 114 uses the data received / collection at block 204 to monitor the sensor 202 for indications of an event. According to one or more embodiments described herein, the decision engine 114 can utilize artificial intelligence (e.g., machine learning) to detect features within the sensor data (e.g., captured images, recorded sound waves, etc.) that are indicative of an event. For example, features commonly associated with emergency vehicles can be detected, such as flashing lights, sirens, markings / symbols on the vehicle, etc.
[0040] More specifically, aspects of the present disclosure can utilize machine learning functionality to accomplish various operations described herein. More specifically, one or more embodiments described herein can accomplish various operations described herein in conjunction with and utilizing rule-based decision making and artificial intelligence (AI) inference. The phrase “machine learning” broadly describes the functionality of an electronic system that learns from data. A machine learning system, module, or engine (e.g., the decision engine 114) can include a trainable machine learning algorithm that can be trained, for example, in an external environment (e.g., an edge processing node, a cloud processing node, etc.) to learn a functional relationship between inputs and outputs that are currently unknown, and the resulting model (e.g., the ML model 118) can be used to determine whether an event has occurred. In one or more embodiments, the machine learning functionality can be implemented using an artificial neural network (ANN) that has the ability to be trained to perform a currently unknown function. In machine learning and cognitive science, an ANN is a family of statistical learning models inspired by the biological neural networks of animals, particularly the brain. ANNs can be used to estimate or approximate systems and functions that depend on a large number of inputs.
[0041] ANNs can be embodied as so-called "neuromorphic" systems of interconnected processor elements that act as simulated "neurons" and exchange "messages" between one another in the form of electronic signals. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that carry messages between biological neurons, connections in an ANN that carry electronic messages between simulated neurons are provided with numerical weights that correspond to the strength or weakness of a given connection. The weights can be adjusted and tuned based on experience, adapting the ANN to the input and enabling learning. For example, an ANN for handwriting recognition is defined by a set of input neurons that can be activated by the pixels of an input image. After being weighted and transformed by a function determined by the network designer, the activations of these input neurons are then passed to other downstream neurons, which are often referred to as "hidden" neurons. The process is repeated until output neurons are activated. The activated output neurons determine which character was read. Similarly, the decision engine 114 can utilize the ML model 118 to detect an event. For example, the decision engine 114 can use image recognition techniques to detect an emergency vehicle in an image captured by the camera 120, can use audio processing techniques to detect a siren of an emergency vehicle in sound waves captured by the microphone 142, etc.
[0042] At decision block 208, it is determined whether an event was detected at block 206. If it is determined at decision block 208 that an event has not occurred, the decision engine 114 continues to monitor the sensors 202 for indications of an event.
[0043] However, if it is determined at decision block 208 that an event has occurred, the control engine 116 initiates recording / storage of data from the sensors 202 at block 210. This can include storing previously captured data and / or causing future data to be captured and stored. The data can be stored locally, such as in the data store 111, and / or remotely, such as in the data store 151 of the remote processing system 150 or another suitable system or device. The control engine 116 can also take action at block 212 and / or issue a notification at block 214 in response to the decision engine 114 detecting the event. Examples of actions that can be taken at block 212 include, but are not limited to: controlling the vehicle 100 (e.g., causing the vehicle 100 to perform a driving maneuver, such as changing lanes, changing speed, etc.; causing the vehicle 100 to turn on one or more of its lights; causing the vehicle 100 to lower / raise one or more of its windows, etc.), causing the recorded data to be modified (e.g., overlaying GPS data, speed / rate data, location data, timestamps, etc. on the recorded video; combining recorded sound waves and recorded video, etc.), and other suitable actions. Examples of notifications that can be issued at block 214 can include, but are not limited to: presenting an audio and / or visual cue to an operator or passenger of the vehicle 100 (e.g., presenting a warning message on a display within the vehicle, playing a warning tone within the vehicle, etc.), alerting a third-party service (e.g., an emergency dispatch service, a known contact of the operator or passenger of the vehicle, etc.), sending an alert to another vehicle 154 and / or the remote processing system 150. The type of action taken and / or the type of notification issued can be based on one or more of: a user preference; the type of event detected; a geographically-based law, regulation, or custom; etc.
[0044] Figure 3 A flowchart of a method 300 for automated event detection for a vehicle, in accordance with one or more embodiments described herein, is depicted. The method 300 can be performed by any suitable system or device, such as the processing system 110 of the vehicle 100, or any other suitable processing system and / or processing device (e.g., a processor). Figure 1 and Figure 2 The method 300 is now described with reference to the elements of the vehicle 100, but is not limited thereto. Figure 1 and / or Figure 2 The method 300 is now described with reference to the elements of the vehicle 100, but is not limited thereto.
[0045] At block 302, the processing system 110 receives first data from a sensor of the vehicle (e.g., one or more of the video cameras 120, 123, 130, 133, the radar sensor 140, the lidar sensor 141, the microphone 142, etc.).
[0046] At block 304, the processing system 110 determines whether an event outside the vehicle has occurred by processing the first data using a machine learning model. For example, as described herein, the decision engine 114 processes data collected by the sensors. In particular, the decision engine 114 monitors the sensors 202 using the data received / collection at block 204 to discover indications of events. According to one or more embodiments described herein, the decision engine 114 can utilize artificial intelligence (e.g., machine learning) to detect features within the sensor data (e.g., captured images, recorded sound waves, etc.) that are indicative of an event. For example, using the data received / collection at block 204, the decision engine 114 can detect the presence of an emergency vehicle or other first responder vehicle, such as an ambulance, fire truck, law enforcement vehicle, etc.
[0047] At block 306, the processing system 110 initiates recording of second data by the sensors in response to determining that an event outside the vehicle has occurred. For example, the processing system 110 initiates recording of video in response to detecting an event based on recorded audio.
[0048] At block 308, the processing system 110 takes an action to control the vehicle in response to determining that an event outside the vehicle has occurred. Taking an action includes the processing system 110 causing another system, device, component, etc. to take an action. In some examples, the processing system 110 controls the vehicle 100, such as performing a driving maneuver (e.g., changing lanes, changing speed, etc.), initiating recording from sensors of the vehicle 100, causing recorded data to be stored in the data store 111 and / or the data store 151 of the remote processing system 150, and performing other suitable actions.
[0049] Additional processes can also be included, and it should be understood that the processes depicted are representative, Figure 3 The processes depicted in the figures represent illustrative stages that can be added, omitted, and / or modified, and the order of stages can be changed, and additional stages can be added, without departing from the scope and spirit of the disclosure.
[0050] Figure 4 An example federated learning system 400 is depicted that includes a central node 402 coupled to any number of different edge nodes 404, 406 through one or more networks (e.g., network 152). It should be understood that although a simplified embodiment with two edge nodes 404, 406 is depicted for purposes of explanation, in practice, the federated learning system 400 can include any number of different edge nodes, and the subject matter described herein is not limited to any particular number of edge nodes. Figure 4 An example federated learning system 400 is depicted that includes a central node 402 coupled to any number of different edge nodes 404, 406 through one or more networks (e.g., network 152). It should be understood that although a simplified embodiment with two edge nodes 404, 406 is depicted for purposes of explanation, in practice, the federated learning system 400 can include any number of different edge nodes, and the subject matter described herein is not limited to any particular number of edge nodes.
[0051] The central node 402 generally represents a central server, a remote server, or any other type of remote processing system (e.g., the remote processing system 150) capable of communicating with the edge nodes 404, 406 over a network. The central node 402 includes a processing system 410, which can be implemented using any kind of processor, controller, central processing unit, graphics processing unit, microprocessor, microcontroller, and / or combinations thereof, which is suitably configured to determine one or more classification models 414 that are general, global, or otherwise location-agnostic, and to update or otherwise adapt the global classification models 414 using federated learning, as described in greater detail below. The central node 402 also includes a data storage element 412, which can be implemented as any kind of memory (e.g., random access memory, read-only memory, etc.), data storage (e.g., solid state drive, hard disk drive, mass storage, etc.), database, or the like, coupled to the processing system 410 and configured to store or otherwise maintain the classification models 414 determined by the processing system 410 for distribution to the edge nodes 404, 406.
[0052] Similar to the central node 402, in example embodiments, the edge nodes 404, 406 include respective processing systems 420, 430 and data storage elements 422, 432, which are suitably configured to receive the global classification model(s) 414 from the central node 402 and store or otherwise maintain local instances of the global classification model(s) 414 at the respective edge nodes 404, 406. As described in greater detail below, the edge nodes 404, 406 also store or otherwise maintain one or more local classification models 424, 434 that are locally determined at the respective edge nodes 404, 406 by training, adapting, or otherwise updating an initial global classification model 414 using respective data sets 426, 436 that are locally available at the respective edge nodes 404, 406. The local data 426 available at the first edge node 404 is different or otherwise distinct from the local data 436 available at the second edge node 406, and thus, the resulting locally adapted classification model 424 trained or otherwise determined at the first edge node 404 can be different from the locally adapted classification model 434 trained or otherwise determined at the second edge node 406.
[0053] In one or more embodiments, federated learning system 400 is implemented using a vehicle communication system, where edge nodes 404, 406 are implemented as different instances of vehicles 100, and local data 426, 436 includes image data, sensor data, radar data, object data, and / or any other kind of data generated, sensed, or otherwise collected from cameras 120-123, 130-133, radar sensors 140, lidar sensors 141, microphones 142, and / or other components located on or otherwise associated with respective vehicles 100. Additionally, in example embodiments, local data 426, 436 used to train, update, or otherwise determine locally adapted classification model(s) 424, 434 includes location data obtained from a global positioning system (GPS) or other navigation system associated with respective vehicles 100. In this regard, locally adapted versions of classification model(s) 424, 434 can be configured to receive location data characterizing a current location of respective edge nodes 404, 406 as one or more input parameters, which in turn results in local classification model(s) 424, 434 generating or otherwise providing location-dependent outputs as a function of input location data associated with respective sensor data input to respective models 424, 434.
[0054] Still referring to Figure 4In example embodiments, the edge nodes 404, 406 send or otherwise provide information related to respective local classification model(s) 424, 434 to the central node 402, which in turn utilizes the information related to the different local classification model(s) 424, 434 to train, update, or otherwise adapt the corresponding general, location-agnostic classification model 414 maintained at the central node 402 according to federated learning principles. For example, the edge nodes 404, 406 can be configured to determine the local classification model(s) 424, 434 using principal component analysis (PCA) or another machine learning technique suitable for use with federated learning, and then send or otherwise provide one or more statistics, performance metrics, or other parameters indicative of the respective local classification model 424, 434 and / or the respective underlying local data 426, 436 associated with the respective local classification model 424, 434 to the central node 402. Thereafter, according to federated learning principles, the central node 402 aggregates or otherwise combines the statistics, performance metrics, or other parameters indicative of updates to the respective local classification model(s) 424, 434, and then updates the general location-agnostic classification model(s) 414 according to the received statistics, performance metrics, or other parameters to better reflect the respective local classification model 424, 434 and / or the respective underlying local data 426, 436 associated with the respective local classification model 424, 434 distributed across the different edge nodes 404, 406. Thereafter, the central node 402 can push, broadcast, or otherwise send the updated general, location-agnostic classification model(s) 414 reflecting federated learning derived from the different local classification models to the edge nodes 404, 406 for future use.
[0055] Figure 5 An example embodiment of a federated learning process 500 suitable for implementation with one or more vehicles is depicted in accordance with one or more embodiments described herein. It should be appreciated that the order of operations within the federated learning process 500 is not limited to the sequential execution as shown, but can be executed in one or more varied orders as applicable and in accordance with the present disclosure. Moreover, one or more of the tasks shown and described in the context of Figure 5 the federated learning process 500 can be omitted from an actual embodiment of the federated learning process 500 while still achieving the overall functionality generally contemplated. The following description can reference elements mentioned above in connection with Figure 5 and FIG. 4 for illustrative purposes. In this regard, while portions of the federated learning process 500 can be performed by different elements of a vehicle system, such as the Figures 1-2 and Figure 1 and Figure 2the processing system 110 or any other suitable processing system) performs, but for purposes of explanation, the subject matter can be primarily described herein in the context of the federated learning process 500 being primarily performed by or supported in coordination with the edge nodes 404, 406 and the center node 402.
[0056] Still referring to Figure 5 The federated learning process 500 is initialized or otherwise started by identifying, receiving, or otherwise obtaining at 502 a local sensor dataset at an edge node that is classified, labeled, or otherwise assigned a particular object type using an object classification model, and identifying, receiving, or otherwise obtaining at 504 a location dataset associated with the labeled local dataset at the edge node. For example, in one or more embodiments, the processing system 420 at the edge node 404 (e.g., the processing system 110 of the vehicle 100) identifies or otherwise obtains a time-stamped sequence of samples of sensor data 426 from one or more on-board sensing devices 120-123, 130-133, 140, 141, 142 that are detected, classified, or otherwise identified as being associated with an emergency vehicle using a generic, location-agnostic emergency vehicle classification model 414 that was initially provided by the center node 402. Additionally, the processing system 420 identifies or otherwise obtains location data that is contemporaneous with the local sensor data 426 from a GPS or other navigation system associated with the edge node 100, 404.
[0057] At 506, after obtaining the labeled sensor dataset and corresponding location information associated with the detected object classified as the particular object type using the corresponding classification model, the federated learning process 500 determines, from the sensor data and corresponding location data, a local classification model associated with the particular object type for classifying or otherwise labeling future detected objects as the object type. For example, in one embodiment, after classifying a detected object in the vicinity of the vehicle 100, 404 as an emergency vehicle using the initial general emergency vehicle classification model 414 provided by the central node 402, the on-board processing system 110, 420 on the vehicle 100, 404 determines a local emergency vehicle classification model 424 using the labeled sensor data from the on-board sensing devices 120-123, 130-133, 140, 141, 142 and contemporaneous location information. In some embodiments, the labeled sensor data from the on-board sensing devices 120-123, 130-133, 140, 141, 142 and contemporaneous location information associated therewith is used as part of a training dataset to train, learn, or otherwise develop the local emergency vehicle classification model 424 for classifying detected objects as emergency vehicles from sensor data associated with the detected object and contemporaneous location information associated with the detected object. That is, in other embodiments, the labeled sensor data from the on-board sensing devices 120-123, 130-133, 140, 141, 142 and contemporaneous location information can be used to adjust or otherwise update the general emergency vehicle classification model 414 to obtain the local emergency vehicle classification model 424 from the initial general emergency vehicle classification model 414 (e.g., by retraining or updating the initial general emergency vehicle classification model 414 using the obtained data as training data).
[0058] It should be noted that steps 502, 504, and 506 can be repeated in practice to iteratively refine and adapt the local classification models 424, 434 at the vehicles 100, 404, 406 to improve the accuracy and reliability of classification of each detected object at the vehicles 100, 404, 406 for labeled sensor data and co-located data. Thus, each time a particular type of object is detected, the local classification model 424, 434 associated with that particular type of object can be updated, adapted, or otherwise refined to improve subsequent classification. In this regard, by defining and adapting local classification models 424, 434 that are based on features but are also location dependent or otherwise weighted using location information, the federated learning process 500 addresses the data diversity challenge by accounting for local data diversity. For example, emergency vehicles in a particular city, state, or region can have a particular color, shape, or other characteristic that is different from the same type of emergency vehicle in another city, state, or region. Thus, by developing local classification models 424, 434 that classify sensor feature data associated with a detected object according to the current location of the owner vehicle 100, 404, 406 and / or the detected object, the local classification models 424, 434 can be used to classify detected objects exhibiting location-specific characteristic features with a higher degree of confidence or accuracy, as the local classification models 424, 434 account for the variability or diversity of data associated with characteristic features of the same type or class of object relative to location or across different geographic regions. Moreover, the local classification models 424, 434 can also evolve to account for changes in characteristic features associated with a particular type or class of object over time (e.g., fleet upgrades, different color schemes, etc.).
[0059] In one implementation, a local classification model is implemented using a support vector machine (SVM) trained using a location and time dependent sequence of sensor data associated with a detected object having a known classification derived from using a global classification model associated with the object type. In this regard, the detected object can be associated with a time dependent sequence of sensor feature data of the detected object and corresponding tracked location of the detected object and / or contemporaneous ego vehicle location suitable for use as an input object data set for a classification model. Each sample of sensor feature data can be input or otherwise provided to an instance of the global classification model 414 at the respective vehicle to classify the detected object with an estimated confidence level associated with the classification as a particular object type associated with the global classification model 414. When a number or percentage (e.g., greater than 20% of the sequence) of sensor feature data samples in the input object data set are assigned an estimated classification confidence greater than a threshold (e.g., greater than 80% probability or confidence), the object type label associated with the global classification model 414 is assigned to the detected object and each sensor feature data sample of the input object data set, thereby extrapolating or otherwise extending the object label across the entire data set associated with the common detected and tracked object. As described in greater detail below, in other implementations, when the ego vehicle is unable to classify or otherwise assign an object type label to a detected object with a desired level of confidence, the ego vehicle can query other nearby vehicles or other external systems or devices using the time and location data associated with the detected object to obtain an object type label assigned to the same object by a nearby vehicle or another external source. The crowdsourced label received from another vehicle or external source can then be assigned to each sensor feature data sample of the input object data set associated with the detected object.
[0060] Thereafter, the entire time dependent sequence of sensor feature data, location data, and classified object type label associated with the detected object can be used to train or otherwise develop a local classification model associated with the particular type of object using a SVM or other model configured to compute, estimate, or otherwise determine a classification confidence or probability of the particular object type from both input sensor feature data and associated location data. The updated support vectors defining classification boundaries for the local model update can also be stored or otherwise maintained for subsequent provision to the central node to support federated learning of the global classification model with respect to the particular object type. In addition to developing a local classification model, the entire time dependent labeled sequence of sensor feature data and location data can also be leveraged to support crowdsourcing of object classification and provide an index of object classification to other nearby vehicles as described herein.
[0061] Still referring to Figure 5 , the federated learning process 500 illustrated continues at 508 by sending or otherwise providing the index of the local classification model determined at 506 to the central node for updating the general location-agnostic classification model at 510. For example, the vehicles 100, 404, 406 can periodically push, upload, or otherwise send statistics, performance metrics, or other parameters indicative of the local classification model(s) 424, 434 to the central server 402, which then aggregates or otherwise combines the statistics, performance metrics, parameters, or other indices of the different local classification model(s) 424, 434 to update the global classification model(s) 414 in a manner influenced by the different local classification models 424, 434 without accessing or otherwise retrieving the local data 426. At the central server 402, the local classification model(s) 424, 434 are based on the classification model. In this regard, in some embodiments, the vehicles 100, 404, 406 can provide statistics or other metrics characterizing the underlying local data 426, 436 (e.g., sensor data and contemporaneous location data) on which the local classification model(s) 424, 434 are adapted to allow the processing system 410 at the central server 402 to update the global classification model(s) 414 in an appropriate manner while preserving privacy with respect to the local data 426, 436 at the vehicles 100, 404, 406.
[0062] For example, a first vehicle 100, 404 trains, learns, or otherwise develops a local ambulance classification model 424 for classifying detected objects as ambulances from detected object sensor data from on-board sensing devices 120-123, 130-133, 140, 141, 142 and contemporaneous location information of the vehicle 100, 404, while a second vehicle 100, 406 in a different city, state, or geographic region similarly develops a local ambulance classification model 434 specific to the locations in which the second vehicle 100, 406 has operated. Thereafter, the first vehicle 100, 404 can push, upload, or otherwise send to the central server 402 statistics, performance metrics, or other parameters indicative of its local ambulance classification model 424 and / or underlying local data 426 specific to the geographic region in which the first vehicle 100, 404 has operated, while the second vehicle 100, 406 can similarly provide to the central server 402 indicia of its local ambulance classification model 434 and / or underlying local data 436 specific to the geographic region in which the second vehicle 100, 406 has operated. Thereafter, the processing system 410 at the central server 402 aggregates or otherwise combines indicia of the local ambulance classification model 424 associated with the geographic region of operation of the first vehicle with the local ambulance classification model 434 associated with the geographic region of operation of the second vehicle and then updates the global location-independent ambulance classification model 414 in a manner influenced by the similarities of the different local classification models 424, 434 and / or the differences between the different local classification models 424, 434 and the previous global ambulance classification model 414. As a result, the updated global ambulance classification model 414 determined at the central server 402 can better account for data diversity and variation between different geographic regions.
[0063] In one or more implementations, when the classification model 414, 424, 434 is implemented using an SVM, support vectors defining the classification boundaries associated with the updates to the different local classification models 424, 434 are uploaded or otherwise provided to the central server 402. At the central server 402, the processing system 410 combines the updated support vectors received from the different vehicle edge nodes 404, 406 corresponding to the different local classification models 424, 434 with the support vectors associated with the global classification model 414 to retrain, update, or otherwise adapt the global classification model 414 using the updated support vectors from the vehicle edge nodes 404, 406. In this regard, by providing support vectors from different vehicle edge nodes 404, 406 operating in different, diverse, or varying geographic locations, the resulting global classification model 414 can be location-independent.
[0064] Still referring to Figure 5After updating the global classification model, the federated learning process 500 continues at 512 by downloading or otherwise obtaining the updated global classification model(s) at the edge nodes from the central node. For example, in some embodiments, the central server 402 can automatically push, broadcast, or otherwise send the updated global classification model that is generic or otherwise location-agnostic to the vehicles 100, 404, 406, which in turn overwrite existing copies of the generic, location-agnostic classification model 414 with the updated global classification model 414. In this regard, when the global classification model 414 achieves a higher level of confidence than the local classification model(s) 424, 434, the vehicles 100, 404, 406 can utilize the global classification model 414 to classify or otherwise assign object type labels to detected objects. For example, when the vehicles 100, 404, 406 are traveling from a home region that is well-adapted to the local emergency vehicle classification model 424, 434 that exhibits different characteristic features of emergency vehicles in different geographic regions, the global classification model 414 can detect, recognize, or otherwise assign an emergency vehicle object type to local sensor data 426, 436 with a higher level of confidence than the local emergency vehicle classification model 424, 434 that has not been trained using data specific to that new geographic region. Thus, because the federated learning process 500 enables the central node 402 to adapt the generic global classification model 414 in a manner that accounts for data diversity, object classification can be improved when the vehicles 100, 404, 406 are operating in new or unfamiliar geographic regions.
[0065] In practice, the federated learning process 500 can be repeated to iteratively update the generic, location-agnostic global classification model determined at the central node 402 over time using federated learning, while also adapting the local classification models to the specific geographic regions in which the respective vehicles 100, 404, 406 are located. As a result of the federated learning process 500, the resulting global classification model derived at the central node 402 can be more robust and achieve better performance by accounting for data diversity or variability exhibited across different geographic regions observed by the vehicle edge nodes 100, 404, 406. At the same time, the local classification models at the respective vehicles 100, 404, 406 are adapted to their respective geographic regions of operation, thereby improving performance locally at the vehicle edge nodes 100, 404, 406.
[0066] Figure 6 depicted in FIG. 6 is adapted to be implemented in conjunction with the federated learning process 500 of FIG. 5, in accordance with one or more embodiments described herein. Figure 5 An example embodiment of an enhanced labeling process 600 implemented with the federated learning process 500 of FIG. 5 with one or more vehicles, in accordance with one or more embodiments described herein. For example, reference is made to FIG. 5. Figures 1-5In an exemplary embodiment, an enhanced labeling process 600 is performed at the vehicle to increase the availability and quantity of the labeled sensor dataset available at the vehicle at step 502 of the federated learning process 500, for developing, training, or otherwise updating the local classification model at step 506. In this respect, the enhanced labeling process 600 addresses the problem of labeling sensor data samples of detected objects, where low confidence or inability to assign labels or classifications to those data samples would otherwise exist, for example, due to variations in the viewing angle of the detected objects and / or the inability to observe or identify the characteristic features of the detected objects from the perspective of the vehicle itself. In addition to the challenges posed by the data diversity and / or location-related variations of characteristic features as described above, the classification of detected objects at the vehicle may also be limited by the orientation of the detected objects relative to the vehicle, for example, by being able to perceive only the rear view of an emergency vehicle instead of multiple different views of the emergency vehicle (e.g., side views and / or front views).
[0067] It should be understood that the order of operations within the enhanced marking process 600 is not limited to, for example... Figure 6 The execution order shown may be modified to allow for one or more variations of the execution order as applicable, based on this disclosure. Furthermore, the execution order may be omitted from the actual embodiment of the enhanced marking process 600. Figure 6 The following description, for illustrative purposes, may refer to one or more of the tasks shown and described in the context of the above, while still achieving the overall functionality generally expected. For illustrative purposes, the following description may refer to the above in conjunction with... Figures 1-2 The elements mentioned in 4. In this regard, although parts of the enhanced labeling process 600 may be performed by different elements of the vehicle system, for illustrative purposes, the subject matter may be described primarily in the context of the enhanced labeling process 600, which is mainly performed or supported by or at the vehicle 100, which serves as edge nodes 404, 406 in the federated learning system 400.
[0068] The enhanced labeling process 600 is initialized or otherwise started by analyzing, at 602, sensor data obtained from one or more sensing devices associated with a vehicle to detect or otherwise identify a set or sequence of sensor data samples corresponding to the presence of an object in the vicinity of the vehicle, and then storing or otherwise maintaining, at 604, a sequence of sensor data samples of the detected object associated with location information and time information of the set of detected object data. For example, as described above, the processing system 110, 420 associated with the vehicle 100, 404 can include a decision engine 114 that analyzes local sensor data 426 generated from sensing devices 120-123, 130-133, 140, 141, 142 on the vehicle 100, 404 to detect and / or track objects relative to the vehicle 100, 404. After identifying a set or series of sensor data samples from one or more on-board sensing devices 120-123, 130-133, 140, 141, 142 associated with a detected and / or tracked object in the vicinity of the vehicle 100, 404, the processing system 110, 420 stores or otherwise maintains the sensor data samples associated with a time (or time stamp) and location information associated with a time (or time stamp) associated with the respective sensor data samples. For example, based on the time stamp associated with the sensor data samples and the time stamp associated with the location data samples obtained from a GPS or other on-board vehicle navigation system, the processing system 110, 420 can associate the detected object sensor data set with a simultaneous or current location of the vehicle 100, 404 at the time the object was detected. In practice, the location data maintained in association with the detected object sensor data set can include a current or contemporaneous location of the vehicle 100, 404 and a corresponding location of the detected object at the respective time, which is estimated, determined, or otherwise derived from the own vehicle location using the sensor data (e.g., based on a distance determined in a particular direction relative to the vehicle using lidar or radar).
[0069] At 606, the enhanced labeling process 600 analyzes the detected object sensor dataset to identify or otherwise determine whether at least one sample of the detected object’s sensor dataset can be classified, labeled, or otherwise assigned a particular object type using a classification model for detecting that object type. For example, the processing system 110, 420 can sequentially analyze the combination of sensor data samples and location information associated with respective timestamps to determine whether a particular view of the detected object (or time slice of object sensor data) can be classified or labeled with a desired level of confidence. In this regard, the processing system 110, 420 can input or otherwise provide the sensor data sample and contemporaneous location information associated with a first timestamp of the detected object sensor dataset to the local ambulance classification model 424 to determine whether the particular view of the detected object (or time slice of object sensor data) can be classified as an ambulance with an associated confidence level or other performance metric that is greater than a classification threshold. When the local ambulance classification model 424 is unable to classify the detected object with the desired level of confidence, the processing system 110, 420 continues to input or otherwise provide the sensor data sample associated with that timestamp to the global ambulance classification model 414 and / or other global or local classification models 414, 424 available at the vehicle 100, 404.
[0070] When the classification model 414, 424 is unable to classify or label the sensor data sample associated with a particular timestamp with a confidence level that is greater than a classification threshold, the processing system 110, 420 can sequentially analyze the combination of sensor data samples and location information associated with the next respective timestamp to determine whether the particular view of the detected object (or time slice of object sensor data) can be classified or labeled with a desired level of confidence until the entire sequence of detected object sensor data has been analyzed.
[0071] At 608, when one of the classification models is able to classify, label, or otherwise assign an object type to a sensor data sample with a desired level of confidence, the augmented labeling process 600 continues by labeling or otherwise assigning other sensor data samples associated with the same tracked and / or detected object with the assigned object type. For example, if a sequence of sensor data samples associated with a detected object includes a rear view and an angled view of the object in addition to a side view, and the processing system 110, 420 is only able to classify the sensor data sample corresponding to the side view as a particular object type (e.g., an ambulance) with a desired level of confidence, the processing system 110, 420 can extend the classification to assign the same object type to previous and / or subsequent data samples of the detected object sensor data set that are identified as belonging to the same object in the vicinity of the vehicle 100, 404 based on temporal and / or spatial relationships between the sensor data samples (e.g., when differences in location, spatial orientation, and / or characteristic features, as well as differences between timestamps, are less than respective object tracking thresholds indicative of the same object). In this way, the augmented labeling process 600 can label more diverse or varied views of a detected object, thereby increasing the size of the labeled data set used to train the local classification model (at 502), which in turn improves the performance and robustness of the local classification model (derived at 504).
[0072] Still referring to Figure 6When the enhanced labeling process 600 is unable to classify or otherwise assign an object type label to a detected object sensor data set, the enhanced labeling process 600 continues by broadcasting or otherwise transmitting a request for classification or labeling of the detected object to one or more other vehicles at 610. In this regard, the vehicle 100, 404 can utilize V2V communications to query other vehicles operating in the vicinity of the vehicle 100, 404 over the network 152 to indicate whether another vehicle operating in the vicinity of the vehicle 100, 404 is able to classify or label a detected object located at substantially the same location at the same time as the respective sensor data sample of the detected object sensor data set. For example, the vehicle 100, 404 can broadcast a request over the V2V communications network 152 including information identifying the timestamp and corresponding location associated with the detected object. A second vehicle 100, 406 operating in the vicinity of the requesting vehicle 100, 404 can receive the request and analyze detected object data stored or otherwise maintained at the second vehicle 100, 406 to determine whether the processing system 110, 430 on the second vehicle 100, 406 is able to classify or label a detected object at substantially the same location at substantially the same time. For example, the second vehicle 100, 406 can have been positioned to observe a side view of an ambulance, which allows the second vehicle 100, 406 to classify the detected object at that location as an ambulance, while the sensor data sample at the requesting vehicle 100, 404 only captures a rear view, an oblique view, a partial view, or otherwise does not allow the processing system 110, 420 at the requesting vehicle 100, 404 to classify the detected object as an ambulance with a desired level of confidence. In response to identifying a classified detected object that matches the location and time associated with the request, the second vehicle 100, 406 can transmit or otherwise provide a response to the query including an index of the object type assigned to the detected object at substantially the same time and location.
[0073] In response to receiving a response at 612 indicating that one or more other vehicles are able to classify the detected object, the augment labeling process 600 continues by labeling or otherwise assigning the crowd-sourced object type at 614 to all sensor data samples associated with the same tracked and / or detected object in a similar manner as described above at 608. In this regard, when a neighboring vehicle 100, 406 is able to classify a detected object as a particular object type (e.g., an ambulance) with a desired level of confidence, the processing system 110, 420 can extend that classification to assign the same object type (and associated level of confidence assigned by the neighboring vehicle 100, 406) to its own local sensor data samples corresponding to the same detected object in the vicinity of the vehicle 100, 404. In this regard, when multiple different vehicles respond to the query request, the processing system 110, 420 can combine or otherwise augment the responses to obtain an aggregate response before assigning the most likely object type to the detected object. For example, when multiple different vehicles assign the same object type to a detected object at substantially the same time and location, the processing system 110, 420 can increase the level of confidence associated with that classification to an aggregate level of confidence that is greater than the individual level of confidence from any one vehicle. On the other hand, when multiple different vehicles assign different object types to a detected object at substantially the same time and location, the processing system 110, 420 can implement one or more voting schemes to arbitrate between the classifications to determine the most likely object type to assign to the detected object as well as a corresponding adjusted level of confidence that accounts for the differences between the different vehicles. In the absence of a response from other vehicles, the augment labeling process 600 exits without classifying or labeling the detected object.
[0074] With reference to Figure 6 With reference to Figure 5It will be appreciated that the enhanced labeling process 600 and the federated learning process 500 can be implemented concurrently and work in conjunction with one another in order to improve object classification in an unsupervised manner. For example, the enhanced labeling process 600 can increase the amount of local data 426, 436 that can be labeled or classified at 502 and used to develop the local classification model 424, 434. The development of the improved local classification model 424, 434 at the vehicle 100, 404, 406 at 506 of the federated learning process 500, in turn, improves the ability of the processing system 110, 420, 430 at the vehicle 100, 404, 406 to classify individual data samples at 606 and assign labels to larger groups of sensor data samples associated with the same object at 608 of the enhanced labeling process 600, thereby increasing the available data at 502 and 504 and refining the local classification model 424. During subsequent iterations of the federated learning process 500 at 506, step 434 is performed. The local classification model 424, 434 also increases the likelihood that another vehicle operating substantially contemporaneously at or around substantially the same location is able to classify or label detected objects at 612 and provide the corresponding index at 610 in response to receiving a query request from the ego vehicle, which allows the ego vehicle to assign labels to detected object sensor data sets that would otherwise go unlabeled at 614. The available data at 502 is improved in order to improve future development of the local classification model at 506 in a manner that better accounts for different, challenging, or otherwise difficult to classify views of the object. Furthermore, the federated learning process 500 allows the global classification model 414 developed at the central node 402 at 510 to also reflect the improved local models that benefit from the enhanced labeling process 600 and are received at 508. The improvement of the global model provided back to the vehicle edge nodes at 512 also increases the likelihood that the vehicle edge nodes are able to use the local models in conjunction with the enhanced labeling process 600 to classify objects that would otherwise be unable to be classified with a desired level of confidence, which similarly improves the development of the local models and future iterations of the global model.
[0075] With reference to Figures 1-6With the federated learning process 500, the federated learning system 400 addresses the data diversity challenge by using local models that are adapted using local data associated with previous location information from the vehicle to provide feature-based classification or clustering, while also being weighted or dependent on location to improve automated event detection. Consistent with the federated learning process 500, the enhanced labeling process 600 addresses the labeling challenge in unsupervised learning by using object tracking and localization to correlate relevant sensor data samples and crowdsource to obtain classifications or labels assigned by other vehicles to address difficult or challenging classification scenarios with unknown detected objects. By improving the decision engine 114 with the ability to detect and classify emergency vehicles or other first responder vehicles with higher confidence, the processing system 110 of the vehicle 100 can more quickly and / or more accurately detect corresponding events outside the vehicle (e.g., at 304 of the method 300) and initiate appropriate actions in response to the detected events (e.g., 306 and / or 308 of the method 300).
[0076] While at least one example embodiment has been presented in the foregoing detailed description, it should be appreciated that a wide variety of modifications exist. It should also be appreciated that the example embodiment(s) are by way of example only and are not intended to limit the scope, applicability, or configuration of the disclosure in any way. Rather, the foregoing detailed description will provide those skilled in the art with a convenient road map for implementing an example embodiment. It should be understood that various changes can be made in the function and arrangement of elements without departing from the scope of the disclosure as set forth in the appended claims and the legal equivalents thereof.
Claims
1. A method comprising: obtaining, at a vehicle, sensor data for a detected object external to the vehicle from a sensor of the vehicle; obtaining, at the vehicle, location data associated with the detected object; obtaining, at the vehicle, a local classification model associated with an object type; assigning, at the vehicle, the object type to the detected object using the local classification model based on an output of the local classification model from the sensor data and the location data; and initiating, at the vehicle, an action responsive to assigning the object type to the detected object; wherein initiating the action comprises: obtaining, at a second time different from a first time associated with the sensor data, second sensor data from the sensor of the vehicle; obtaining, at the second time, second location data associated with the second sensor data, wherein the second location data is different from the location data; associating the second sensor data and the second location data with the detected object based on a first relationship between the second location data and the location data and a second relationship between the second time and the first time; and assigning the object type to the second sensor data and the second location data when a confidence value output by the local classification model for a combination of the sensor data and the location data is greater than a classification threshold.
2. The method of claim 1, wherein initiating the action comprises updating, at the vehicle, the local classification model using the sensor data and the location data associated with the sensor data. after updating the local classification model, transmitting, by the vehicle, a signature of the local classification model to a remote server over a network.
3. The method of claim 2, further comprising: after transmitting the signature of the local classification model to the remote server, receiving, at the vehicle, an updated global classification model associated with the object type from the remote server over the network, wherein the remote server determines the updated global classification model associated with the object type using the signature of the local classification model.
4. The method of claim 3, further comprising:
5. The method of claim 1, further comprising: when the confidence value output by the local classification model is less than a classification threshold, transmitting, by the vehicle, a request for a classification associated with the detected object over a network, wherein the request includes the location data associated with the detected object; and receiving, at the vehicle, an indication of the object type from a second vehicle responsive to the request over the network. assigning the object type indicated by the second vehicle for the detected object to the sensor data and the location data, wherein initiating the action comprises updating, at the vehicle, the local classification model using the sensor data and the location data associated with the sensor data.
6. The method of claim 5, further comprising:
7. The method of claim 5, further comprising: at the vehicle, obtaining second location data associated with the second sensor data at the second time, wherein the second location data is different from the location data; at the vehicle, obtaining second location data associated with the second sensor data at the second time, wherein the second location data is different from the location data; based on a first relationship between the second location data and the location data and a second relationship between the second time and the first time, associating the second sensor data and the second location data with the detected object; and after associating the second sensor data and the second location data with the detected object, assigning the object type indicated by the second vehicle for the detected object to the second sensor data and the second location data.
8. The method of claim 7, wherein, initiating the action includes: after assigning the object type indicated by the second vehicle for the detected object to the second sensor data and the second location data, updating, at the vehicle, the local classification model using the second sensor data and the second location data associated with the second sensor data.
9. A vehicle comprising: a sensor to provide sensor data of a detected object external to the vehicle; a navigation system to provide location data of the vehicle contemporaneously with the sensor data; a memory comprising computer readable instructions and a local classification model associated with an object type; and a processing device to execute the computer readable instructions, the computer readable instructions controlling the processing device to perform operations comprising: assigning, using the local classification model, the object type to the detected object based on an output of the local classification model from the sensor data and the location data; and initiating an action in response to assigning the object type to the detected object; wherein initiating the action includes: obtaining second sensor data from the sensor of the vehicle at a second time different from a first time associated with the sensor data; obtaining second location data associated with the second sensor data at the second time, wherein the second location data is different from the location data; based on a first relationship between the second location data and the location data and a second relationship between the second time and the first time, associating the second sensor data and the second location data with the detected object; and assigning the object type to the second sensor data and the second location data when a confidence value output by the local classification model for the combination of the sensor data and the location data is greater than a classification threshold.
Citation Information
Patent Citations
Embedded vehicle perception with machine learning classification of sensor data
CN110582778A