Method of and system for perceiving dynamic targets
The method and system enhance autonomous vehicle navigation by using multiple sensing modalities to detect and predict dynamic targets, addressing the limitations of existing navigation systems in managing moving obstacles.
Patent Information
- Application Number
- PCT/US2024/061360
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
Existing navigation systems struggle with managing dynamic obstacles, such as moving vehicles or living beings, due to limitations in sensing technology, data processing, and the cost-effectiveness of obstacle perception and motion analytics.
A method and system that utilize multiple sensing modalities with variable frequencies to detect, perceive, and predict the trajectories of dynamic targets, incorporating pseudo-synchronization of sensor data and track histories to enhance navigation.
Enables timely and cost-effective autonomous vehicle navigation by accurately tracking and predicting the motion of dynamic targets, improving the vehicle's ability to navigate through ordinary traffic.
Smart Images

Figure US2024061360_03072025_PF_FP_ABST
Abstract
Description
METHOD OF AND SYSTEM FOR PERCEIVING DYNAMIC TARGETSCROSS REFERENCE TO RELATED APPLICATION(S)
[0001] This application claims the benefit of U.S. Provisional Patent Application Serial No. 63 / 614,952 filed December 27, 2023, entitled Method of and System for Perceiving Dynamic Targets (Attorney Docket No. AB100), which is incorporated herein by reference in its entirety.FIELD OF THE INVENTION
[0002] The invention generally relates to navigation, and more particularly to perception of dynamic targets that may impact navigation.RESERVATION OF COPYRIGHTS
[0003] Portions of the disclosure of this document contain material that is subject to copyright protection. The copyright owner has no objection to any reproduction of the document or disclosure as it appears in official records, but reserves all remaining rights under copyright.BACKGROUND
[0004] Improvements to navigation of vehicles, not limited to self-driving or autonomous vehicles, have increased demand and potential for them. Onboard control systems are capable of making real time decisions independently without relying on human assistance. These decisions relate to determining paths to target locations and, to some extent, altering paths in response to encountered static obstacles. However, a persistent problem remains with managing dynamic obstacles, such as other vehicles or living beings that are moving nearby with undetermined velocity and trajectory.
[0005] Dynamic obstacle perception and motion analytics have been limited by the cost and performance limitations of equipment used for sensing dynamic targets, a lack of useable data respecting common targets, such as bicycles, and time-consuming processing of real-time data respecting sensed targets. However, such limits can be overcome with intelligent analytics and prioritized computing of data from sensors of diverse capabilities.
[0006] What are needed are a method of and a system for perceiving and estimating dynamic target velocity that are timely and cost effective to enable autonomous vehicle navigation in ordinary traffic.SUMMARY OF THE INVENTION
[0007] The invention overcomes the difficulties of dynamic obstacle perception and motion analytics with a method of and system for perceiving dynamic targets that are timely and cost effective to enable autonomous vehicle navigation in ordinary traffic. The method includes detecting targets, perceiving target locations and motions, predicting target trajectories and planning a travel route based on the trajectories.
[0008] Detecting preferably employs multiple sensing modalities with variable and unique frequencies.
[0009] Perceiving preferably includes collecting contemporaneous detections and “pseudo synchronizing” or combining them into a single message or muxed detection, into which is factored a confidence or reliability of the constituent detections. Perceiving preferably also includes tracking the muxed detections. The projected position of a target is compared with a track history thereof and the target heading corrected accordingly. Target headings are corrected based on a track history of position and velocity estimates.
[0010] Predicting includes developing trajectories of tracks for targets and the probabilities thereof.
[0011] Planning develops a path based on a trajectory from the predicted trajectories for navigating a vehicle.
[0012] The system includes more than one sensor configured for sensing local targets to which a processor is responsive.
[0013] The details of one or more embodiments are set forth in the accompanying drawings and the description below. Other features and advantages will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figs. 1 and 13 are diagrammatic views of a vehicle navigating among dynamic targets according to principles of the invention;
[0015] Figs. 2-8 are flow charts of processes configured according to principles of the invention;
[0016] Figs. 9-12 are diagrammatic views of target trajectory predictions; and
[0017] Fig 14 is a top, front, right side elevational view of an embodiment of a vehicle configured according to principles of the invention.
[0018] Like reference symbols in the various drawings indicate like elements.DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] The examples shown in drawings are presented to demonstrate examples of the disclosure. The drawings are illustrative and non-limiting. In the drawings, for illustrative purposes, the size of some of the elements may be exaggerated and not drawn to a particular scale. Additionally, elements shown within the drawings that have the same numbers may be identical elements or may be similar elements, depending on the context.
[0020] Where the term "comprising" is used in the present description and claims, it does not exclude other elements or steps. Where an indefinite or definite article is used when referring to a singular noun, e.g., "a", "an", or "the", this includes a plural of that noun unless something otherwise is specifically stated. Hence, the term "comprising" should not be interpreted as being restricted to the items listed thereafter; it does not exclude other elements or steps, and so the scope of the expression "a device comprising items A and B" should not be limited to devices consisting only of components A and B. Furthermore, to the extent that the terms “includes”, “has”, “possesses”, and the like are used in the present description and claims, such terms are intended to be inclusive in a manner similar to the term “comprising,” as “comprising” is interpreted when employed as a transitional word in a claim.
[0021] Furthermore, the terms "first", "second", "third", and the like, whether used in the description or in the claims, are provided to distinguish between similar elements and not necessarily to describe a sequential or chronological order. It is to be understood that the terms so used are interchangeable under appropriate circumstances (unless clearly disclosed otherwise) and that the aspects of the disclosure described herein are capable of operation in other sequences and / or arrangements than are described or illustrated herein.
[0022] In the following description, numerous specific details are set forth to provide a thorough understanding of various aspects and arrangements. It will be recognized, however, that the techniques described herein can be practiced without one or more of the specific details, or with other methods, components, materials, etc. In other instances, well known structures, materials, or operations may not be shown or described in detail to avoid obscuring certain aspects.
[0023] Reference throughout this specification to “an aspect,” “an arrangement,” “a configuration,” or “an example” indicates that a particular feature, structure, or characteristic is described. Thus, appearances of phrases such as “in one aspect,” “in one arrangement,” “in a configuration,” “in some examples,” or the like in various places throughout this specification do not necessarily each refer to the same aspect, feature, configuration, example, or arrangement. Furthermore, the particular features, structures, and / or characteristics described may be combined in any suitable manner.
[0024] To the extent used in the present disclosure and claims, the terms “component,” “system,” “platform,” “layer,” “selector,” “interface,” and the like are intended to refer to a computer- related entity or an entity related to an operational apparatus with one or more specific functionalities, wherein the entity may be either hardware, a combination of hardware and software, software, or software in execution. As an example, a component may be, but is not limited to being, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration and not limitation, both an application running on a server and the server itself can be a component. One or more components may reside within a process and / or thread of execution and a component may be localized on one computer and / or distributed between two or more computers. In addition, components may execute from various computer-readable media, device-readable storage devices, or machine-readable media having various data structures stored thereon. The components may communicate via local and / or remote processes such as in accordance with a signal having one or more data packets (e.g., data from one component interacting with another component in a local system, a distributed system, and / or across a network such as the Internet with other systems via the signal). As another example, a component can be an apparatus with specific functionality provided by mechanical parts operated by electric or electronic circuitry, which may be operated by a software or firmware application executed by a processor, wherein the processor can be internal or external to the apparatus and executes at least a part of the software or firmware application. As yet another example, a component can be an apparatus that provides specific functionality through electronic components without mechanical parts; the electronic components can include a processor therein to execute software or firmware that confers at least in part the functionality of the electronic components.
[0025] To the extent used in the subject specification, terms such as “store,” “storage,” “data store,” data storage,” “database,” and the like refer to memory components, entities embodied in a memory, or components comprising a memory. It will be appreciated that the memory components described herein can be either volatile memory or nonvolatile memory, or can include both volatile and nonvolatile memory.
[0026] In addition, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless specified otherwise, or clear from context, “X employs A or B” is intended to mean any of the natural inclusive permutations. That is, if X employs A, X employs B, or X employs both A and B, then “X employs A or B” is satisfied under any of the foregoing instances. Moreover, articles “a” and “an” as used in the subject disclosure and claims should generally be construed to mean “one or more” unless specified otherwise or clear from context to be directed to a singular form.
[0027] The words “exemplary” and / or “demonstrative,” to the extent used herein, mean serving as an example, instance, or illustration. For the avoidance of doubt, the subject matter disclosed herein is not limited by disclosed examples. In addition, any aspect or design described herein as “exemplary” and / or “demonstrative” is not necessarily to be construed as preferred or advantageous over other aspects or designs, nor is it meant to preclude equivalent exemplary structures and techniques. Furthermore, to the extent that the terms “includes,” “has,” “contains,” and other similar words are used in either the detailed description or the claims, such terms are intended to be inclusive, in a manner similar to the term “comprising” as an open transition word, without precluding any additional or other elements.
[0028] As used herein, the term “infer” or “inference” refers generally to the process of reasoning about, or inferring states of, the system, environment, user, and / or intent from a set of observations as captured via events and / or data. Captured data and events can include user data, device data, environment data, data from sensors, application data, implicit data, explicit data, etc. Inference can be employed to identify a specific context or action or can generate a probability distribution over states of interest based on a consideration of data and events, for example.
[0029] The disclosed subject matter can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques to produce software, firmware, hardware, or any combination thereof to control a computer to implement the disclosed subject matter. The term "article of manufacture," to the extent used herein, is intended to encompass a computer program accessible from any computer-readable device, machine-readable device, computer-readable carrier, computer-readable media, or machine- readable media. For example, computer-readable media can include, but are not limited to, a magnetic storage device, e.g., hard disk; floppy disk; magnetic strip(s); an optical disk (e.g., compact disk (CD), digital video disc (DVD), Blu-ray Disc (BD)) ; a smart card; a flash memory device (e.g., card, stick, key drive); a virtual device that emulates a storage device; and / or any combination of the above computer-readable media.
[0030] Generally, program modules include routines, programs, components, data structures, etc., that perform particular tasks or implement particular abstract data types. The illustrated aspects of the subject disclosure may be practiced in distributed computing environments where certain tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules can be located in both local and remote memory storage devices.
[0031] Computing devices can include at least computer-readable storage media, machine- readable storage media, and / or communications media. Computer-readable storage media or machine-readable storage media can be any available storage media that can be accessedby the computer and includes both volatile and nonvolatile media, removable and non-remov- able media. By way of example, and not limitation, computer-readable storage media or machine-readable storage media can be implemented in connection with any method or technology for storage of information such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data.
[0032] Computer-readable storage media can include, but are not limited to, random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other memory technology, compact disk read only memory (CD ROM), digital versatile disk (DVD), Blu-ray disc (BD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, solid state drives or other solid state storage devices, or other tangible and / or non-transitory media that can be used to store desired information. In this regard, the terms “tangible” or “non-transitory” herein as applied to storage, memory, or computer-readable media, are to be understood to exclude only propagating transitory signals per se as modifiers, and do not exclude any standard storage, memory, or computer-readable media that are not only propagating transitory signals per se.
[0033] Computer-readable storage media can be accessed by one or more local or remote computing devices, e.g., via access requests, queries, or other data retrieval protocols, for a variety of operations with respect to the information stored by the medium.
[0034] A system bus, as may be used herein, can be any of several types of bus structure that can further interconnect to a memory bus (with or without a memory controller), a peripheral bus, and a local bus using any of a variety of commercially available bus architectures. A database, as may be used herein, can include basic input / output system (BIOS) that can be stored in a non-volatile memory such as ROM, EPROM, or EEPROM, with BIOS containing the basic routines that help to transfer information between elements within a computer, such as during startup. RAM can also include a high-speed RAM such as static RAM for caching data.
[0035] As used herein, a computer can operate in a networked environment using logical connections via wired and / or wireless communications to one or more remote computers. The remote computer(s) can be a workstation, server, router, personal computer, portable computer, microprocessor-based entertainment appliance, peer device, or other common network node. Logical connections depicted herein may include wired / wireless connectivity to a local area network (LAN) and / or larger networks, e.g., a wide area network (WAN). Such LAN and WAN networking environments are commonplace in offices and companies, and facilitate enterprise-wide computer networks, such as intranets, any of which can connect to a global communications network, e.g., the Internet.
[0036] When used in a LAN networking environment, a computer can be connected to the LAN through a wired and / or wireless communication network interface or adapter. The adapter can facilitate wired or wireless communication to the LAN, which can also include a wireless access point (AP) disposed thereon for communicating with the adapter in a wireless mode.
[0037] When used in a WAN networking environment, a computer can include a modem or can be connected to a communications server on the WAN via other means for establishing communications over the WAN, such as by way of the Internet. The modem, which can be internal or external, and a wired or wireless device, can be connected to a system bus via an input device interface. In a networked environment, program modules depicted herein relative to a computer or portions thereof can be stored in a remote memory / storage device.
[0038] When used in either a LAN or WAN networking environment, a computer can access cloud storage systems or other network-based storage systems in addition to, or in place of, external storage devices. Generally, a connection between a computer and a cloud storage system can be established over a LAN or a WAN, e.g., via an adapter or a modem, respectively. Upon connecting a computer to an associated cloud storage system, an external storage interface can, with the aid of the adapter and / or modem, manage storage provided by the cloud storage system as it would other types of external storage. For instance, the external storage interface can be configured to provide access to cloud storage sources as if those sources were physically connected to the computer.
[0039] As employed in the subject specification, the term “processor” can refer to substantially any computing processing unit or device comprising, but not limited to comprising, singlecore processors; single-core processors with software multithread execution capability; multicore processors; multi-core processors with software multithread execution capability; multicore processors with hardware multithread technology; vector processors; pipeline processors; parallel platforms; and parallel platforms with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application specific integrated circuit (ASIC), a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), a state machine, a discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. Processors can exploit nano-scale architectures such as, but not limited to, molecular and quantum-dot based transistors, switches, and gates, in order to optimize space usage or enhance performance of user equipment. A processor may also be implemented as a combination of computing processing units. For example, a processor may be implemented as one or more processors together, tightly cou-pled, loosely coupled, or remotely located from each other. Multiple processing chips or multiple devices may share the performance of one or more functions described herein, and similarly, storage may be effected across a plurality of devices. A processor may be implemented to reside in a cloud-based network such as, e.g., the Internet.
[0040] The actions of a method or algorithm described in connection with the arrangements disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. A storage medium may be coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in functional equipment such as, e.g., a computer, a robot, a user terminal, a mobile telephone or tablet, a car, or an IP camera. In the alternative, the processor and the storage medium may reside as discrete components in such functional equipment. Additionally or alternatively, at least one of the processor and / or the storage medium may reside in a cloud-based network such as, e.g., the Internet.
[0041] Configurations of the present teachings are directed to computer systems for accomplishing the methods discussed in the description herein, and to computer readable media containing programs for accomplishing these methods. The raw data and results can be stored for future retrieval and processing, printed, displayed, transferred to another computer, and / or transferred elsewhere. Communications links can be wired or wireless, for example, using cellular communication systems, military communications systems, and satellite communications systems. Parts of the system can operate on a computer having a variable number of CPUs. Other alternative computer platforms can be used.
[0042] The present configuration is also directed to software / firmware / hardware for accomplishing the methods discussed herein, and computer readable media storing software for accomplishing these methods. The various modules described herein can be accomplished on the same CPU, or can be accomplished on different CPUs. In compliance with the statute, the present configuration has been described in language more or less specific as to structural and methodical features. It is to be understood, however, that the present configuration is not limited to the specific features shown and described, since the means herein disclosed comprise preferred forms of putting the present configuration into effect.
[0043] Methods can be, in whole or in part, implemented electronically. Signals representing actions taken by elements of the system and other disclosed configurations can travel over at least one live communications network. Control and data information can be electronicallyexecuted and stored on at least one computer-readable medium. The system can be implemented to execute on at least one computer node in at least one live communications network. Common forms of at least one computer-readable medium can include, for example, but not be limited to, a floppy disk, a flexible disk, a hard disk, magnetic tape, or any other magnetic medium, a compact disk read only memory or any other optical medium, punched cards, paper tape, or any other physical medium with patterns of holes, a random access memory, a programmable read only memory, and erasable programmable read only memory (EPROM), a Flash EPROM, or any other memory chip or cartridge, or any other medium from which a computer can read. Further, the at least one computer readable medium can contain graphs in any form, subject to appropriate licenses where necessary, including, but not limited to, Graphic Interchange Format (GIF), Joint Photographic Experts Group (JPEG), Portable Network Graphics (PNG), Scalable Vector Graphics (SVG), and Tagged Image File Format (TIFF).
[0044] Various arrangements are described herein. For simplicity of explanation, the methods or algorithms are depicted and described as a series of steps or actions. It is to be understood and appreciated that the various arrangements are not limited by the actions illustrated and / or by the order of actions. For example, actions can occur in various orders and / or concurrently, and with other actions not presented or described herein. Furthermore, not all illustrated actions may be required to implement the methods. In addition, the methods could alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, the methods described hereafter are capable of being stored on an article of manufacture, as defined herein, to facilitate transporting and transferring such methodologies to computers.
[0045] Referring to Fig. 1 , the invention is a method of and a system for perceiving a moving target T relative to a vehicle V. A controller (not shown) processes measurements from sensors having diverse modalities. Based on the nature and content of the information synthesized from the sensors, the controller prioritizes targets for enhanced analysis as needed and enables timely, cost effective autonomous vehicle navigation in ordinary traffic.
[0046] Referring to Fig. 2, a preferred embodiment of a method configured according to principles of the invention includes detecting 100 targets, perceiving 200 target locations and motions, predicting 300 target trajectories and planning 400 a travel route based on the target trajectories and intended route and / or destination data, which can guide navigation of vehicle V.
[0047] Detecting
[0048] Referring also to Fig. 3, detecting 100 relies on input from multiple sensors 105. Each sensor 105 has a modality 1 10 that provides different information about targets within anenvironment for determining the positions and trajectories thereof. Employing sensors of diverse modalities enables maximizing the individual strengths and minimizing the individual shortcomings of each modality. Aggregating data from temporally-related detections are processed to obtain more accurate estimates of the states of targets or dynamic targets of interest.
[0049] The sensors 105 preferably include a LIDAR sensor 105a detects targets in three dimensions (three-dimension (3D) mode) and can be positioned so as to capture a plan view in addition to side elevational views (two-dimension (2D) mode) of a target. The detection algorithm of the LIDAR based detection also provides three-dimensional information respecting a bounding box assumed for a detected target, in terms of length, width and height also with reference to an adopted coordinate system (3D Bounding Box (L, W, H)). A bounding box essentially is a rectangle that surrounds a target, that specifies its position, class (eg: car, person) and confidence (how likely the target is to be at that location). Bounding boxes mainly are used in the task of target detection where the aim is identifying the position and type of multiple targets in an image. Although not entirely reliable, LIDAR based detection also provides bounding box orientation, that is, how a rectangle assigned to a detected target is positioned or arranged at its location. While able to identify some target classes, namely automobiles and pedestrians, LIDAR may not be able to identify others for lack of training and data.
[0050] Other sensors 105 preferably include but are not limited to a camera 105b that operates in 2D mode and a radar sensor 105c that operates in 3D mode. Camera Based Detections (2DOD) provide only two-dimensional position information within the ground plane (x, y) regarding the location of a target in space relative to the detector. Camera detectors also provide only two-dimensional bounding box information respecting a detected target (2D Bounding Box (L, W)). However, camera detections extend to more target classes than other detection modes, not limited to automobile, pedestrian, bicycle, bus, motorcycle and large vehicles.
[0051] Other sensors 105 preferably include radar based detections that provide for three- dimensional position information regarding the location of a target in space relative to the detector (3D Position (x, y, z)). Radar based detections also provide velocity information in the radial direction regarding the motion of a target in space relative to the detector (Velocity (vx, vy, vz)). However, radar technology does not provide data that allows for sufficient definition of the configuration of a detected target, thus is unable to discern classes of targets or associated bounding boxes and orientations thereof. Nevertheless, the longer range capability of radar enables awareness of approaching targets for subsequent identification via other detection modalities.
[0052] Perceiving / Collecting
[0053] A difficulty with multi-mode sensor platforms is the variable measurement and transmission rates of each sensing modality. For example, LIDAR sensor 105a may sweep or sense a target field every 0.1 seconds, thus have a sampling rate of 0.1 seconds or 10 hz. Radar sensor 105c detections may occur at roughly 17hz and camera sensor 105b detections may occur at roughly 9hz. Accordingly, the invention includes a step of perceiving 200 for receiving detections from each sensing modality and, factoring in their respective temporal differences, combining them into a single message. Perceiving 200 includes “pseudosynchronizing” or temporally normalizing by combining detections that occur at about the same time or within a timeframe. The timeframe may be selected according to any criteria but, preferably, is based on a sampling time of a primary or driving sensor having a driving modality.
[0054] The Driving Modality is the slowest modality. As an example, based on the detection rates noted above, which may differ depending on particular sensors employed, the camera modality is the slowest modality. Accordingly, based on the foregoing detection rates, the camera modality is designated as the preferred Driving Modality in this example.
[0055] Sensor detections are collected in their respective containers 220. In this example, LIDAR sensor 105a detections are received in container 220a, camera sensor 205b detections are received in container 220b and radar sensor 205c detections are received in container 220c. At this point in the process, detections within the range of each sensor are collected regardless of relevance to navigation of the vehicle. Detections not considered relevant or considered in determining navigation for a vehicle V from a first position Pi to a second position P2are filtered out during predicting 300, as described below.
[0056] At step 223, perceiving 200 includes awaiting reception of a detection of the driving modality. Once received, that is, once a sensor 105 detection is received in container 220, then, at step 225, perceiving 200 includes awaiting reception of detections from the other modalities. Once at least one detection is received in each of the non-driving modalities, preferably, at step 230, perceiving 200 includes placing a Mutex lock on the detection containers to prevent a race condition.
[0057] At step 235, perceiving 200 includes identifying the detections of the non-driving modalities having timestamps closest to the timestamp of the driving modality. In this case, based on the detection rates noted above, perceiving 200 might identify a LIDAR detection having a timestamp closest to the timestamp of the camera detection and a radar detection having a timestamp closest to the timestamp of the camera detection. Perceiving 200 then includes combining the data of the driving modality detection and the temporally-nearest nondriving modality detections into a single message or muxed detection. In this case, perceiving200 combines the data of the camera detection with the data of the LIDAR detection having a timestamp closest to the timestamp of the camera detection and the data of the radar detection having a timestamp closest to the timestamp of the camera detection.
[0058] At step 240, perceiving 200 includes publishing or transmitting the muxed detection for subsequent processing.
[0059] At step 245, preferably, perceiving 200 includes clearing the containers 220 and prepares for collecting new detections.
[0060] At step 250, perceiving 200 includes unlocking the mutex lock on the containers 220 and permits reception of detections therein.
[0061] Referring to Fig. 4, at step 255, perceiving 200 includes detecting or identifying for individual processing according to respective modality the components of the muxed detection message published at step 240 in Fig. 3. For example, where available, a component respecting a modality may be associated with an attribute including one or more of: a classification identification; a confidence level respecting the reliability of the detection or identification; a bounding box; a velocity; and a dynamic state (static / dynamic).
[0062] At step 270, based on the modality, perceiving 200 includes converting the detection or signal associated with each component into a respective measurement. For example, the component from LIDAR sensor 105a is processed through algorithm 260a having a step 270 that converts the data of the LIDAR into a point in space relative to vehicle V.
[0063] Step 275 includes assigning a noise covariance or uncertainty to each measurement based on the mode associated therewith.
[0064] Referring also to Fig. 5, step 280 includes aggregating the measurements and assigned values (noise covariance and attributes) and defining a “measurements vector” 285.
[0065] Referring again to Fig. 4, perceiving 200 includes a step 290 of converting measurements vector 285 into a measurement space. Step 290 factors into each measurement conversion the noise covariance associated with the detection. Step 290 essentially converts raw data of the measurements vector 285 into a useable format, with common units, such as meters.
[0066] Perceiving / Tracking
[0067] Referring to Fig. 6, with a measurement vector 285 converted into a measurement space, perceiving 200 progresses to tracking 201 .
[0068] Referring also to Fig. 7, perceiving 200 includes step 204 of initializing tracker, which may include setting initial values. After step 204 comes a step 206 of selecting a tracking strategy. Depending on anticipated conditions, the invention provides for selecting a strategy from various strategies at run time Some exemplary tracking strategies include: a nearest-neighbor strategy; a nearest-neighbor JPDAF strategy; and a JPDAF strategy. The Nearest- Neighbor JPDAF tracking strategy is preferred.
[0069] Also following step 204 is step 207 of choosing a state estimation algorithm. The algorithm can be a Kalman Filter, an Unscented Kalman Filter or combinations thereof. The Kalman Filter is preferred.
[0070] Referring to Fig. 6, with the measurement space from step 290 and the tracking strategy and state estimation algorithm type set, the invention next initiates a step 202 of associating measurements with a new or existing target. Step 202 begins with a step 205 of gating measurements. Referring also to Fig. 1 , measurement gating 205 defines a gating region R about a target T that is relevant to navigation of vehicle V, as described above. An absolute location of region R is based on a projected position Ptfor target T as described below. The size of region R depends on the classification and tracking data for the target T. Region R is an empirically predetermined value based on the sampling rate and the top expected target speed of the target being tracked. A nonlimiting example of a formula for determining R is Rate x Vmax. For example, if a sensor sampled at 10hz and the tracked target could go as fast as 40m / s, then R would be set to 4 meters.
[0071] Once the location and size of region R are determined, step 205 includes identifying detections from the measurements vector that fall within region R. As shown in Fig. 1 , detections Daand Db, respectively having absolute positions Paand Pb, fall within, while detection Dc, having absolute position Pc, falls without region R.
[0072] Once the in-region detections are identified, in this case detections Daand Db, a step 210 provides for determining “cost calculations” for each of only the in-region detections relative to projected position Ptof target T. This limiting of cost calculations to only in-region detections, or gating, minimizes computations needed by not considering detections outside of gating region R of a given target / track. Detections or measurements that lie outside of gating region R are considered sufficiently unlikely to belong to the respective target.
[0073] Preferably, a cost calculation involves determining a Gaussian Likelihood that a measurement belongs to or is associable with a target T. Because each of the detection modalities operates according to different principles, each tends to impact the cost or position of a target differently from other modalities. For example, as shown on Fig. 1 , detections Daand Dbmay be detected by the LIDAR and camera sensors respectively at positions Paand Pb. Positions Paand Pb define different respective costs Caand Cb from position Pt. Cost calculation also can involve determining a Mahalanobis distance, Euclidian distance or other approaches consistent with the principles of the invention.
[0074] At step 215, perceiving 200 provides for assigning the lowest-cost calculation to the target T. For example, detection Dahaving a position Pawith the least cost Cafrom positionPtwould be the least of the associated costs, thus associated as the new position for target T for tracking.
[0075] Once a detection is associated with a target T, perceiving 200 includes determining whether a target T is new or previously tracked. New targets are stored as new tracks without further processing while new positions associated with older tracks are evaluated for track updating and prediction.
[0076] More specifically, with respect to new or unassociated measurements, step 220 includes creating tentative tracks for the new / unassociated measurements. For example, as shown in Fig. 1 , since detection Dcdoes not lie within the gating region R of any target / track, then Dcis unable to be associated to an existing target / track. Accordingly, the step 220 would provide for creating a tentative track for detection Dc. If a detected target T has only one associated measurement space, then step 202 proceeds to step 220 which provides for creating a tentative track in a tentative track memory 225 with which to associate that initial measurement space.
[0077] After the initial detection of a target T, perception 200 progresses to step 235 including state estimating 235 the detections associated with a target or track to project position Ptfor target T. Preferably, state estimating 235 employs Kalman filtering, also known as linear quadratic estimation (LQE). LQE processes measurements observed over time, including statistical noise and other inaccuracies, to obtain estimates of unknown variables that tend to be more accurate than those based on a single measurement alone, by estimating a joint probability distribution over the variables for each timeframe.
[0078] Kalman filtering also includes estimating the velocity of detected targets. Although radar sensors can measure velocity, estimated velocities without radar measurements are reasonably accurate. The behavior of these velocity estimates depends on the motion model choice and the accuracy of position measurements. For example, the Constant Velocity (CV) motion model assumes that a target is moving at a constant velocity. The preferred motion model is a Mean-Adaptive Acceleration Model, preferably a Singer Model with an adaptive mean. This motion model has a much wider coverage than traditionally used motion models, such as Constant Velocity and Constant Acceleration.
[0079] In addition to predictions respecting targets, state estimating 235 includes predicting a track / target ahead in time for use in step 202 associating measurements with and step 205 gating measurements using a joint probabilistic data-association filter (JPDAF). JPDAF is a statistical approach to the problem of data association in a target tracking algorithm. Rather than choosing the most likely assignment of measurements to a target or declaring the target not detected or a measurement to be a false alarm, the PDAF takes an expected value, which is the minimum mean square error (MMSE) estimate for the state of each target. At each time,the estimate of the target state is maintained as the mean and covariance matrix of a multivariate normal distribution. However, unlike the PDAF, which is only meant for tracking a single target in the presence of false alarms and missed detections, the JPDAF can handle multiple target tracking scenarios.
[0080] If a detected target T has more than one associated measurement space or existing track, then step 202 proceeds to step 255 of applying a filter, preferably a JPDAF filter, that updates the position associated with the existing track. Processing then proceeds to step 240 of managing tracks.
[0081] One function of track management 240 involves track maintenance 245. A component of track maintenance 245 involves classifying the dynamics of a tracked target, that is whether the tracked target is dynamic or static. Dynamics Classification provides the prediction node, described below, with contextual information about the targets being tracked. If a target is classified as static, then the prediction node need not perform tasks for predicting a future location. This classification can be fairly difficult to establish because of measurement noise that causes measurements of targets that are truly static to suggest large estimated velocities. The invention is configured to provide accurate classifications despite this noise by utilizing the position and velocity history of a given target. It also determines and utilizes static and dynamic probabilities based on an interacting multiple model (IMM) filter.
[0082] Another function of track management 240 involves track life management 250. Track life management essentially confirms tentative tracks and removes stale tracks. To this end, once a target associated with a tentative track has been detected a number of times within a duration of time, such as 5-10 times per second, then the tentative track associated with that target is classified as a confirmed track stored in a confirmed track memory 230.
[0083] An important function of track maintenance 245 is estimating a true heading of a target T. The invention incorporates a target heading correction function to overcome the inherent unreliability of heading measurements received from 3DOD sensors. The measurements are unreliable because LIDAR detections lack context, such as the front and back of a vehicle and which direction the vehicle is travelling.
[0084] Because of this lack of reliable heading measurements, target headings are estimated based on a target / track history including the estimated velocity associated with the track associated with the target.
[0085] Predicting
[0086] Referring to Figs. 2 and 8, from the estimated positions and velocities derived from the measurement vectors of perceiving 200, predicting 300 also provides forecasts as to the future motion of moving targets, i.e. forecasting the future trajectories of moving targets (e.g. future horizon of 6 seconds), based on their track history (e.g. 2 seconds) and bird’s eye viewmaps around targets. Because of multi-modal distribution of future trajectories, for each target, predicting 300 provides multiple probable trajectories and their respective probabilities to the planner.
[0087] Predicting 300 has an External Communications Node 305 that handles communication with nodes external to the prediction pipeline. On the input side, node 305 listens to tracking 200 for state estimates of targets around vehicle V. The state estimates include an identification number, position, velocity, acceleration, orientation and category. Node 305 also operates as a bookkeeper of relevant objects or objects of concern. Old objects for which state data no longer exists, signifying disappearance from a scene, are dropped. New objects associated with data different from existing tracks are added, and existing tracks associated with targets persisting in the scene are updated.
[0088] Referring also to Fig. 9, node 305 passes the packaged estimates to a Deterministic Predictions node 310. Deterministic Predictions node 310 is a fast node that produces a set of all possible predictions for a target given its location, orientation and lane information. A “lane” is an area, typically modeled as a polygon, to which is assigned metadata, such as: a direction of movement or traffic; class (parking, pedestrian walkway, motor vehicle roadway); intersection. The lane in which the target is determined to exist depends on the location of nearby targets and pre-mapped information. The lane to which a target is assigned is determined by its (x,y) proximity and whether it falls within an angle threshold of the target orientation and the lane orientation. Preferably, the angle threshold should not exceed 90sfrom the direction in which the target travels.
[0089] The deterministic prediction algorithm is a depth first search algorithm that generates a lane tree based on termination conditions and predictions along each branch of the lane tree. The algorithm begins by retrieving all lanes that contain a target. If no such lane exists, (for example, due to inaccurate tracking of a dynamic obstacle, the position of the obstacle is estimated to be outside of a lane polygon) the algorithm would include lanes that exist within a threshold distance, such as 5 meters, and within the angle threshold of the target. If no such lanes exist, then the method generates a straight-line prediction. If such lanes exist, then, starting with the lane closest to the target, for each of these initializing lanes, the algorithm:
[0090] 1 . Retrieves all the lanes that are:
[0091] a. physically connected to the initializing lane; and
[0092] b. within 90 degrees of the initializing lane.
[0093] For example, as shown in Fig. 9, a target T traversing road R can continue along line C or switch to lanes aligned with lines Li or l_2. At a four-way intersection, target T can continue in the lane along line C, switch to lanes aligned with lines Li or L2, turn left along line Y or turnright along line Z. Accordingly, the algorithm retrieves all of lanes C, Li , L2, Y and Z for further processing.
[0094] 2. The algorithm adds the newly retrieved lanes as child lanes of the initializing lane to the lane tree. For each of the child lanes, the algorithm locates the children thereof that satisfy conditions 1 .a. and 1 .b. above. The algorithm expands a branch until the either of the following termination conditions are met for a child:
[0095] a. Referring to Fig. 10, the child lane C does not have any children of its own and has a dead end D. In this case, the algorithm creates a constant velocity prediction along all the lanes that are in the same branch as the child starting from the current location L and ending at the dead end D.
[0096] b. Referring to Fig. 11 , the child lane C lies outside of the distance threshold H of prediction. In this case, the algorithm creates a prediction along all of the lanes that are in the same branch as the child starting from the current location L of the target and ending at the distance threshold H. The location of the distance threshold depends on the velocity of the target. For example, for a 6-second prediction, the distance threshold should be 6 times target velocity. Preferably, predictions assume constant velocity and are limited to 6 seconds, which balances the increased computational requirements of longer the predictions and the need for longer predictions in view of assumed urban environment speed limits of about 45 mph.
[0097] Referring to Fig. 12, an example prediction tree for a vehicle at an intersection is shown. In the example, seven branches in the prediction tree are identified by the relatively associated encircled numerals 1 -7. For each branch, the algorithm generates seven corresponding constant velocity predictions. Example predictions Gi and G2are shown in dashed lines for respective branches 2 and 6.
[0098] Referring again to Fig. 8, once predictions are completed for all of the targets supplied by node 305, if the target category associated with a prediction set falls within a critical group of categories, then node 310 passes the prediction sets to a Target Criticality Score node 315. The critical group of categories may include, for example, large vehicles, including automobiles, buses, trucks, etc., because greater harm is likely to arise from collision with a bus than, e.g. an errant kickball. Targets not within the critical group are associated with a straight-line prediction.
[0099] Node 315 receives critical group predictions from node 310 and further reduces computational demands and prioritizes computational resources for a predetermined number of those targets that should be accorded the greatest concern. First, node 315 filters out stationary targets, such as parked cars, trees, etc. Then node 315 evaluates collision likelihood of the remaining targets. The predetermined number depends only on the amountof processing power employed in implementing the invention. Node 315 then passes the predetermined number of target sets to a machine learning node 320.
[0100] Referring also to Fig. 13, node 320 receives the prioritized target sets from node 315 along with rasterized maps pertaining to the localized area within a predefined area of each target. For example, as shown in Fig. 13, with further processing, the data allow for planning for the vehicle V on the map M pertaining to the localized area. In the map M, vehicle V travels with critical targets Ti and T2on their respective rasterized maps Mi and M2. The maps Mi and M2pertain to the localized area in which critical targets Ti and T2travel. From the model, node 320 employs conventional machine learning techniques, not limited to of deep leaning, inverse reinforcement learning and machine learning classification approaches, to develop and cluster trajectories and probabilities that are passed to an External Communications Node 325 described below. Several parameters factor into the motion prediction model, including map data, class, velocity / acceleration / heading angle / heading angle rate of the target, and position / velocity / acceleration of surrounding objects.
[0101] Target sets not passed from node 315 to node 320 are passed directly to node 325.
[0102] If, after node 310, the target category associated with a prediction set falls without the critical group of categories, then node 310 passes the prediction sets directly to node 325.
[0103] Referring to Fig. 8, node 325 receives the deterministic and machine learning predictions respecting targets of interest and passes same to a path planning module 400. Path planning module 400 generates a path that a navigation controller (not shown) of vehicle V uses for guiding vehicle V to an intended destination.
[0104] Referring to Fig. 14, a system configured according to principles of the invention may be deployed on an autonomous vehicle (AV) 500. Preferably, AV 500 includes a processor (not shown) that receives communications from sensors having diverse modalities, not limited to a LIDAR sensor 505, a camera 510 and a radar sensor 515.
[0105] The foregoing focuses on innovative methods for understanding the locations and trajectories of moving targets. The purpose of the invention is to inform vehicle navigation so as to avoid or accommodate the moving targets. The intricacies of the navigation strategies and processes do not impact target location perception, tracking and projection, thus are beyond the scope of this application.
[0106] While the principles of the invention have been described herein, the foregoing description is only an example and not a limitation on the scope of the invention. Other embodiments are contemplated within the scope of the present invention in addition to the exemplary embodiments shown and described herein. Modifications and substitutions by one of ordinary skill in the art are within the scope of the present invention. The invention is notlimited to the particular embodiments described and depicted herein, rather only to the following claims.
Claims
CLAIMSWe claim:1 . Method of perceiving a dynamic target comprising: receiving multiple detections of diverse modalities respecting the target; identifying detections occurring within a period of time and defining contemporary detections; associating a measurement with each of the contemporary detections; for each measurement, calculating a distance therefrom to a predicted target position; and defining as an accepted target position the measurement associated with the least distance.
2. Method of claim 1 further comprising: associating an attribute with each of the contemporary detections; and calculating each measurement based on a detection and the attribute associated therewith.
3. Method of claim 2 wherein the attribute is selected from a noise covariance; a bounding box; a velocity; a dynamic; a class; an uncertainty; and combinations thereof.
4. Method of claim 1 wherein said calculating is limited to a measurement that falls within an area.
5. Method of claim 1 wherein: one of the modalities is a driving modality associated with a driving detection; and said identifying comprises the driving detection and detections other than the driving detection temporally closest to the driving detection.
6. Method of claim 1 wherein said receiving is from: a LIDAR sensor; a camera; and / or a radar sensor.
7. Method of claim 1 wherein the area, said calculating and / or said assigning derives from Joint Probabilistic Data Association.
8. Method of claim 1 wherein the predicted target position derives from a stored prediction and the least distance and / or Kalman filtration.
9. Method of claim 1 further comprising generating a track based on the accepted target position.
10. Method of claim 1 further comprising updating a stored position with the accepted target position and defining an updated position.11 . Method of claim 10 further comprising updating a track with the updated position.
12. Method of claim 1 further comprising navigating a vehicle is based on the accepted target position.
13. System for perceiving a dynamic target comprising: a plurality of sensors having unique modalities and configured for generating detections respecting a target; and a processor configured for: identifying detections occurring within a period of time and defining contemporary detections; associating a measurement with each of the contemporary detections; for each measurement, calculating a distance therefrom to a predicted target position; and defining as an accepted target position the measurement associated with the least distance.
14. System of claim 13 wherein said processor is configured for: associating an attribute with each of the contemporary detections; and calculating each measurement based on a detection and the attribute associated therewith.
15. System of claim 14 wherein the attribute is selected from a noise covariance; a bounding box; a velocity; a dynamic; a class; an uncertainty; and combinations thereof.
16. System of claim 13 wherein the calculating is limited to a measurement that falls within an area.
17. System of claim 13 wherein: one of the modalities is a driving modality associated with a driving detection; and the identifying comprises the driving detection and detections other than the driving detection temporally closest to the driving detection.
18. System of claim 13 wherein said sensors are selected from: a LIDAR sensor; a camera; a radar sensor; and combinations thereof.
19. System of claim 13 wherein the area, the calculating and / or the assigning derives from Joint Probabilistic Data Association.
20. System of claim 13 wherein the predicted target position derives from a stored prediction and the least distance and / or Kalman filtration.21 . System of claim 13 wherein said processor is configured for generating a track based on the accepted target position.
22. System of claim 13 wherein said processor is configured for updating a stored position with the accepted target position and defining an updated position.
23. System of claim 22 wherein said processor is configured for updating a track with the updated position.
Citation Information
Patent Citations
Detection system for predicting information on pedestrian
EP4036892A1
System and method for modeling advanced automotive safety systems
US20160245949A1