Multimodal sensor fusion method for cross-domain correspondence based on object trajectory

By calculating the spatio-temporal object trajectory matching of cameras and radar/lidar sensors, the problem of inaccurate data integration and high recognition error rate in multimodal sensor fusion system is solved, and high accuracy of moving object recognition and vehicle environment perception is achieved, and real-time decision-making of autonomous driving and early warning systems is supported.

CN113261010BActive Publication Date: 2025-08-22YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980087727.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-01-10
Publication Date
2025-08-22
Estimated Expiration
2039-01-10

AI Technical Summary

Technical Problem

The existing multimodal sensor fusion system has the problem of inaccurate data integration in vehicle perception and high error rates for identification and classification. Especially in the process of data integration of different sensors, it is difficult to effectively match the trajectory of moving objects.

Method used

By calculating the distance measurement and similarity between the space-time object trajectory of the camera and the radar/lidar sensor, using techniques such as optimization methods and Kalman filters, the matching and fusion of different sensor data is achieved to generate high-quality estimation of motion objects.

Benefits of technology

It improves the accuracy of the identification and classification of moving objects, enhances the perception ability of the vehicle's surrounding environment, and supports real-time decision-making of autonomous driving and early warning systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113261010B_ABST
    Figure CN113261010B_ABST
Patent Text Reader

Abstract

A system for calculating correspondences between multiple sensor observations includes: an interface for receiving a first spatiotemporal dataset and a second spatiotemporal dataset of multiple moving objects based on signals from a first sensor and a second sensor that capture different types of signals; a processing circuit for: generating a plurality of first spatiotemporal object trajectories and a plurality of second spatiotemporal object trajectories based on the first spatiotemporal dataset and the second spatiotemporal dataset, respectively; calculating a distance metric between the first and second spatiotemporal object trajectories by corresponding mappings from corresponding domains to a predetermined domain; performing a matching calculation between the mappings of the plurality of first and second spatiotemporal object trajectories using the distance metric so as to calculate the similarity between each pair of mappings of the first and second spatiotemporal object trajectories, and outputting the result of the matching calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] In some embodiments of the invention, the invention relates to multimodal sensor fusion of moving object data, but not exclusively, to a multimodal sensor fusion system for moving object data generated by vehicle sensors.

[0002] Automotive applications often require a clear perception of the area surrounding the vehicle. Perception includes, for example, identifying moving objects such as other vehicles and pedestrians, as well as static objects such as road signs and debris. To generate appropriate perception, vehicles are equipped with sensors such as cameras, wireless detection and ranging (radar), and / or light detection and ranging (lidar) sensors. Advanced warning systems use the information generated by the sensors to warn of potential hazards, while autonomous vehicle control systems use it to safely maneuver the vehicle. Advanced warning systems need to operate in real time and minimize errors in the identification and classification of moving objects while integrating the data received from different sensors. Summary of the Invention

[0003] It is an object of some embodiments of the present invention to provide a system and method for multimodal sensor fusion of moving object data.

[0004] The aforementioned and other objects are achieved by the features of the independent claims. Further embodiments are apparent from the dependent claims, the description and the drawings.

[0005] According to a first aspect of the present invention, a system for calculating correspondences between multiple sensor observations of a scene comprises: an interface for receiving a first spatiotemporal dataset and a second spatiotemporal dataset of multiple moving objects in the scene, the first spatiotemporal dataset and the second spatiotemporal dataset being based on signals reflected from the multiple moving objects and originating from a first sensor and a second sensor, respectively, wherein the first sensor and the second sensor capture different types of signals; a processing circuit for: generating a plurality of first spatiotemporal object trajectories and a plurality of second spatiotemporal object trajectories based on the first spatiotemporal dataset and the second spatiotemporal dataset, respectively; calculating a distance metric between the first and second spatiotemporal object trajectories by corresponding mappings from corresponding domains to a predetermined domain; performing calculations between each of the first and second spatiotemporal object trajectories and calculating the Euclidean distance between feature sequences.

[0006] According to a second aspect of the present invention, a method for calculating the correspondence between multiple sensor observations of a scene includes: an interface for receiving a first spatiotemporal dataset and a second spatiotemporal dataset of multiple moving objects in the scene, the first spatiotemporal dataset and the second spatiotemporal dataset being based on signals reflected from the multiple moving objects and originating from a first sensor and a second sensor, respectively, wherein the first sensor and the second sensor capture different types of signals; generating a plurality of first spatiotemporal object trajectories and a plurality of second spatiotemporal object trajectories based on the first spatiotemporal dataset and the second spatiotemporal dataset, respectively; calculating a corresponding mapping from a corresponding domain to a predetermined domain for each of the first and second spatiotemporal object trajectories; using a distance metric to perform a matching calculation between the mappings of the plurality of first and second spatiotemporal object trajectories so as to calculate the similarity between each pair of mappings of the first and second spatiotemporal object trajectories, and outputting the result of the matching calculation.

[0007] With reference to the first and second aspects, optionally, the first spatiotemporal dataset originates from a camera sensor, for example, in many vehicles a camera is mounted in front of the cab to record activities in front of the vehicle.

[0008] With reference to the first and second aspects, optionally, the second spatiotemporal dataset is derived from a radar and / or light detection and ranging (LiDAR) sensor. Radar and / or LiDAR sensors are widely available and effective in tracking the position and velocity of moving objects and can supplement the tracking information provided by camera sensors.

[0009] With reference to the first and second aspects, the matching calculation optionally pairs the trajectories so that the sum of the distances between the trajectories is minimized across all possible combinations of trajectory matches. Determining which trajectories match each other is a computationally difficult and complex problem. Therefore, using optimization methods is an effective way to overcome this difficulty.

[0010] With reference to the first and second aspects, optionally, the predetermined domain is a domain comprising a coordinate system according to the first spatiotemporal trajectory originating from the camera sensor. This approach may be advantageous when the velocity of the moving object is relatively easy to estimate and the terrain is expected to be relatively flat, as the height of the object projected from the radar / lidar domain can be roughly estimated.

[0011] With reference to the first and second aspects, optionally, the method further comprises assigning a velocity to each moving object associated with the first spatiotemporal trajectory by estimating the difference between the positions of the corresponding moving objects within the first spatiotemporal trajectory. Estimating the velocity of the moving object can be achieved by known effective solutions such as optical flow.

[0012] With reference to the first and second aspects, optionally, the height of each moving object corresponding to the trajectory of the second spatiotemporal dataset is estimated using a predetermined constant. This estimation is suitable if the terrain is expected to be relatively flat, such as a road or track, and the moving objects are expected to be similar, such as cars, then assigning the same height to each moving object.

[0013] With reference to the first and second aspects, optionally, the predetermined domain is a domain including a coordinate system based on the second spatiotemporal trajectory derived from the radar and / or lidar sensor. Since camera depth information can be used for velocity estimation and improve the accuracy of the matching calculation, this method can be used to address issues such as uneven terrain.

[0014] With reference to the first and second aspects, optionally, the corresponding mapping uses a bounding box representation of each moving object represented in the first spatiotemporal dataset, so as to estimate the velocity of the corresponding moving object by calculating the corresponding difference between the projections of the centroid of the corresponding bounding box onto the vehicle coordinates. Estimating the velocity in this manner is simple and computationally efficient.

[0015] With reference to the first and second aspects, optionally, the predetermined domain includes a domain of latent vehicle coordinates, and the corresponding mapping maps each possible combination of a pair of the first spatiotemporal moving object trajectory and the second spatiotemporal moving object trajectory to a corresponding latent state. The latent state-based approach can combine precise radar / lidar velocity detection with precise camera spatial detection to achieve a more stable matching calculation than projection onto the corresponding domain.

[0016] With reference to the first and second aspects, optionally, the corresponding mapping employs a Kalman filter in order to estimate the plurality of latent states. Kalman filters comprise a group of effective existing solutions for estimating properties of moving objects.

[0017] With reference to the first and second aspects, optionally, the matching calculation includes: estimating the likelihood that the two trajectories correspond to the same object by first calculating a sequence (trajectory) of latent states, and then for each estimated latent state: calculating projections in the first and second domains corresponding to the first and second spatiotemporal datasets, respectively. The likelihood can be estimated by measuring the Euclidean distance between the sequence of projections and the actual observed signal trajectory. For example, under the assumption of a Gaussian distribution with a known covariance matrix, the matching between multiple first and second spatiotemporal moving object trajectories is calculated in the following manner: calculating the sum that maximizes the calculated probabilities of all estimated latent states, and matching between the first and second spatiotemporal moving object trajectories corresponding to the corresponding latent states in the calculated sum. A useful method for calculating the latent space is to use a nonlinear Kalman filter on the two datasets simultaneously, and optionally smoothing it. After the distance metric is calculated, it can be used for matching, for example by minimizing the sum of the distances between all matching trajectories.

[0018] After reviewing the following drawings and detailed description, other systems, methods, features and advantages of the present disclosure will become apparent to those skilled in the art. It is intended that all such additional systems, methods, features and advantages be included within this description, be within the scope of the present disclosure, and be protected by the following claims.

[0019] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention relates. Exemplary methods and / or materials are described below, but methods and materials similar or equivalent to those described herein can also be used in the practice or testing of embodiments of the present invention. In the event of a conflict, the present patent specification, including definitions, will prevail. In addition, the materials, methods, and examples are illustrative only and are not intended to be necessarily limiting. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Some embodiments of the present invention are herein described, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is emphasized that the details shown are by way of example and for purposes of illustrative discussion of embodiments of the invention. In this regard, the description taken in conjunction with the drawings will make apparent to those skilled in the art how the embodiments of the invention may be practiced.

[0021] In the diagram:

[0022] Figure 1 is an exemplary layout of various components of a system for multimodal sensor fusion of moving object data according to some embodiments of the present invention;

[0023] Figure 2 is an exemplary data flow of a process for multimodal sensor fusion of moving object data according to some embodiments of the present invention;

[0024] Figure 3 is an exemplary data flow of a process for mapping a spatiotemporal trajectory onto a domain corresponding to a camera domain according to some embodiments of the present invention;

[0025] Figure 4 is an exemplary data flow for a process of mapping a spatiotemporal trajectory onto a domain corresponding to a radar / lidar domain according to some embodiments of the present invention;

[0026] Figure 5 is a data stream for a process of matching spatiotemporal trajectories using a latent state model according to some embodiments of the present invention;

[0027] Figure 6 is a table of attribute values ​​of potential states and update functions for evaluating the probabilities of the potential states using a Kalman filter according to some embodiments of the present invention;

[0028] Figure 7 is an exemplary depiction of camera and radar / lidar trajectories, each in a respective domain of the same scene, according to some embodiments of the present invention; and

[0029] Figure 8 is an exemplary depiction of the application of an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In some embodiments of the present invention, the present invention relates to multimodal sensor fusion of moving object data, and more particularly, but not exclusively, to a multimodal sensor fusion system for moving object data generated by vehicle sensors.

[0031] According to some embodiments of the present invention, multimodal sensor fusion systems and methods are provided in which datasets generated from different sensors are integrated to present the data to a user in a combined manner. For example, the systems and methods can fuse spatiotemporal streaming data from sensors such as a camera, one or more radars, and / or one or more lidars.

[0032] Multimodal sensor fusion systems present many computational challenges, such as overcoming the different coordinate systems of different sensors, incompatible detection times (e.g., false positives and false negatives), cross-temporal alignment between different modalities (e.g., different sensing frequencies for different sensors), and noise. Overcoming these computational challenges is important because multimodal sensor fusion systems are often used in time-critical and sensitive situations, such as in early warning systems for preventing vehicle accidents.

[0033] Existing multimodal sensor fusion solutions exhibit some shortcomings. For example, in "On-road vehicle detection and tracking using MMW radar and monovision fusion" (IEEE Transactions on Intelligent Transportation Systems 17.7 (2016): 2075-2084), a system is described that fuses radar and image modalities using tracks. However, the fusion is performed on a frame-by-frame level, and the tracks are only used as an additional verification layer.

[0034] In Lynen, Simon et al., "Arobust and modular multi-sensor fusion approach applied to MAV navigation" (Intelligent Robots and Systems (IROS), 2013 IEEE / RSJ International Conference, IEEE, 2013), a system is described that tracks multiple modalities across time to obtain better obstacle estimates than when using a single modality, but no solution is provided for matching observations between modalities.

[0035] In contrast, the multimodal sensor fusion system described in this paper performs matching between different modalities by matching spatiotemporal trajectories, which are calculated separately for each modality and then input into the matching calculation for matching between complete trajectories. Matching between trajectories detected by cameras and radar can improve robustness because each trajectory contains more information about the moving object than a single camera frame.

[0036] Another potential advantage is to use the system described herein offline to assist in offline training and evaluation in order to generate high quality estimates of the match between modalities.

[0037] The following is a brief description of the data flow for spatiotemporal data processing according to some embodiments of the present invention. For the sake of brevity, 'spatiotemporal trajectories' may be referred to herein as 'trajectories', 'radar and / or lidar' may be referred to herein as 'radar / lidar', and 'variables' and 'attributes' may be used interchangeably herein.

[0038] Each of the described systems includes processing circuitry for executing code, and each of the described methods is implemented using processing circuitry for executing code. The processing circuitry may include hardware and firmware and / or software. For example, the processing circuitry may include one or more processors and a non-transitory medium connected to the one or more processors and carrying code. When the code is executed by the one or more processors, the system performs the operations described herein.

[0039] For example, spatiotemporal data from a camera and / or radar / lidar is cached by counting a predetermined number of camera frames and caching data from the radar / lidar corresponding to a time interval according to the predetermined number of camera frames. Next, code is executed for generating a spatiotemporal trajectory for each moving object detected by each sensor.

[0040] Next, for each pair of spatiotemporal trajectories, a "distance" metric is calculated to quantify the similarity between each pair of spatiotemporal trajectories. Optionally, the trajectories are processed by code to map them into a predetermined domain. According to some embodiments of the present invention, the predetermined domain may include a coordinate system based on data from a camera sensor, or a coordinate system based on data from a radar / lidar sensor, or a potential temporal domain including a set of position, orientation, and velocity attributes for each moving object.

[0041] Processing spatiotemporal trajectories may include estimating velocity and / or estimating the size and / or height of the corresponding moving objects. For example, a camera typically produces two-dimensional data without depth in a field of view corresponding to the camera lens position and orientation, while a radar / lidar typically produces two-dimensional data with depth but without height depending on the radar / lidar orientation and mounting position on the vehicle. In addition, a calibration matrix that provides the relative positions of the camera and radar with respect to the ego-vehicle center (the center of the vehicle on which the camera and radar are mounted) and camera projection properties (intrinsic and extrinsic properties) are used to align the data from the different sensors.

[0042] Next, after estimating the speed and height of the corresponding tracks of the corresponding moving objects and aligning the coordinates, a track matching calculation is performed on the mapped tracks to identify each moving object by combining the two corresponding tracks. Optionally, the matching calculation uses a Euclidean distance metric, and this metric is used to match the tracks, for example, by minimizing the sum of the distances between pairs in all possible track matching combinations, which can be performed by applying an optimization algorithm.

[0043] According to some embodiments of the present invention, in which a latent time domain is used as the predetermined domain for mapping trajectories, a latent state is estimated for each pair of camera and radar / lidar trajectories used as the latent model, and a statistical model is assumed to relate it to the latent state variables. In the matching calculation, the statistical model is employed to match trajectories whose joint latent state has a high probability of originating from a single moving object.

[0044] After the matching calculation, the matching results are output to corresponding components, such as one or more vehicle controllers and / or one or more output displays.

[0045] Before explaining at least one embodiment of the present invention in detail, it should be understood that the present invention is not necessarily limited in its application to the details of construction and arrangement of components and / or methods set forth in the following description and / or illustrated in the drawings and / or examples. The present invention is capable of other embodiments or can be practiced or carried out in various ways.

[0046] The present invention may be a system, method, and / or computer program product. The computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to perform various aspects of the present invention.

[0047] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof.

[0048] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to a corresponding computing / processing device or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network).

[0049] The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or an external computer may be connected (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits, such as programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), may execute computer-readable program instructions by personalizing the electronic circuit using state information of the computer-readable program instructions to perform various aspects of the present invention.

[0050] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0051] The flow charts and block diagrams in the figure illustrate the architecture, functionality and process of the possible implementation schemes of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent the part of a module, section or instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions mentioned in each box may not occur in the order mentioned in the figure. For example, depending on the functionality involved, the two boxes shown in succession can actually be executed substantially simultaneously, or these boxes can sometimes be executed in the opposite order. It should also be noted that each box in the block diagram and / or flow chart description, and the combination of the boxes in the block diagram and / or flow chart description can be implemented by a dedicated hardware-based system that performs a specified function or action, or implements a combination of dedicated hardware and computer instructions.

[0052] See now Figure 1 , which is a depiction of system components and related vehicle components in a multimodal sensor fusion system 100 according to some embodiments of the present invention. According to some embodiments of the present invention, the system 100 is used for multimodal sensor fusion, wherein the spatiotemporal sensor 102 is composed of a camera and a radar / lidar. For example, the system 100 can be integrated into a vehicle as part of a warning mechanism, accessed by the driver through a heads-up display (HUD) to present the driver with a fused sensor information stream, such as the speed, direction, and highlighted indications of moving objects, which helps the driver shorten the reaction time when there is a risk of collision with debris, pedestrians, and other vehicles. Alternatively, the system 100 can present fused data about moving objects in front of the vehicle to a vehicle controller responsible for autonomous driving of the vehicle, so that the vehicle controller can decide actions regarding vehicle navigation based on the moving object fused data.

[0053] The I / O interface 104 receives raw spatiotemporal data from the spatiotemporal sensor 102, and one or more processors 108 execute code stored in the memory 106 to process the raw spatiotemporal data. The code includes instructions for performing multimodal sensor fusion of moving object data by generating spatiotemporal trajectories from the raw spatiotemporal data according to predetermined data sampling criteria and matching the generated spatiotemporal trajectories across different modalities, for example, matching camera trajectories with radar / lidar trajectories.

[0054] The matching results of the spatiotemporal trajectory are output by one or more processors executing code instructions through the I / O interface 104, and the output can be directed to the output display 110 and / or the controller 112. For example, if the vehicle using the system is controlled by a driver, the output can be directed to the output display 110, and if the vehicle is autonomously controlled by the controller 112, the controller 112 can use the output for obstacle avoidance and route planning.

[0055] See now Figure 2 , which is an exemplary data flow of a process for multimodal sensor fusion of moving object data according to some embodiments of the present invention. First, as shown at 200 and 202, raw camera and radar / lidar data are received, for example, camera frame data is received as a data stream at a certain rate and / or radar / lidar detections are received at a certain frequency.

[0056] Next, as shown in 204 and 206, the received raw data is checked for a predefined sampling rate, which may be different for the camera and / or radar / lidar sensor. The predefined sampling rate can be, for example, a predefined number of seconds, a predefined number of camera frames, and / or a predefined number of radar / lidar detections. If the sampling rate is achieved, a spatiotemporal trajectory is generated for each raw data stream, as shown in 208. The spatiotemporal trajectory is generated by an extended object tracking process over time. According to some embodiments of the present invention, a Kalman tracking process is used to track and generate the spatiotemporal trajectory.

[0057] Next, as shown in 210, the generated spatiotemporal trajectories are mapped onto a predetermined domain. These are represented by the X, Y, and Z horizontal, vertical, and depth axes, respectively. The radar / lidar domain describes moving objects in the XZ plane, including each moving object's velocity, while the camera domain describes moving objects in the XY plane. Furthermore, a calibration matrix can be used to align trajectories from different sensors, based on their positions relative to the center of the vehicle.

[0058] Optionally, the predetermined domain may include a domain corresponding to a camera domain or a radar / lidar domain, or may include a latent time domain. For different embodiments of the present invention, each choice of predetermined domain may exhibit certain advantages / disadvantages compared to other choices. For example, mapping spatiotemporal trajectories to a latent time domain may increase trajectory matching accuracy, but may also increase computational complexity.

[0059] Next, as shown in 212, after the spatiotemporal trajectories are mapped onto the predetermined domain, a matching calculation is performed between the mapped spatiotemporal trajectories to match trajectories corresponding to the same moving object. The matching calculation is performed by using a predefined distance metric, such as the Euclidean distance, and by applying an optimization algorithm to minimize the sum of distances. For example, the Kuhn-Munkres ("Hungarian") algorithm can be used to select matches (i.e., a subset of all pairs in which each trajectory appears at most once) such that the sum of distances is minimized.

[0060] Next, as shown in 214 , the result of the matching calculation is output. The matching trace may be output to the output display 110 and / or the controller 112 via the I / O interface 104 .

[0061] Also refer to Figure 3 , which is an exemplary data flow of a process for mapping a spatiotemporal trajectory to a domain corresponding to a camera domain according to some embodiments of the present invention, as depicted at 208. After the spatiotemporal trajectory is generated, as shown at 206, the camera and radar / lidar trajectories are received, as shown at 300 and 302, respectively. Next, as shown at 304, for each camera trajectory given in XY coordinates, a velocity is estimated. Optionally, the velocity is estimated by averaging the optical flow within the object's boundaries. Optical flow is the pattern of apparent motion of an image object between two consecutive frames due to the motion of the object or camera. There are many ways to calculate optical flow, from vision processing algorithms (e.g., the Lucas-Kanade method) to deep neural networks (e.g., using FlowNet 2.0). The output of an optical flow algorithm typically includes motion vectors in the camera domain for the image pixels. The average optical flow within the object's boundaries provides an estimate of the object's motion projected onto the camera plane. The series of velocity estimates calculated over the camera trajectory can be weighted summed with the velocity calculation based on the bounding box to obtain a more reliable estimate. The exact values ​​of the weights in the weighted sum can be learned from a sample training dataset using machine learning techniques.

[0062] As shown at 306, the estimated height of each moving object is mapped to the radar / lidar track given in XZ coordinates. Optionally, each radar / lidar track is assigned a predefined, identical height, for example, one meter. Next, as shown at 308, the radar / lidar track, along with its assigned height, is mapped to the camera domain, optionally using a calibration matrix for coordinate alignment. Finally, as shown at 310, the mapped spatiotemporal track is output for further processing by the mapping calculation 210.

[0063] Also refer to Figure 4, which is an exemplary data flow for a process of mapping a spatiotemporal trajectory to a domain corresponding to the radar / lidar domain, as depicted at 208, according to some embodiments of the present invention. First, as shown at 400, the process receives a camera trajectory. Next, as shown at 402, for each camera trajectory, a temporal sequence of two-dimensional spatial bounding boxes is retrieved, where each bounding box corresponds to a corresponding moving object in the corresponding camera frame. According to some embodiments of the present invention, the bounding boxes are retrieved by applying a predefined camera object detection algorithm. Next, as shown at 404, for each camera trajectory, the velocity of the corresponding moving object is estimated. The velocity can be estimated in a similar manner as explained with respect to 304, or optionally, include bounding box-based calculations. Next, as shown at 406, the camera trajectory is mapped to the radar / lidar domain using a coordinate calibration matrix. Finally, as shown at 408, the mapped camera and radar / lidar trajectories are output for further processing by the mapping calculation 210.

[0064] Next, refer to Figure 5 , which is an exemplary data flow for a process for matching spatiotemporal trajectories using a latent state model according to some embodiments of the present invention, related to 208 and 210. First, as shown at 500, after generating the trajectories at 206, the process receives the spatiotemporal trajectories. Next, as shown at 502, a latent state is estimated for each possible combination of paired camera and radar / lidar trajectories. Each spatiotemporal trajectory can be assigned a numerical index for reference, and the pairing of trajectories can be performed according to the ascending index.

[0065] Optionally, the latent state is estimated using a nonlinear dual-model Kalman filter with smoothing and used as observations for estimating the corresponding bounding box and the corresponding radar / lidar detection for each corresponding trajectory in the corresponding trajectory pair used in the latent state.

[0066] Each potential state has attributes (x, y, z, V(x), V(y), A(x), A(y), h, w), including: distance from the ego vehicle on the axis of motion, position on the axis orthogonal to the motion, velocity X component, velocity Y component, acceleration X component, acceleration Y component, height of the object in the plane facing the ego vehicle, and width of the object in the plane facing the ego vehicle. Figure 6 Exemplary details of the attribute values ​​and update functions used by the Kalman filter are described in FIG.

[0067] Next, as shown in 504, for each estimated latent state, a mapping is performed into the camera domain and the radar / lidar domain. The purpose of the mapping is to estimate, for each mapping, a conditional probability, as shown in 506, which represents the event that the corresponding latent state corresponds to a moving object detected by the spatiotemporal sensor, given the corresponding mapping. For example, the camera observations are preprocessed into a bounding box represented by an upper left value and a lower right value in the camera domain. Given a latent state, a camera observation with upper left and lower right values ​​can be calculated by projecting the values ​​derived from the latent state onto the camera domain. Next, the difference from the observed camera properties is evaluated, and the product of a Gaussian density function with a predetermined variance is calculated (e.g. Figure 6 described in detail in ).

[0068] Next, the two conditional probabilities corresponding to the corresponding latent states are summed, and the resulting sum is designated as the distance between the corresponding pairings of camera and radar / lidar tracks.

[0069] Next, as shown at 508 , it is checked whether there are any unmapped camera and radar / lidar track pairs, and if so, the selection of a track pair continues at 502 .

[0070] After completing the pairing trajectories and calculating the conditional probability of each pair, as shown in 510, the probability matrix P is generated. i,j , where P(i,j) equals the conditional probability estimated in 506 that the corresponding camera track with assigned index i and the corresponding radar / lidar track with assigned index j originate from observations of the same object.

[0071] Next, as shown in 512, the camera and radar / lidar trajectories are matched by maximizing the sum of the probabilities in P(i, j) so that at most one entry is selected from each row and column of P(i, j). According to some embodiments of the present invention, matching is performed using the Kuhn-Munkres ("Hungarian") algorithm. Next, as shown in 514, the matched trajectories are output to 212.

[0072] See now Figure 6 , which is a table of attribute values ​​of potential states and update functions for evaluating the probabilities of the potential states using a Kalman filter according to some embodiments of the present invention;

[0073] As shown at 600, each potential state attribute is assigned an initial value, which can be determined experimentally. As shown at 602, each attribute is assigned an initial probability, where the Kalman filter uses the initial probability in a standard manner, where VarForDeltaAndP(delta,p) represents the variance satisfying Pr[|X - mean| < delta] = p, where X is distributed according to the normal distribution N(mean,Var).

[0074] Now refer to Figure 7 , which depicts camera and radar / lidar trajectories according to some embodiments of the present invention, each trajectory in a corresponding domain of the same scene. As shown at 700, a camera frame is displayed, which shows cars in front of the vehicle. As shown at 700, each car is surrounded by a bounding box, and the camera trajectory of the moving cars is mapped onto the frame in the form of a sequence of points, representing car detections over multiple time frames. As shown at 702, a trajectory derived from radar data is displayed, where the cone represents the boundary of the radar sensor in front of the vehicle. The radar trajectory corresponds to the same time interval as the camera trajectory in 700. The application of the potential state model is demonstrated at 704 and 706, where two trajectories marked by ellipsoids are matched in the corresponding domains as originating from the same car.

[0075] Now refer to Figure 8 , which depicts an exemplary application of an embodiment of the present invention. As shown at 800, the camera frame displays a corresponding bounding box and an estimated distance for each car, which is obtained by fusing camera and radar data by matching the corresponding trajectories. For example, the camera display can be an HUD designated for the driver to warn the driver of any potential hazards, such as a sudden change in the direction of surrounding cars.

[0076] After reviewing the following drawings and detailed description, other systems, methods, features, and advantages of the present disclosure will be apparent or will become apparent to those skilled in the art. All such additional systems, methods, features, and advantages are intended to be included within this description, within the scope of the present disclosure, and protected by the appended claims.

[0077] The description of the various embodiments of the present invention has been presented for purposes of illustration, but the description is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, the practical application, or the technical improvement of the technology found in the marketplace, or to enable one of ordinary skill in the art to understand the embodiments disclosed herein.

[0078] It is expected that during the life of the patent expiring on this application, many related systems, methods, and computer programs will be developed, and the scope of the term sensor fusion is intended a priori to include all such new technologies.

[0079] As used herein, the term "about" refers to ±10%.

[0080] The terms "comprises," "comprising," "includes," "including," "having," and their conjugates mean "including but not limited to." This term encompasses the terms "consisting of" and "consisting essentially of.

[0081] The phrase "consisting essentially of means that the composition or method may include additional ingredients and / or steps, but only if the additional ingredients and / or steps do not materially alter the basic and novel characteristics of the claimed composition or method.

[0082] Unless the context clearly indicates otherwise, as used herein, the singular forms "a," "an," and "the" include plural references. For example, the term "compound" or "at least one compound" may include a plurality of compounds, including mixtures thereof.

[0083] The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments, and / or does not exclude the incorporation of features from other embodiments.

[0084] The word “optionally” is used herein to mean “is provided in some embodiments and not provided in other embodiments.” Any particular embodiment of the present invention may include a number of “optional” features unless such features are incompatible.

[0085] Throughout this application, various embodiments of the present invention may be presented in range format. It should be understood that the description in range format is for convenience and brevity only and should not be construed as a fixed limitation on the scope of the invention. Therefore, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values ​​within the range. For example, a range description of 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within the range, such as 1, 2, 3, 4, 5, and 6. This applies regardless of the breadth of the range.

[0086] Whenever a numerical range is indicated herein, it is meant to include any cited numeral (fractional or integer) within the indicated range. The phrases "a range between a first indicated numeral and a second indicated numeral" and "a range from a first indicated numeral to a second indicated numeral" are used interchangeably herein and are intended to include the first indicated numeral and the second indicated numeral and all fractional and integer numerals therebetween.

[0087] It will be appreciated that certain features of the invention described in the context of separate embodiments for the sake of clarity may also be provided in combination in a single embodiment. Conversely, various features described in the context of a single embodiment for the sake of brevity may also be provided individually or in any suitable subcombination or in any other described embodiment of the invention as appropriate. Certain features described in the context of various embodiments are not considered essential features of those embodiments unless the embodiment would not function without those elements.

[0088] All publications, patents, and patent applications mentioned in this specification are hereby incorporated by reference in their entirety. Similarly, each individual publication, patent, or patent application is specifically and individually indicated as being incorporated herein by reference. In addition, citation or identification of any reference in this application should not be construed as an admission that such reference may be used as prior art for the present invention. Where section headings are used, they should not be construed as necessarily limiting.

Claims

1. A system for calculating correspondences between multiple sensor observations of a scene, characterized in that: The system comprises: an interface for receiving a first spatiotemporal dataset and a second spatiotemporal dataset of a plurality of moving objects in a scene, wherein the first spatiotemporal dataset and the second spatiotemporal dataset are based on signals reflected from the plurality of moving objects and originate from a first sensor and a second sensor, respectively, the first sensor and the second sensor capturing different types of signals; Processing circuitry for: generating a plurality of first spatiotemporal object trajectories and a plurality of second spatiotemporal object trajectories based on the first spatiotemporal dataset and the second spatiotemporal dataset, respectively; computing a distance metric between the first and second spatiotemporal object trajectories by corresponding mappings from the respective domains to a predetermined domain; wherein the predetermined domain comprises a domain according to a coordinate system of the first spatiotemporal object trajectory originating from a camera sensor, or the predetermined domain comprises a domain according to a coordinate system of the second spatiotemporal object trajectory originating from a radar and / or lidar sensor, or the predetermined domain comprises a latent temporal domain; performing a matching calculation between the plurality of mappings of the first and second spatiotemporal object trajectories using a distance metric so as to calculate a similarity between each pair of mappings of the first and second spatiotemporal object trajectories, and The result of the matching calculation is output.

2. The system according to claim 1, wherein: The first spatiotemporal dataset originates from a camera sensor.

3. The system according to claim 1, wherein: The second spatiotemporal dataset originates from a radar and / or light detection and ranging (lidar) sensor.

4. The system according to claim 3, characterized in that The matching calculation pairs trajectories so as to minimize the sum of the distances between the trajectory pairs among all possible trajectory matching combinations.

5. The system according to claim 4, characterized in that The system further includes assigning a velocity to each moving object associated with the first spatiotemporal object trajectory by estimating a difference between respective moving object positions within the first spatiotemporal object trajectory when the predetermined domain includes a domain according to a coordinate system of the first spatiotemporal object trajectory originating from the camera sensor.

6. The system according to claim 5, characterized in that When the predetermined domain comprises a domain according to a coordinate system of the first spatiotemporal object trajectory originating from the camera sensor, the height of each moving object corresponding to the trajectory of the second spatiotemporal dataset is estimated by a predetermined constant.

7. The system according to claim 4, wherein: When the predetermined domain includes the domain of the coordinate system of the second spatiotemporal object trajectory originating from the radar and / or lidar sensor, the corresponding mapping uses a bounding box representation of each moving object represented in the first spatiotemporal dataset to estimate the speed of the corresponding moving object by calculating the corresponding difference between the center of mass of the corresponding bounding boxes.

8. The system according to claim 4, wherein: When the predetermined domain includes the latent temporal domain, the corresponding mapping maps each possible combination of pairs of first and second spatiotemporal moving object trajectories to a corresponding latent state of a plurality of latent states.

9. The system according to claim 8, characterized in that When the predetermined domain includes the latent time domain, the corresponding mapping employs a Kalman filter in order to estimate the plurality of latent states.

10. The system according to claim 9, characterized in that When the predetermined domain includes the potential time domain, the matching calculation includes: For each estimated potential state: computing corresponding projections in first and second domains corresponding to the first and second spatiotemporal datasets, respectively; Calculating the probability that a corresponding pair of first and second spatiotemporal moving object trajectories originate from the same moving object based on the corresponding projections; and performing matching between the plurality of first and second spatiotemporal moving object trajectories by the following operations: Compute the sum that maximizes the calculated probabilities of all estimated potential states, Matching is performed between the first and second spatiotemporal moving object trajectories corresponding to respective latent states in the calculated sum.

11. A method for calculating the correspondence between multiple sensor observations of a scene, characterized in that: The method comprises: receiving a first spatiotemporal dataset and a second spatiotemporal dataset of a plurality of moving objects in a scene, wherein the first spatiotemporal dataset and the second spatiotemporal dataset are based on signals reflected from the plurality of moving objects and originate from a first sensor and a second sensor, respectively, the first sensor and the second sensor capturing different types of signals; generating a plurality of first spatiotemporal object trajectories and a plurality of second spatiotemporal object trajectories based on the first spatiotemporal dataset and the second spatiotemporal dataset, respectively; computing a distance metric between the first and second spatiotemporal object trajectories by corresponding mappings from the respective domains to a predetermined domain; wherein the predetermined domain comprises a domain according to a coordinate system of the first spatiotemporal object trajectory originating from a camera sensor, or the predetermined domain comprises a domain according to a coordinate system of the second spatiotemporal object trajectory originating from a radar and / or lidar sensor, or the predetermined domain comprises a latent temporal domain; performing a matching calculation between the plurality of mappings of the first and second spatiotemporal object trajectories using a distance metric so as to calculate a similarity between each pair of mappings of the first and second spatiotemporal object trajectories, and The result of the matching calculation is output.

Citation Information

Patent Citations

  • Multi-source target fusion method based on track matching

    CN108280442A