System and method for predicting future position of object
By using deep learning models and Kalman filters in ADAS systems, the future location of pedestrians can be predicted based on 2D camera image data. This solves the problems of high cost and complexity in pedestrian detection in ADAS systems and achieves accurate and robust pedestrian location prediction with simple hardware configuration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FARCOSAI AUTOMOBILE CO LTD
- Filing Date
- 2024-09-23
- Publication Date
- 2026-04-21
AI Technical Summary
The high cost and complexity of existing pedestrian detection systems in advanced driver assistance systems (ADAS) limit their application in simple hardware configurations, and existing technologies struggle to predict the future location of pedestrians based on 2D vision.
A deep learning model based on two-dimensional image data is used to detect objects, and the future position of the objects is predicted by a Kalman filter. Image data is captured using a simple 2D camera, and the predicted state data of the objects is generated by combining the posture history and movement characteristics to achieve the prediction of the future position.
It reduces system complexity and cost, enables accurate pedestrian position prediction with simple hardware configuration, and improves the ADAS system's ability to avoid collisions and activate emergency braking mechanisms, especially showing more robust and stable prediction performance in occluded scenarios.
Smart Images

Figure CN121909497A_ABST
Abstract
Description
[0001] This application claims the benefit of European patent application EP23382972.0, filed on September 25, 2023.
[0002] This disclosure relates to a method for predicting the future location of objects, such as living organisms, based on two-dimensional (2D) camera vision. Furthermore, a system operating by said method is disclosed. Background Technology
[0003] It is estimated that more than half of all traffic fatalities are pedestrians on public roads. Therefore, one of the current safety challenges that Advanced Driver Assistance Systems (ADAS) aim to address is those related to vulnerable users, especially pedestrians.
[0004] Motor vehicles already incorporate ADAS systems with various features for safety and accident prevention. In fact, current ADAS systems include a set of sensors and devices that enable these features, which are key features for pedestrian detection. Therefore, pedestrian detection systems can enable ADAS systems to proactively avoid collisions and activate emergency braking mechanisms.
[0005] Currently, pedestrian detection systems can be equipped with visual cameras, stereo cameras, LiDAR, radar, depth sensors, and computer vision / analysis software.
[0006] However, in recent years, vehicle purchases have become increasingly driven by price and safety considerations. This has particularly impacted the development of vehicles in the lower price range, presenting a challenge in equipping them with complex sensors while maintaining low costs. Therefore, the aforementioned stereo cameras, LiDAR, radar, and depth sensors, along with computer vision / analysis software, can increase the cost of pedestrian detection systems. Furthermore, ADAS systems are becoming increasingly complex, limiting their application in simple hardware configurations.
[0007] Most existing technologies for predicting the future location of pedestrians are configured to use 3D or depth information by combining signals from at least two or more of 3D camera arrays, LiDAR, radar, depth sensors, GPS, etc. Therefore, most existing technologies are not based solely on images captured by simple 2D vision cameras.
[0008] US20210150193 A1 discloses a method for recognizing the movement intention of a pedestrian from camera images. It utilizes an object detector to detect pedestrians in at least one camera image, determines a skeletal model of the detected pedestrians to serve as a pose representative, thus uses the pedestrian model to obtain the pose of the detected pedestrians, classifies the movement contours of the detected pedestrians, and outputs a class describing the object's movement intention. It does not disclose a method for predicting the future location of an object (e.g., a pedestrian).
[0009] Furthermore, Kalman filters are known in computer vision for tracking objects and performing prediction and correction tasks in 2D images of video feeds. Kalman filters can be used to predict and correct the future position of rigid objects. However, the applicant is unaware of any existing techniques for predicting the future position of a pedestrian, taking into account their pose. Summary of the Invention
[0010] The purpose of this invention is to provide a method for predicting the future location of an object, the method comprising: - Capture video streams that include two-dimensional (2D) image data; - Detect objects in two-dimensional (2D) image data; - Obtain the pose of an object that includes multiple key points; - The pose history of the generated object; - Determine the movement characteristics of two or more key points of an object; - To obtain predicted state data for each of the two or more key points by estimating the estimated state of each key point based at least on the movement characteristics and the attitude history; and - Predict the future location of at least a reference point of an object from predictive state data, wherein the future location is predicted on two-dimensional (2D) image data of a video stream.
[0011] The future position of the reference point is the position that the reference point is predicted to have at a predefined time. As described below, the predefined time can be between 1 second and 3 seconds. Therefore, this method is configured to predict the position that an object will have on the two-dimensional (2D) image data of the video stream within approximately two seconds.
[0012] The step of predicting future location can be performed by a location prediction module deployed on a platform, as further explained below. The location prediction module may include a Kalman filter configured to receive predicted state data (e.g., position, velocity, and acceleration) of two or more keypoints of an object in two-dimensional (2D) image data. Preferably, the object may be a movable object (e.g., a living organism). More preferably, the keypoints are joints of a living organism (e.g., a human). The two or more keypoints include a first keypoint and a second keypoint. Furthermore, the two-dimensional (2D) image data extends along a horizontal axis and a vertical axis. Therefore, the predicted state data may include the following information (e.g., items) for the first keypoint: position on the horizontal axis, position on the vertical axis, velocity on the horizontal axis, velocity on the vertical axis, and preferably, acceleration on the horizontal axis and acceleration on the vertical axis. Furthermore, the predicted state data may include the following information (e.g., items) for the second keypoint: position on the horizontal axis, position on the vertical axis, velocity on the horizontal axis, velocity on the vertical axis, and preferably, acceleration on the horizontal axis and acceleration on the vertical axis. These items may be arranged in the form of a state vector. That is, the predicted state data may include a state vector, which in turn includes the velocity, speed, and acceleration of each keypoint in the two-dimensional (2D) image data. The step of predicting the future position of the reference point may include extrapolating the state vector, preferably using the transition matrix of a Kalman filter.
[0013] The video stream can be captured by a vision camera of the vehicle. In practice, the vision camera can be mounted on the body of the motor vehicle. The vehicle vision camera can preferably capture a field of view including the external frontal area of the motor vehicle. More preferably, the vision camera can include a front-facing monocular camera, but other applications are certainly possible within the meaning of this disclosure. For example, the vision camera can be used for a surrounding view system or a top-view system, and other applications. Furthermore, the video stream can include multiple sequential two-dimensional (2D) image data (e.g., images). Preferably, the video stream can include at least two image frames (more preferably, three image frames). It is likely preferred that the vision camera has a frame rate of at least fifteen frames per second (15 fps). Furthermore, the step of detecting objects in the two-dimensional (2D) image data described above can be performed in the first image frame, wherein the step of estimating the estimated state of each of two or more keypoints can also be performed on subsequent second image frames. Additionally, the future position of a reference point of the object is predicted on image frames after the first image frame.
[0014] The visual camera is a simple 2D camera that does not include a 3D camera, a time-of-flight (TOF) camera, sensor fusion, or any device that senses depth. That is, the 2D image data does not include 3D image data or depth information. The simplicity of the 2D camera allows for a reduction in the complexity of this method and also allows for the integration of this method into systems. Furthermore, cost reduction can be achieved. It should be emphasized that no global positioning signal (i.e., GPS signal) or any other positioning signal or input containing depth information is required to predict the future position of an object on image frames of a video stream.
[0015] Objects can be detected in captured two-dimensional (2D) image data using a deep learning model. Preferably, multiple objects can be detected in the captured two-dimensional (2D) image data using a deep learning model.
[0016] In any case, objects can preferably be detected in association with bounding boxes. A bounding box (also called a boundary region) is a geometry in two-dimensional (2D) image data that surrounds or encloses an object or, more optionally, a group of objects. Preferably, the bounding box is a rectangle or a square, but of course, different shapes are possible. The bounding box can be used as a guide point for detected objects and to create collision boxes for said objects. Reference points can be predefined in the bounding box as described below.
[0017] Preferably, the method further includes the step of classifying the objects such that they are assigned to a specific object class. Thus, this method can detect two or more objects in the same two-dimensional (2D) image data, wherein a first object is assigned to a first object class and a second object is assigned to a second object class. An object class is a grouping of objects that can be described based on attributes common to their members. For example, an object can be a movable object. Preferably, movable objects can be distinguished between movable items and living beings. Movable items and living beings are generally moving or likely to move in the near future (i.e., the object will move soon (e.g., within the next four seconds, preferably within the next two seconds)). Furthermore, living beings can be subclasses that include animals and humans (e.g., pedestrians). Additionally, object classes or subclasses within the human category (or generic class) can be, for example, males, adults, children, and people with reduced mobility (such as disabled persons). In this case, when the object class is human, there may be certain commonalities. For example, a person can move, for example, by crossing a street, when the intention is to move at least one part of the human body.
[0018] Preferably, the detected object and pose are obtained in a single step of the deep learning model. More preferably, this method can utilize a single deep learning model and the same steps (i.e., one stage) to obtain the following: (i) Detect objects within the bounding box (e.g., pedestrians); and (iii) The pose of an object (e.g., a pedestrian).
[0019] A pose can include multiple keypoints linked or connected to each other. The pose of a (detected) object represents at least the object's position and orientation in two dimensions, i.e., on two-dimensional (2D) image data. The pose can be internally stored as a transformation matrix or a measurement matrix. Therefore, pose can preferably be a lightweight, general-purpose solution for detecting the position and orientation of a living organism's body in real time (i.e., at a given moment). Preferably, if the object is a living organism such as a pedestrian, the pose can be described in real time using a set of skeletal landmarks (also referred to herein as multiple keypoints) to describe the living organism's body position at a given moment. More preferably, the keypoints correspond to or represent joints of the human body. The pose can be a kinematic model.
[0020] A kinematic model (also known as a skeleton-based model) is used in this invention for 2D pose estimation. This flexible and intuitive human body model includes a set of joint positions and limb orientations to represent human anatomy. Therefore, the skeleton pose estimation model can be used to capture the relationships between different body parts.
[0021] Therefore, preferably, keypoints among a plurality of keypoints are connected to form a skeletal representation of a living organism. Furthermore, the plurality of keypoints includes two or more keypoints and adjacent keypoints, such that two or more keypoints are connected to adjacent keypoints. In practice, the two or more keypoints may include at least a first keypoint and a second keypoint that are interconnected. Furthermore, the first keypoint is connected to a (first) adjacent keypoint. Further, the plurality of keypoints may include a second adjacent keypoint connected to a second keypoint. Thus, this method allows for the prediction and tracking of the location of a living organism, as described below.
[0022] The bounding box may include reference points. Preferably, reference points of the bounding box can be predefined in a specific region of the bounding box. That is, the reference point can be a predefined representative point of the bounding box. More preferably, the predefined representative point can be located at the center of the bounding box. Other locations are not excluded. Furthermore, the representative point is not a keypoint, nor is it part of a plurality of keypoints. Preferably, there can be a certain distance between the current point and the aforementioned keypoints. In addition, the reference point can be fixed within the bounding box, i.e., immovable. Conversely, two or more keypoints have mobility within the bounding box. In fact, each of the two or more keypoints can move relative to each other. That is, each keypoint, as or representing a joint of a living organism, can have a position, velocity, or preferably acceleration that is different from each other and within the bounding box.
[0023] Preferably, this method includes comparing the detected object with a previous detection from a previous frame. Furthermore, each detected object can be tracked across different image frames. More preferably, an identifier, also known as a unique identifier (UID), is assigned to the tracked detection.
[0024] Preferably, the pose history of the aforementioned objects can be generated in real time for each object. This is created based on the set of incoming detections (bounding boxes + pose estimates) and the association of active entities with specified UIDs. This allows for approximation of motion characteristics and modeling of weights for certain key points to take into account the reliability of associated previous detections for prediction.
[0025] Preferably, the aforementioned motion characteristics may include the position and velocity of each of two or more keypoints. More preferably, the motion characteristics may include the position, velocity, and acceleration of each of two or more keypoints in the two-dimensional (2D) image data. More preferably, the motion characteristics of adjacent keypoints are determined, and more preferably, a second adjacent keypoint is determined. That is, the motion characteristics of the first keypoint may include the position, velocity, and acceleration of the first keypoint. The motion characteristics of the second keypoint may include the position, velocity, and acceleration of the second keypoint. The motion characteristics of the (first) adjacent keypoints may include the position, velocity, and acceleration of the (first) adjacent keypoints. Further, the motion characteristics of the second adjacent keypoints may include the position, velocity, and acceleration of the second adjacent keypoints. Perhaps even more preferably, motion characteristics are determined or estimated for multiple keypoints. As described above, the video stream may include multiple sequential two-dimensional (2D) image data. Therefore, the aforementioned keypoints can be identified or located in the first two-dimensional (2D) image data that can be tracked by subsequent second two-dimensional (2D) image data. In this way, position and velocity can be determined.
[0026] Furthermore, the steps described above for determining the motion characteristics of two or more key points of an object may include estimating the motion characteristics of two or more key points of the object in two-dimensional (2D) image data.
[0027] Preferably, this is achieved by approximating the velocity and acceleration of an object (e.g., a pedestrian) at the pixel scale.
[0028] As described above, predicted state data for each of two or more keypoints is obtained by estimating the estimated state of each of the keypoints based at least on motion characteristics and pose history. The predicted state data can include the position, velocity, and acceleration of the keypoints in two-dimensional (2D) data. Similarly, the estimated state can include the position, velocity, and acceleration of the keypoints in two-dimensional (2D) data. Therefore, position includes two items: position on the horizontal axis and position on the vertical axis. Similarly, velocity can include two items: velocity on the horizontal axis and velocity on the vertical axis. Likewise, acceleration can include two items: acceleration on the horizontal axis and acceleration on the vertical axis. Since position, velocity, and acceleration can be performed at the pixel scale, predicted state data for each of two or more keypoints can be obtained without GPS and depth information. In practice, predicted state data (e.g., pixels per second) can be obtained within an image reference system.
[0029] Preferably, the step of obtaining the predicted state data may further include estimating the future states of adjacent keypoints. The future states may include the future positions of the adjacent keypoints or at least parameters obtained from them. More preferably, the future states of any adjacent keypoint may include future positions, velocities, and accelerations in two-dimensional (2D) image data. Furthermore, the predicted state data includes the preliminary future states of each of two or more keypoints obtained through the estimated states of each keypoint and the future states of adjacent keypoints. This allows for the estimation of human pose even in occluded scenes, resulting in more robust and stable predictions compared to other solutions.
[0030] Preferably, the step of obtaining the predicted state data may further include determining the confidence or determinism of the estimated position. This is done by determining the uncertainty through the confidence of the estimated state for each keypoint by adding or subtracting the importance of the incoming detections. Confidence values are obtained from the deep learning model and added to the position prediction to add or subtract the uncertainty used to predict N+1 frames or N+X. Thus, the preliminary future state can be further based on the confidence or determinism of the estimated state. Similarly, the preliminary future state may include the position, velocity, and acceleration of keypoints on two-dimensional (2D) image data. Thus, position includes two items: position on the horizontal axis and position on the vertical axis. Similarly, velocity may include two items: velocity on the horizontal axis and velocity on the vertical axis. Likewise, acceleration may include two items: acceleration on the horizontal axis and acceleration on the vertical axis.
[0031] Preferably, the step of obtaining the predicted location data may further include the step of assigning weights to the estimated states. Therefore, the preliminary future state can be further based on the weights of the estimated states.
[0032] Preferably, the step of obtaining the predicted state data may further include: correcting the estimated state of each of two or more key points to obtain a corrected state of each of the two or more key points. Optionally, the step of obtaining the predicted state data may further include: correcting the aforementioned preliminary future state or at least the data obtained from the estimated state.
[0033] More preferably, the corrected estimated state or preliminary future state is obtained by performing a Kalman filter. A Kalman filter (also known as a linear quadratic estimate (LQE)) uses a series of measurements observed over time (e.g., the estimated or preliminary future position of each of two or more keypoints), which may include statistical noise and other inaccuracies, and produces estimates of unknown variables (e.g., the corrected position of each keypoint tends to be more accurate than those based solely on a single measurement) by estimating the joint probability distribution of the variables over each time frame. Furthermore, the Kalman filter keeps track of the estimated state (or any data derived from it) of the estimated position or preliminary future position, as well as the variance or uncertainty of said estimate. The estimate is updated using a state transition model and the measurements.
[0034] Preferably, the step of obtaining the predicted state data may further include the step of predicting the predicted position of each of two or more keypoints at a predefined time by extrapolating and correcting the position. More preferably, the predefined time may be 1 to 3 seconds. Even more preferably, the predefined time may be between 1.5 seconds and 2.5 seconds, approximately two seconds.
[0035] In short, this method includes a deep learning model capable of detecting both the position and pose of each living being (e.g., a pedestrian). As described above, the deep learning model can be configured to output an estimated state (e.g., estimated position, velocity, and acceleration) for each of a plurality of keypoints, based at least on the position and velocity of each keypoint and its pose history. Subsequently, tracking based on two-dimensional (2D) image data can be performed, for example, by employing a Kalman filter to obtain the corrected positions of the multiple keypoints. Thus, the Kalman filter allows for improved accuracy in the prediction of each of the multiple keypoints. Due to limited features, the predictions made by the Kalman filter are substantially insensitive to occlusion, faults, and association challenges. The characteristic of this method is that the future position of a living being (e.g., a pedestrian) can be dynamically predicted based not only on position or distance but also on pose. This is achieved by predicting the future position of a reference point of the bounding box from the corrected position, approximately two seconds from the moment the video stream was captured.
[0036] Some or all of the above steps can be performed by embedded software. Furthermore, the embedded software can be integrated into an embedded system. This embedded software can simply run on an 8-bit microcontroller with several gigabytes of memory and a suitable level of processing complexity determined by an approximately correct computational framework. As shown in the example below, embedded software including a deep learning model and a Kalman filter can be integrated into an embedded platform that executes in real time.
[0037] The purpose of this invention is to provide a system for predicting the future location of an object operated by the above-described method. Attached Figure Description
[0038] Non-limiting examples of the system and method for predicting the future location of an object will now be described with reference to the accompanying drawings.
[0039] In the picture: Figure 1 This invention relates to an automated production line; Figure 2 This is an illustration showing the use of multiple types of detection for pedestrians and electric vehicles; Figure 3 It is an example of two skeleton representations and two people, each surrounded by a skeleton representation; Figure 4 This is an example of a tracking module with multiple pedestrians; Figure 5 This is an example of real-time prediction at pedestrian crossings; and Figure 6 This is a diagram illustrating the steps of an example method according to the present invention. Detailed Implementation
[0040] Referring to the non-limiting examples shown in the accompanying drawings, this document describes a method for predicting the future location of an object, and a system operated by said method.
[0041] The method is a computer-implemented invention performed in an Advanced Driver Assistance System (ADAS) within the field of motor vehicles. The claimed method focuses on safety, particularly regarding pedestrians and other vulnerable road users (VRUs). Furthermore, the claimed method is deployed (e.g., embedded) on a simple hardware configuration while maintaining or even improving its accuracy and performance. It is configured to obtain pedestrian position prediction based on deep learning and is ready for integration into an embedded platform (e.g., Texas Instruments TDA4VM) for real-time execution. In the example, the deep learning model is based on YOLO-Pose, capable of detecting both the position and pose of each pedestrian without constraining the number of pose keypoints. Subsequently, tracking (and association) based on a multidimensional Kalman filter is performed, which ultimately allows for dynamic prediction of the pedestrian's future position based not only on position or distance but also on pose.
[0042] Therefore, this invention relates to a method for pedestrian detection, tracking, and location prediction that allows vehicles to take better and more timely actions to prevent potential collisions, thereby helping to reduce accidents and improve overall pedestrian safety. It enables the prediction of pedestrian movement within 2D spatial image coordinates (i.e., on two-dimensional (2D) data of a video stream captured by, for example, a camera-based visual camera array). Therefore, it enables the prediction of pedestrian positions within said two-dimensional (2D) data, allowing ADAS systems to proactively avoid collisions and activate emergency braking mechanisms as needed.
[0043] This invention differs from other existing methods by focusing on pedestrian behavior (rather than distance or location). This allows for the estimation of human pose even in occluded scenarios, resulting in more robust and stable predictions compared to other solutions.
[0044] This invention provides a comprehensive method and system for achieving accurate state estimation of individual joint positions and generating accurate predictions of VRU movement while addressing occlusion challenges. This method and system can be integrated independently or seamlessly into the data input pipeline, enhancing overall system effectiveness and adaptability, especially when integrated with other ADAS systems.
[0045] As explained, the present invention aims to track and predict the movement of detected moving objects (such as living beings, particularly pedestrians). Figure 1 As shown, this method includes four modules: image dedistortion module 10, person detection and pose estimation module 20, pedestrian tracking module 30, and position prediction module 40.
[0046] In this example, the video stream is captured by a vehicle vision camera (such as a front-facing monocular camera system) 100, serving as a cost-effective alternative to methods based on high-end sensors (such as stereo cameras, LiDAR, and radar). This design ensures scalability and meets the stringent requirements of practical embedded automotive hardware. The system includes a monocular fisheye camera and an electronic control unit (ECU) based on the Texas Instruments (TI) DA4VM, dedicated to deep learning and computer vision.
[0047] Image dedistortion module 10 is used to correct image distortion. Automotive cameras with wide field of view (FoV) suffer from distortion, thus compromising the detection task of pinhole cameras. This is addressed by employing the Kannala method, a nonlinear algorithm that uses a pre-computed lookup table based on inherent camera parameters to correct distortion. This accurately determines how light is re-projected onto the sensor, thus mimicking pinhole camera behavior and enabling the observation of straight lines in the real world without compromising object size. Furthermore, image dedistortion module 10 is capable of applying the proposed method to fisheye cameras, which are currently the most widely used in real-world ADAS systems. Results demonstrate enhanced regression metric accuracy and robust tracking performance compared to similar systems, providing a lightweight solution that can be easily deployed to real-world vehicle applications.
[0048] Therefore, this method and system are configured to predict the future position of a pedestrian in 2D space using at least four image coordinate samples for each observed pedestrian. In the example, a 133ms observation time per pedestrian is required for robust and stable state estimation at a camera frame rate of 30 FPS. This enables accurate prediction, which is crucial for rapid response in collision avoidance situations.
[0049] The person detection and pose estimation module 20, executed by a control unit deployed with deep learning, is configured to simultaneously detect different types of pedestrians and extract their poses in a single inference. This method maintains high prediction accuracy while allowing for real-time implementation on a cost-effective platform.
[0050] In this example, a comprehensive data processing pipeline was developed, adapting input and output data to the needs of each module. The modules are implemented in C++ and leverage deep learning accelerators on their respective platforms. The PC pipeline uses Nvidia's TensorRT, and the embedded TDA4VM platform uses TDL-RT. This programming strategy optimizes performance and resource management during execution.
[0051] The person detection and pose estimation module 20 is used to acquire the detection and pose of each pedestrian. In the example, a custom AI model based on YOLO-Pose developed by TI is proposed. Unlike the original YOLO-Pose method, this method allows for multi-class detection, and the obtained pose is not limited to a specific number of keypoints 23 or links 23''''.
[0052] The objective of this invention is to achieve a flexible and efficient technical solution applicable to various embedded platforms, prioritizing fast inference and equalization accuracy. To achieve this, the method uses the YOLOv5s6-ti-lite model—a modified version with layer tuning for compatibility with TI TDA4VM hardware. The model complexity of this invention is advantageously reduced to enhance latency performance.
[0053] In the example, the first stage of pose estimation is pedestrian detection. The trained model of the person detection and pose estimation module 20 can be configured to detect eight classes: people, cars, trucks, bicycles, motorcycles, buses, traffic signs, and traffic lights. This allows for seamless integration with other modules that need to detect different classes, reducing latency compared to using multiple AI models. Of course, other classes or different numbers of classes are also possible. The inference of the person detection and pose estimation module 20 includes mechanisms to remove unreliable and inaccurate detections by applying constraints to the 22 dimensions of the bounding box, thereby ensuring reliable detection results.
[0054] To enhance detection reliability, the human detection and pose estimation module 20 uses two datasets: COCO and BDD100K. COCO, with 118K training images and 5K evaluation images, is used for both detection and pose estimation tasks. BDD100K has 70K training images and 10K evaluation images. By combining these datasets, our network achieves robust and accurate detection due to the comprehensive training environment.
[0055] The method of this invention focuses on acquiring a sufficient set of image-based features for each pedestrian, ensuring robust state estimation without additional cost. By performing pose estimation 230, our system leverages the interconnections between joints 23, increasing the available features. Combining bottom-up and top-down approaches, we achieve both efficiency and accuracy. Employing a single inference strategy for pedestrian detection enables efficient handling of high-density scenes without increasing execution latency.
[0056] Figure 3 An example skeleton comprising thirteen joints 23 and fourteen links is shown. The pose configuration is determined by assigning weights 720 and incorporating the standard deviation into each keypoint 23 during coordinate regression. These weights indicate the importance of the joints and their contribution to the overall pose determination. The inclusion of the standard deviation accounts for the inherent variability and potential inaccuracies in joint position estimation.
[0057] The human detection and pose estimation module is also configured to merge the COCO and BDD100K datasets, remove neck, hand, and foot keypoints from BDD100K, and add new head keypoints in COCO using triangulation of eye, ear, and nose keypoints. Keypoint 23 is classified into three types based on occlusion level (clear, partially occluded, and completely occluded) to effectively balance training weights.
[0058] This method includes steps to generate a pedestrian pose history generated by a human detection and pose estimation module. It also includes steps to determine motion characteristics such as the position, velocity, and acceleration of key points.
[0059] The method includes the step of obtaining predicted location data for key points. In the example, the predicted location data is obtained at least in the following ways: - The estimated state of keypoints is estimated based on movement characteristics and posture history. Since keypoints are connected to form a skeletal representation of a person (e.g., a pedestrian), keypoints are connected to adjacent keypoints, where each keypoint can be an adjacent keypoint of another keypoint.
[0060] - Estimate the preliminary future state of the key point based on the estimated state of the key point and the future states of adjacent key points.
[0061] - Correct the estimated state of key points or the data obtained from the estimate by performing a Kalman filter as described below, such as a preliminary future position.
[0062] The pedestrian tracking module 30 includes a Kalman filter 400 (e.g., a first Kalman filter) for its effectiveness in achieving accurate and computationally efficient tracking. This method allows for the association of past predictions with current measurements, thereby achieving stable paths and unique identifiers 430 for each pedestrian. Although more advanced deep learning solutions are available, the Kalman filter 400 is chosen to maintain real-time tracking performance without sacrificing accuracy by tracking only predefined points (e.g., the centroids of the bounding boxes) of detected pedestrians to reduce the complexity of the motion model. In this example, the predefined points of the bounding boxes may be reference points.
[0063] A Kalman filter (e.g., a first Kalman filter) is used to predict the future positions of pedestrians (e.g., the positions of predefined points in the bounding box), and these predictions are correlated with the corresponding detections. The Kalman filter is a widely used method for estimating the state of objects in a system affected by measurement noise. In our method, the reference position of each pedestrian is defined as the centroid of the bounding box surrounding them.
[0064] The first Kalman filter in the tracking prediction module 30 uses a state vector to represent the current predicted state of each keypoint 23. The Kalman filter algorithm uses this state vector to iteratively refine the prediction based on new measurements, thereby ensuring robust and accurate tracking of pedestrian poses in the system.
[0065] The association process involves evaluating the proximity between new detections and existing tracked entities. Bounding boxes of closely tracked detections are associated using matching with corresponding entities, updating the state matrix, and predicting future locations. The Hungaris algorithm optimizes the matching between entities and detections, ensuring accurate association in dense pedestrian environments, such as... Figure 4 As shown. Unrelated detection creates new entities, and the tracking module is continuously adapted to merge the newly detected entities. Lost frames are added for unrelated detection, and if the count exceeds sixty, equivalent to a duration of two seconds at a frame rate of 30fps, the entity is considered lost.
[0066] The location prediction module 40 integrates pedestrian information from pose estimation and tracking to predict future pedestrian positions. Therefore, the location prediction module 40 is configured to receive the estimated position of each keypoint 23 of the proposed pose. The predicted pose is processed, a spatial reference is established, and area filters and scaling constraints are applied to ensure accurate prediction while mitigating potential errors.
[0067] Throughout the processing, the Kalman filter in the future prediction module 40 uses a state vector to represent the predicted current state. This vector combines the position, velocity, and acceleration components of each coordinate, allowing for accurate capture of the dynamics of pedestrian movement. The Kalman filter algorithm leverages this state vector to iteratively refine the prediction based on new measurements, ensuring robust and accurate tracking of pedestrians in the system.
[0068] The position prediction module 40 includes a predictive Kalman filter for each pose (e.g., pedestrian pose) based on the movement of each keypoint. This allows for enhanced pedestrian prediction, improving system accuracy and reliability. The predictive Kalman filter considers image coordinates, providing position (i.e., position on the horizontal and vertical axes), and provides velocity and acceleration estimates in each iteration. Personalized processing for each pedestrian ensures accurate motion tracking while minimizing interference from other detected pedestrians.
[0069] In the described example, the predictive Kalman filter effectively handles all 13 keypoints of the pose. This involves parallel operations with large matrices. As shown in Equation (1), the state vector Xn comprises 78 terms, which include the 2D coordinates of the 13 keypoints and their acceleration and velocity values.
[0070] .
[0071] A solution is proposed that considers a longer time interval than that between consecutive samples to generate the transition matrix for prediction. As shown in Equation (2), the location prediction module 40 generates the state transition matrix (F... prediction The value is approximated as an Nth power to calculate the prediction.
[0072] .
[0073] Once the state transition matrix (F) for future times is obtained prediction If a new incoming detection is received, it is multiplied by the current state vector (Xn) to obtain a prediction for each keypoint at a given time. When a new incoming detection is received, the Kalman model is updated to accommodate the new conditions for each keypoint 23.
[0074] .
[0075] To achieve accurate pedestrian location prediction, a minimum observation time is set before initiating the prediction. The observation time is calculated separately for each keypoint 23, and only valid observations with a confidence level higher than a threshold (default 20%) are considered to avoid abrupt changes between detections (which lead to deviations in the motion model of the final pose).
[0076] To enhance the accuracy and reliability of pedestrian motion prediction, we propose scaling the uncertainty of keypoints and enhancing the Kalman gain based on measurement confidence. Furthermore, we propose adjusting the measurement covariance matrix by considering the covariance between adjacent keypoints in the proposed pose. This innovative strategy tailors a conventional Kalman filter for skeleton scenes, prioritizing interactions between joints for improved motion prediction, particularly in cases involving occlusion or outliers.
[0077] In short, this invention is a novel method for customizing conventional Kalman filters for skeleton scenes, where the interconnections between joints play a crucial role in motion prediction. Therefore, once the estimated state of each keypoint (or any data derived from it, such as the corrected state described above) has been obtained, the estimated states of the keypoints are used to predict the future position of a reference keypoint (e.g., the centroid of the pose, which will be considered the centroid of the midpoint of the future pedestrian). This reference point is considered the reference point for processing information and the output of the prediction algorithm.
[0078] Development focuses on the TDA4VM processor designed by TI, targeting various computer vision applications, particularly in the automotive sector. The TDA4VM processor integrates a Dual Arm® Cortex® A72, C7x DSP, and multimedia accelerator.
[0079] Once validated, the AI model was deployed to the TI deep learning library and optimized for inference on the C7x+MMA DSP. Latency was improved by using mixed-precision quantization, 8-bit precision for some layers, and 16-bit precision for critical layers (the first and last). This optimization resulted in a model speedup of up to 2.2x compared to the base version with 16-bit precision, achieving inference speeds of up to 101fps.
[0080] The module implementation targets TIOVX deployment, using TI's OpenVX v1.1 specification framework. Each submodule acts as a black-box component, allowing dynamic allocation to various hardware targets and supporting multiprocessing. Supported hardware includes two DSP C66x, one DSP C7X, two MCU R5F, dual-arm Cortex-A72, and one GPU PowerVR.
[0081] This invention provides high accuracy in predicting the future location of pedestrians, ensuring the real-time execution of the method.
[0082] In the example, the well-known multi-object tracking KITTI dataset was chosen for its focus on ADAS and its higher refresh rate (10 FPS) compared to other datasets. It contains 21 training sequences, which have been labeled with 8 distinct classes, including pedestrians, which are the only training sequences to be considered.
[0083] The tracking method was evaluated using the High-Order Tracking Accuracy (HOTA) metric, which combines localization, detection, and association into a single metric. To assess its performance, the tracking algorithm was compared to other real-time items from the top 50 MOT data sets from the KITTI dataset, and the results are presented in Table \ref{tab:trackingComp}.
[0084] The OC-SORT algorithm uses a single camera, but it is computationally demanding and unsuitable for embedded platforms. Opm-NC2 and NC2 algorithms achieve high accuracy in object detection, but require expensive setups with LiDAR data. CAT detects pedestrians from image deep learning models and then matches them in 3D with a stereo camera. RMOT and NOMT-HM have similar 2D camera setups to our approach, but focus on modeling the relative motion between objects.
[0085] As evaluated by the LocA metric, our proposed solution demonstrates superior detection accuracy compared to other existing methods. It also outperforms alternatives with similar sensor setups in terms of HOTA metrics while maintaining efficient computational requirements. These results validate the effectiveness of our method for achieving real-time inference on resource-constrained platforms.
[0086] In short, the novel system and method are configured to detect, estimate pose, and predict the future location of pedestrians within the context of ADAS. To achieve this, it comprises a two-stage processing: the first stage consists of a model network for detecting and estimating pedestrian pose, and the second stage is based on a Kalman filter with pose features to predict the estimate for each pedestrian.
[0087] This invention has been designed for implementation in automotive embedded platforms, achieving high-value metrics relative to prior art methods using an equivalent sensor configuration without compromising execution latency on PCs and embedded TDA4VM platforms. This design provides flexible deployment capabilities for integration into existing systems without significant changes to the required setup and computing platform.
[0088] This invention provides a robust and affordable system designed for short-term deployment because its integrated approach is a viable alternative to the current state of the automotive industry.
[0089] Figure 6 The steps according to an example of the present invention are illustrated. As shown in the figure, this non-limiting example of a method for predicting the future location of an object includes the following steps: - A video stream 100, comprising multiple sequential two-dimensional (2D) image data, is captured by vision cameras deployed in the vehicle; - Image distortion is corrected by the image dedistortion module 10 for multiple sequential two-dimensional (2D) image data; - Perform a single step of the deep learning model 200 for: (i) Obtain the detection of object 21 in bounding box 22, where object 21 is a living being; (ii) Obtain the reference point of bounding box 220 (specifically, located at the center of bounding box 22), and (iii) Obtaining the pose of an object comprising a plurality of keypoints 23, the plurality of keypoints 23 comprising two or more keypoints 23', 23'' connected to each other by links 23'''' and adjacent keypoints 23''' connected to at least one of the two or more keypoints 23', 23'', wherein each keypoint 23 is or represents a joint of a living organism, and optionally, the keypoints in the plurality of keypoints 23 are connected to form a skeletal representation of a living organism; The method further includes the step of classifying the object 21 300 such that the object 21 is assigned to a specific object class. Therefore, if multiple (different) objects 21 are detected on two-dimensional (2D) image data, the method can be configured to classify the (different) objects 21 into different classes such that each of the (different) objects 21 is assigned to a specific object class. - Compare the detection of object 410 with the previous detection of the object; - For example, by using a first Kalman filter 400 to track the detection of 420 objects in order to track reference points through multiple sequential two-dimensional (2D) image data; - The identification code (UID) is assigned to the object for tracking and detection by using the first Kalman filter 400; - Based on the tracked detection identifier (UID), the pose history of 500 objects is generated in real time from multiple sequential two-dimensional (2D) image data; - Determine the movement characteristics 600 (e.g., position, velocity, and optional acceleration) of key points 23 of an object (e.g., two or more key points 23', 23'' and adjacent key points 23'''); - Obtain predicted state data 700 for at least each of two or more keypoints 23', 23'' in the two-dimensional (2D) image data of the video stream by means of at least the following: (a) Estimating the estimated state of 710 keypoints (e.g., each of two or more keypoints 23', 23'', and optionally adjacent keypoints 23''') based on movement characteristics and pose history. For clarity, the estimated state of adjacent keypoints is referred to herein as the future state of adjacent keypoints; (b) Optionally, the confidence of the estimated state of the 720 key points is determined by assigning weights to the estimated state of each key point, such that each key point is designated as more important based on the confidence of the estimated state. (c) Estimate the preliminary future state of 730 key points (e.g., each of two or more key points) based on the estimated state or any data obtained from the estimated state and the future states of adjacent key points; (d) Correct the estimated state or any data obtained from the estimated state (such as the preliminary future state) to obtain the corrected position of each of two or more key points, for example, by performing a second Kalman filter; - The future positions of reference points for objects 86, 87, and 88 are predicted based on predicted state data. These future positions are predicted in two-dimensional (2D) image data of the video stream, where the future positions of reference points 86, 87, and 88 are the predicted positions of the reference points at a predefined future time. The predefined future time is between 1 second and 3 seconds. Specifically, the future positions 86, 87, and 88 are predicted by the position prediction module 40. The position prediction module 40 receives predicted state data, which includes the state of a first keypoint 23' out of two or more keypoints and the state of a second keypoint 23'' out of two or more keypoints. Furthermore, the state of the first keypoint 23' includes its position on the horizontal axis, its position on the vertical axis, its velocity on the horizontal axis, its velocity on the vertical axis, and optionally, its acceleration on the horizontal axis and its acceleration on the vertical axis. Furthermore, the state of the second keypoint 23'' includes position on the horizontal axis, position on the vertical axis, velocity on the horizontal axis, velocity on the vertical axis, and optionally acceleration on the horizontal axis and acceleration on the vertical axis. More specifically, the position prediction module 40 includes a predictive Kalman filter configured to receive predicted state data, such as position, velocity, and acceleration, of two or more keypoints 23', 23'' of an object in two-dimensional (2D) image data. The predicted state data is arranged in the form of a state vector; and - Display the predicted future positions 86, 87, 88 of 900 reference points via a display device and / or input the predicted future positions 86, 87, 88 of the reference points into an ADAS system configured, for example, to prevent a collision between the detected object 21 and the vehicle and / or to activate an emergency braking mechanism.
[0090] Although only a few examples of the electronic device and its assembly methods have been disclosed herein, other alternatives, modifications, uses, and / or equivalents are possible. Furthermore, all possible combinations of the described examples are covered. Therefore, the scope of this disclosure should not be limited to any particular example, but should be determined solely by a fair reading of the appended claims. Reference numerals associated with the drawings placed in brackets within the claims are used only to attempt to increase the comprehensibility of the claims and should not be construed as limiting the scope of the claims.
Claims
1. A method for predicting the future location of an object, the method comprising: - Capture video streams that include two-dimensional (2D) image data; - Detect objects in the two-dimensional (2D) image data; - Obtain the pose of the object, including multiple key points; - Generate the pose history of the object; - Determine the movement characteristics of two or more key points of the object; - Predicted state data for each of the two or more key points is obtained by estimating the estimated state of each of the two or more key points based on the movement characteristics and the attitude history; as well as - Predict the future location of a reference point of the object from the predicted state data, wherein the future location is predicted on two-dimensional (2D) image data of the video stream.
2. The method according to claim 1, wherein, The future position of the reference point is the position that the reference point is predicted to have at a predefined future time, preferably, the predefined future time is between 1 second and 3 seconds.
3. The method according to any of the preceding claims, wherein, Detecting the object and obtaining the pose are achieved in a single step of the deep learning model.
4. The method according to any of the preceding claims, wherein, The method further includes the step of classifying the object so that the object is assigned to a specific object class.
5. The method according to any of the preceding claims, wherein, The pose history of the object is generated in real time.
6. The method according to any of the preceding claims, wherein, The motion characteristics include the position, velocity, and acceleration of each of the two or more keypoints.
7. The method according to any of the preceding claims, wherein, The object is associated with a bounding box, wherein the two or more key points are movable within the bounding box, and the reference point is fixed within the bounding box, the reference point being a predefined representative point of the bounding box, the predefined representative point being preferably located at the center of the bounding box.
8. The method according to any of the preceding claims, wherein, The object is a living being, such as a person, preferably a pedestrian.
9. The method according to claim 8, wherein, The key points among the plurality of key points are connected to form a skeletal representation of a living organism, wherein the plurality of key points includes two or more key points and adjacent key points, such that the two or more key points are connected to the adjacent key points.
10. The method according to claim 9, wherein, The step of obtaining the predicted state data further includes: estimating the future state of the adjacent key points, wherein the predicted state data includes the preliminary future state of each of the two or more key points obtained by the estimated state of each key point and the future state of the adjacent key points.
11. The method according to claim 10, wherein, The step of obtaining the predicted state data further includes: correcting the estimated state or the preliminary future state to obtain the corrected position of each of the two or more key points.
12. The method according to claim 11, wherein, The corrected estimated state or the preliminary future state is obtained by performing a Kalman filter.
13. The method according to any of the preceding claims, wherein, The step of obtaining the predicted state data further includes: determining the confidence level of the estimated state, wherein the preliminary future state is also based on the confidence level of the estimated state.
14. The method according to claim 13, wherein, The step of obtaining the predicted location data further includes: assigning weights to the estimated state, wherein the preliminary future state is also based on the weights of the estimated state.
15. A system for predicting the future location of an object by means of the method according to any of the preceding claims.
Citation Information
Patent Citations
Recognizing the movement intention of a pedestrian from camera images
US20210150193A1