Three-dimensional space track identification method and system based on multi-view three-dimensional reconstruction
Through multi-view three-dimensional reconstruction technology, multiple cameras and internal and external parameters of the camera are used to recognize and reconstruct pedestrian three-dimensional trajectories, solving the problem that single-camera two-dimensional images is difficult to provide depth information, and improving the accuracy and adaptability of trajectory recognition.
Patent Information
- Application Number
- CN202510008730.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-03
AI Technical Summary
The existing pedestrian trajectory recognition technology is based on two-dimensional images of a single camera, which is difficult to provide depth information, and it is difficult to handle occlusion and viewing angle changes, resulting in tracking errors and trajectory interruptions.
Multi-view three-dimensional reconstruction technology is adopted to capture the same scene from different angles through multiple cameras, combining the camera's internal and external parameters and multi-view triangulation method to realize the trajectory recognition and reconstruction of pedestrians in three-dimensional space.
The accurate identification and reconstruction of pedestrian three-dimensional trajectories is realized, the labor and calculation costs are reduced, and the dependence on known sizes or three-dimensional structures is avoided, and it has strong versatility and adaptability.
Smart Images

Figure CN119941801A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the research field of pedestrian travel characteristics, and specifically relates to a three-dimensional space trajectory recognition method and system based on multi-view stereo reconstruction. Background Art
[0002] With the rapid development of intelligent monitoring, security and urban management, the study and analysis of pedestrian travel behavior has gradually become an important application demand. Traditional pedestrian trajectory recognition technology usually performs target detection and tracking based on the two-dimensional image of a single camera, but these methods have some limitations. The main problems include lack of depth information, difficulty in dealing with occlusion and perspective changes. In order to overcome these problems, in recent years, three-dimensional space trajectory recognition methods based on multi-view stereo reconstruction technology have gradually become a research hotspot.
[0003] At present, many pedestrian trajectory recognition systems rely on the planar image data of a single camera and use target detection and tracking algorithms (such as YOLO and SORT) to extract pedestrian trajectories. However, this method can only provide two-dimensional trajectories, lacks accurate capture of object depth information, and is difficult to accurately reflect the position and motion trajectory of objects in actual three-dimensional space. Especially in complex environments and high-density scenes, two-dimensional trajectories may be affected by occlusion and changes in perspective, resulting in tracking errors and trajectory interruptions.
[0004] In order to solve these problems, stereo reconstruction methods based on multi-camera perspectives have received widespread attention in recent years. By using multiple cameras to shoot the same scene from different angles, rich disparity data can be obtained, thereby achieving more accurate 3D space reconstruction. In particular, through the camera's internal and external parameters and multi-view triangulation technology, the precise position of the object in the 3D space can be restored, thereby constructing a more realistic 3D trajectory.
[0005] However, existing multi-view trajectory recognition methods still face some technical challenges, such as how to effectively handle image registration between different viewpoints, feature matching accuracy, and how to convert two-dimensional feature points into three-dimensional coordinates. With the development of computer vision technology and deep learning algorithms, multi-view stereo reconstruction and three-dimensional trajectory recognition methods are gradually overcoming these technical difficulties and showing broad application prospects in multiple fields. Summary of the invention
[0006] The purpose of the present invention is to provide a three-dimensional space trajectory recognition method based on multi-view stereo reconstruction, which can accurately recognize and reconstruct the three-dimensional trajectory of pedestrians through the view information of multiple cameras. This method can reduce manpower and computing costs, while avoiding dependence on objects of known size or known three-dimensional structure, and has strong versatility and adaptability.
[0007] The technical solution of the present invention is as follows:
[0008] A three-dimensional space trajectory recognition method based on multi-view stereo reconstruction, the method comprising the following steps:
[0009] Use multiple cameras to shoot the same scene from different angles, and determine the intersection of the images of multiple cameras as the research scene;
[0010] Perform frame extraction and pedestrian detection on the video of each camera, identify the pedestrians and their location information in each frame of video, and assign a unified identity number to the same pedestrian;
[0011] The video frames of different cameras are registered to obtain the pixel position of each pedestrian in each camera. The three-dimensional point coordinates of each pedestrian in the camera coordinate system are solved by multi-view triangulation method combined with the internal and external parameters of the camera.
[0012] Convert the three-dimensional point coordinates into coordinates in the world coordinate system to obtain the world coordinate system coordinates of each numbered pedestrian point in each frame;
[0013] According to the time information, the numbered pedestrian points in each frame are connected to form a trajectory, completing the reconstruction of the pedestrian's three-dimensional spatial trajectory in the world coordinate system.
[0014] Furthermore, the video of each camera is subjected to frame extraction and pedestrian detection, pedestrians and their position information in each frame of video are identified, and a unified identity number is assigned to the same pedestrian. Specifically,
[0015] First, the YOLOv8 model is used to perform target detection and instance segmentation on pedestrians in each frame of the video to obtain the bounding box and segmentation mask of each human body. Then, the detected bounding box is input into the BoT-SORT model for pedestrian tracking. The BoT-SORT model assigns a unified identity number to each pedestrian in the entire video sequence based on their position and motion trajectory.
[0016] Furthermore, assigning a unified identity number to the same pedestrian also includes matching pedestrian features in different cameras.
[0017] Furthermore, the video frame images of different cameras are registered to obtain the pixel position of each pedestrian in each camera, and the three-dimensional point coordinates are solved by multi-view triangulation method in combination with the internal and external parameters of the camera:
[0018] First, the two cameras are calibrated using camera calibration software and calibration plates to obtain the camera's intrinsic and extrinsic parameters, thereby determining the camera's position and posture in the world coordinate system.
[0019] Then, feature points are extracted within the human body bounding box in the two view images using a feature extraction algorithm, and the same feature points in the images from different view images are matched using a feature matching algorithm;
[0020] Next, according to the position of the feature point in each camera image and the camera internal and external parameters, a spatial ray passing through the feature point is constructed; the intersection of the spatial ray in the three-dimensional space is the three-dimensional point coordinate of the object.
[0021] Furthermore, the intrinsic parameters of the camera include focal length, principal point coordinates, pixel ratio coefficient and distortion coefficient.
[0022] Furthermore, the camera extrinsic parameters include a rotation matrix and a translation vector.
[0023] Furthermore, the three-dimensional points are converted into coordinates in the world coordinate system, and the world coordinate system coordinates of the numbered pedestrian points in each frame are obtained as follows:
[0024] The camera's rotation matrix and translation vector are used to transform the three-dimensional point coordinates in the camera coordinate system to obtain the world coordinate system coordinates of each numbered pedestrian point in each frame.
[0025] The present invention also provides a three-dimensional space trajectory recognition system based on multi-view stereo reconstruction, wherein the system comprises:
[0026] Multi-camera acquisition module: used to shoot the same scene from different angles using multiple cameras, and determine the intersection of the images of multiple cameras as the research scene;
[0027] Video processing and pedestrian detection module: used to extract frames from the video of each camera, use pedestrian detection algorithm to identify pedestrians and their location information in each frame of video, and assign a unified identity number to each pedestrian;
[0028] Video registration and feature point extraction module: used to register video frame images from different cameras and extract the pixel position of each pedestrian in each camera;
[0029] 3D reconstruction and depth calculation module: Based on the multi-view triangulation method and combined with the internal and external parameters of the camera, the 3D point coordinates of each pedestrian in the camera coordinate system are calculated;
[0030] Coordinate conversion module: used to convert the coordinates of the three-dimensional point from the camera coordinate system into coordinates in the world coordinate system, and obtain the world coordinate system coordinates of each pedestrian in each frame;
[0031] Trajectory reconstruction module: According to the time information, the numbered pedestrian points in each frame are connected to form the pedestrian's three-dimensional spatial trajectory, completing the reconstruction of the pedestrian's three-dimensional spatial trajectory in the world coordinate system.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] Compared with a single camera perspective, the present invention can cover a larger area, reduce occlusion, provide more parallax data, and improve the accuracy of object detection and depth estimation in complex scenes;
[0034] The present invention can identify pedestrian travel behavior and reconstruct the trajectory in the world coordinate system so that the trajectory has real spatial information; subsequently, combined with three-dimensional modeling, the characteristics of pedestrian travel trajectories and influencing factors can be intuitively analyzed. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] The accompanying drawings generally illustrate various embodiments by way of example and not limitation, and together with the description and claims, serve to illustrate the embodiments of the invention. Where appropriate, the same reference numerals are used throughout the drawings to refer to the same or similar parts. Such embodiments are illustrative and are not intended to be exhaustive or exclusive embodiments of the present apparatus or method.
[0036] Figure 1 A schematic diagram of the process of the present invention is shown;
[0037] Figure 2 A schematic diagram of three-dimensional reconstruction of the present invention is shown. DETAILED DESCRIPTION
[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0039] like Figure 1-Figure 2 As shown, an embodiment of the present invention provides a three-dimensional space trajectory recognition method based on multi-view stereo reconstruction, and the method includes the following steps:
[0040] Use multiple cameras to shoot the same scene from different angles, and determine the intersection of the images of multiple cameras as the research scene;
[0041] Perform frame extraction and pedestrian detection on the video of each camera, identify the pedestrians and their location information in each frame of video, and assign a unified identity number to the same pedestrian;
[0042] The video frames of different cameras are registered to obtain the pixel position of each pedestrian in each camera. The three-dimensional point coordinates of each pedestrian in the camera coordinate system are solved by multi-view triangulation method combined with the internal and external parameters of the camera.
[0043] Convert the three-dimensional point coordinates into coordinates in the world coordinate system to obtain the world coordinate system coordinates of each numbered pedestrian point in each frame;
[0044] According to the time information, the numbered pedestrian points in each frame are connected to form a trajectory, completing the reconstruction of the pedestrian's three-dimensional spatial trajectory in the world coordinate system.
[0045] The technology used for pedestrian detection is: using a preset tracking model to perform multi-target detection and instance segmentation on the video captured by the camera, and obtaining the identity number, location information, bounding box and segmentation mask of each pedestrian in each frame of the picture through frame-by-frame recognition and tracking;
[0046] The tracking model is the YOLOv8+BoT-SORT model. The YOLOv8 model first performs target detection and instance segmentation on the human body in the video frame to obtain the bounding box and segmentation mask of each human body. The human body bounding box is then sent to the BoT-SORT model. The BoT-SORT model tracks the detected human body and assigns a consistent human identity number to the same human body, thereby ensuring coherent and stable tracking of the target throughout the video sequence and accurately depicting its motion trajectory.
[0047] The technology used to unify the pedestrian ID numbers in the two cameras is: multi-camera pedestrian re-identification; extracting pedestrian visual features through a deep learning model; using metric learning technology to train the network to bring the feature vectors of the same object closer and distinguish the feature vectors of different objects;
[0048] The process of calculating the world coordinate system of the human body point is:
[0049] Use camera calibration software and calibration board to calibrate the two cameras respectively to obtain camera intrinsic parameters and extrinsic parameters; camera intrinsic parameters include focal length, principal point coordinates, pixel ratio coefficient, distortion coefficient; camera extrinsic parameters include rotation matrix and translation vector. Determine the coordinates of the two cameras in the world coordinate system.
[0050] Find the feature points of the target object in the two view images and match them in different images; specifically, use the feature extraction algorithm to extract obvious features such as corners and edges within the human body boundary box; use the feature matching algorithm to match the same feature points in the multi-view images;
[0051] The basic matrix and essential matrix of the two cameras are calculated through the feature point pairs in the two-view images; the feature points in the image are restored to three-dimensional space through multi-view triangulation to obtain the three-dimensional point coordinates of each human body; specifically, according to the position of the feature point in each camera image and the internal and external parameters of the camera, a spatial ray passing through the feature point is constructed; these rays should intersect at a certain point in the three-dimensional space, which is the three-dimensional coordinate of the object.
[0052] Using the rotation and translation matrices of the camera, the three-dimensional point coordinates are converted into a unified world coordinate system to obtain the world coordinate system coordinates of the human body in each video frame.
[0053] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A three-dimensional space trajectory recognition method based on multi-view stereo reconstruction, characterized in that: The method comprises the following steps: Use multiple cameras to shoot the same scene from different angles, and determine the intersection of the images of multiple cameras as the research scene; Perform frame extraction and pedestrian detection on the video of each camera, identify the pedestrians and their location information in each frame of video, and assign a unified identity number to the same pedestrian; The video frames of different cameras are registered to obtain the pixel position of each pedestrian in each camera. The three-dimensional point coordinates of each pedestrian in the camera coordinate system are solved by multi-view triangulation method combined with the internal and external parameters of the camera. Convert the three-dimensional point coordinates into coordinates in the world coordinate system to obtain the world coordinate system coordinates of each numbered pedestrian point in each frame; According to the time information, the numbered pedestrian points in each frame are connected to form a trajectory, completing the reconstruction of the pedestrian's three-dimensional spatial trajectory in the world coordinate system.
2. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 1 is characterized in that: The process of extracting frames and detecting pedestrians from the video of each camera, identifying pedestrians and their location information in each frame of video, and assigning a unified identity number to the same pedestrian is as follows: First, the YOLOv8 model is used to perform target detection and instance segmentation on pedestrians in each frame of the video to obtain the bounding box and segmentation mask of each human body. Then, the detected bounding box is input into the BoT-SORT model for pedestrian tracking. The BoT-SORT model assigns a unified identity number to each pedestrian in the entire video sequence based on their position and motion trajectory.
3. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 1 is characterized in that: The step of assigning a uniform identity number to the same pedestrian also includes matching features of the pedestrians in different cameras.
4. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 1 is characterized in that: The video frames of different cameras are registered to obtain the pixel position of each pedestrian in each camera. The three-dimensional point coordinates are solved by multi-view triangulation method combined with the internal and external parameters of the camera: First, the two cameras are calibrated using camera calibration software and calibration plates to obtain the camera's intrinsic and extrinsic parameters, thereby determining the camera's position and posture in the world coordinate system. Then, feature points are extracted within the human body bounding box in the two view images using a feature extraction algorithm, and the same feature points in the images from different view images are matched using a feature matching algorithm; Next, according to the position of the feature point in each camera image and the camera internal and external parameters, a spatial ray passing through the feature point is constructed; the intersection of the spatial ray in the three-dimensional space is the three-dimensional point coordinate of the object.
5. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 4 is characterized in that: The internal parameters of the camera include focal length, principal point coordinates, pixel ratio coefficient and distortion coefficient.
6. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 4 is characterized in that: The camera extrinsic parameters include a rotation matrix and a translation vector.
7. The three-dimensional space trajectory recognition method based on multi-view stereo reconstruction according to claim 1 is characterized in that: The three-dimensional points are converted into coordinates in the world coordinate system, and the world coordinate system coordinates of the numbered pedestrian points in each frame are obtained as follows: The camera's rotation matrix and translation vector are used to transform the three-dimensional point coordinates in the camera coordinate system to obtain the world coordinate system coordinates of each numbered pedestrian point in each frame.
8. A three-dimensional space trajectory recognition system based on multi-view stereo reconstruction, characterized in that: The system comprises: Multi-camera acquisition module: used to shoot the same scene from different angles using multiple cameras, and determine the intersection of the images of multiple cameras as the research scene; Video processing and pedestrian detection module: used to extract frames from the video of each camera, use pedestrian detection algorithm to identify pedestrians and their location information in each frame of video, and assign a unified identity number to each pedestrian; Video registration and feature point extraction module: used to register video frame images from different cameras and extract the pixel position of each pedestrian in each camera; 3D reconstruction and depth calculation module: Based on the multi-view triangulation method and combined with the internal and external parameters of the camera, the 3D point coordinates of each pedestrian in the camera coordinate system are calculated; Coordinate conversion module: used to convert the coordinates of the three-dimensional point from the camera coordinate system into coordinates in the world coordinate system, and obtain the world coordinate system coordinates of each pedestrian in each frame; Trajectory reconstruction module: According to the time information, the numbered pedestrian points in each frame are connected to form the pedestrian's three-dimensional spatial trajectory, completing the reconstruction of the pedestrian's three-dimensional spatial trajectory in the world coordinate system.
Citation Information
Patent Citations
Multi-view gait identification method based on adaptive three dimensional human motion statistic model
CN106056050A
Intelligent vehicle track measurement method based on binocular stereo vision system
CN110285793A
Cross-multi-camera pedestrian track recognition method
CN111027462A
Multi-camera pedestrian recognition method and recognition system
CN111079600A
Space target motion state identification method based on cooperative observation image sequence
CN112508999A
Cited By
Three-dimensional tennis track real-time reconstruction method and system based on multiple cameras
CN120997256A
Three-dimensional model construction method and device based on image sequence processing
CN121170166A
Three-dimensional scene reconstruction data set generation method based on multi-view first-person video
CN121259175A
Three-dimensional space groove track generation method based on multi-view image
CN121304944A
A three-dimensional space groove track generation method based on multi-view images
CN121304944B