Multi-person collision detection method and system for VR large-space immersive sightseeing
By acquiring RGB and depth images in VR devices for instance segmentation and feature extraction, constructing factor graphs for residual operations, optimizing object poses and motion trajectories, solving the real-time and accuracy of multi-person collision detection in large-space immersive VR tours, and improving the security and immersion of VR experience.
Patent Information
- Application Number
- CN202510408176.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-02
AI Technical Summary
In large-space immersive VR tours, multi-person collision detection is difficult to monitor and deal with in real time, resulting in a decrease in safety hazards and experience quality, especially when multiple VR exploration teams are active at the same time.
RGB images and depth images are obtained through VR devices, instance segmentation and feature extraction are performed, and a factor diagram is constructed for residual operations, the object pose and motion trajectory are optimized, and multiple collision detection is finally performed.
It realizes high-precision and real-time multi-person collision detection, ensuring that the system accurately estimates object movement in complex dynamic environments, improving the security and immersion of the VR experience.
Smart Images

Figure CN120355751A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence, computer vision, and computer graphics, and particularly relates to a multi-person collision detection method and system for VR large-space immersive tours. Background Art
[0002] In large-space immersive tour experiences, it is usually necessary to form a group of 3 to 5 people to conduct team VR exploration within a specific LBE (Location-Based Entertainment) venue. However, during the tour, participants may encounter various collision problems, such as collisions between people, collisions between people and walls or props, etc. These problems not only pose safety hazards but also affect the quality of the VR experience. Especially in the case of a large number of tourists, multiple VR exploration teams will be active on the same LBE venue at the same time, further increasing the risk of collisions.
[0003] To solve the above problems and ensure that each participant can safely and happily enjoy the VR experience, it is particularly important to monitor and process the postures and movement trajectories of all dynamic objects in the venue in real time. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a multi-person collision detection method and system for VR large-space immersive tours to solve the problems existing in the above prior art.
[0005] To achieve the above object, the present invention provides a multi-person collision detection method for VR large-space immersive tours, including:
[0006] Obtain an RGB image and a depth image through a VR device, and obtain an instance segmentation image based on the RGB image;
[0007] Extract features based on the instance segmentation image and the depth image to obtain a static optical flow image, static 3D feature points, a dynamic optical flow image, and dynamic 3D feature points;
[0008] Obtain an initial camera pose based on the static optical flow image and the static 3D feature points; calculate the motion residuals between consecutive frames on each dynamic object based on the dynamic 3D feature points; obtain an initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals;
[0009] Construct a factor graph based on the initial camera pose, the initial object motion, the static 3D feature points, and the dynamic 3D feature points; perform residual operations based on the factor graph, and obtain the postures and movement trajectories of each dynamic object based on the results of the residual operations, the initial camera pose, and the initial object motion;
[0010] Perform multi-person collision detection based on the postures and motion trajectories of each dynamic object.
[0011] Optionally, the process of feature extraction based on the instance segmentation image and the depth image includes:
[0012] Select a feature detection algorithm based on the environmental map, perform sparse feature point detection on the instance segmentation image based on the feature detection algorithm, and screen through the adaptive non-maximum suppression algorithm. Then, use the optical flow tracking algorithm to match and track the screened feature points between consecutive frames to generate a static optical flow image, and back-project the 2D feature points on the static optical flow image into 3D feature points to obtain static 3D feature points;
[0013] Use the instance segmentation mask image to distinguish dynamic objects from the static background, assign unique instance labels to the dynamic objects, perform dense feature point detection and screening, use the optical flow tracking algorithm to generate a dynamic optical flow image, and perform feature point back-projection to obtain dynamic 3D feature points.
[0014] Optionally, the process of obtaining the initial camera pose based on the static optical flow image and the static 3D feature points includes:
[0015] Construct the correspondence between the 2D static feature points at the current time step and the static 3D feature points at the previous time step, perform a preliminary estimation of the camera pose by minimizing the reprojection error through the PnP algorithm, and verify the estimation result; Combine the static optical flow information with the verified estimation result, and obtain the initial camera pose through non-linear least squares optimization.
[0016] Optionally, the motion residuals between consecutive frames on each dynamic object are calculated as follows:
[0017]
[0018] where represents the position coordinates of the 3D feature points in the world coordinate system at time step τ, represents the object motion from time step τ - 1 to τ.
[0019] Optionally, the process of obtaining the initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals includes:
[0020] Build the correspondence between the 2D dynamic feature points at the current time step and the dynamic 3D feature points at the previous time step, perform a preliminary estimation of the dynamic object motion by minimizing the reprojection error using the PnP algorithm, and verify the estimation result; combine the dynamic optical flow information with the verified estimation result, optimize the initially estimated dynamic object motion through non-linear least squares, and perform a secondary optimization of the optimized dynamic object motion by minimizing the weighted sum of squares of the motion residuals between consecutive frames on each dynamic object to obtain the initial object motion.
[0021] Optionally, the process of building the factor graph includes:
[0022] Construct a factor graph with the camera pose, object motion, static 3D feature points, and dynamic 3D feature points as nodes, and the optical flow information and reprojection error as edge constraints.
[0023] Optionally, the residual operations include 3D point measurement residuals, camera motion residuals, 3D point position residuals, and object motion smoothness residuals.
[0024] Optionally, construct an objective function based on the residual operation results, and obtain the poses and motion trajectories of each dynamic object based on the objective function; the objective function is as follows:
[0025]
[0026] Where, represents the residual of the initial camera pose, represents the camera motion residual, represents the 3D point measurement residual, represents the 3D point position residual under object motion, represents the object motion smoothness residual; ∑0, ∑ 3D,R , represent the covariance matrices of the corresponding residuals respectively; ρ h represents the Huber function, represents the set of static 3D points at time step τ, represents the set of dynamic 3D points at time step τ, represents the set of dynamic 3D points on the j-th object at time step τ, represents the set of objects observed at time step τ.
[0027] Optionally, it is characterized in that it further includes: the current VR device performs contour rendering according to the poses and motion trajectories of each dynamic object, and after cross-verifying the data transmitted by multiple VR devices, the accurate poses and motion trajectories of all moving objects in the scene are transmitted back to the current VR device to update the participant's contour.
[0028] The present invention also provides a multi-person collision detection system for VR large-space immersive tours, including:
[0029] A VR device for acquiring RGB images and depth images, and obtaining an instance segmentation image based on the RGB images;
[0030] A feature extraction module for performing feature extraction based on the instance segmentation image and the depth image to obtain a static optical flow image, static 3D feature points, a dynamic optical flow image, and dynamic 3D feature points;
[0031] An attitude and motion estimation module for obtaining an initial camera attitude based on the static optical flow image and the static 3D feature points; calculating a motion residual between consecutive frames on each dynamic object based on the dynamic 3D feature points; and obtaining an initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residual;
[0032] A factor graph construction module for constructing a factor graph based on the initial camera attitude, the initial object motion, static 3D feature points, and dynamic 3D feature points;
[0033] An optimization module for performing residual calculations based on the factor graph, and obtaining the attitudes and motion trajectories of each dynamic object based on the results of the residual calculations, the initial camera attitude, and the initial object motion;
[0034] A collision detection module for performing multi-person collision detection based on the attitudes and motion trajectories of each dynamic object.
[0035] Compared with the prior art, the present invention has the following advantages and technical effects:
[0036] (1) The present invention directly parameterizes the motion and attitude of each object through new residual calculation and optimization formulas, ensuring that the constraints of rigid body kinematics are maintained. This formulaic approach enables the system to more accurately estimate the motion of objects, especially in complex dynamic environments.
[0037] (2) The present invention jointly optimizes the camera attitude, object motion, and environmental map through multiple joint optimization algorithms, ensuring the consistency and accuracy of the system state. This joint optimization strategy enables the system to better handle complex scenarios in dynamic environments.
[0038] (3) The present invention adopts a sliding window optimization, parallel optimization strategy for dynamic objects and static backgrounds, which can significantly reduce the computational amount and improve the real-time performance of the system; adopting a modular design enables the system to flexibly integrate different modules and optimization algorithms to adapt to different application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0040] Figure 1 It is a schematic diagram of the method according to an embodiment of the present invention;
[0041] Figure 2 It is a schematic diagram of the feature extraction process according to an embodiment of the present invention;
[0042] Figure 3 It is a schematic diagram of the algorithm optimization process according to an embodiment of the present invention. Detailed implementation manners
[0043] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.
[0044] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0045] Embodiment 1
[0046] As Figures 1-3 shown, in this embodiment, a multi-person collision detection method for VR large-space immersive tours is provided, including:
[0047] Obtain RGB images and depth images through VR devices, and obtain instance segmentation images based on the RGB images;
[0048] Perform feature extraction based on the instance segmentation images and the depth images to obtain static optical flow images, static 3D feature points, dynamic optical flow images, and dynamic 3D feature points;
[0049] Specifically, the process of performing feature extraction based on the instance segmentation images and the depth images includes:
[0050] Select a feature detection algorithm based on the environmental map, perform sparse feature point detection on the instance segmentation images based on the feature detection algorithm, and perform screening through an adaptive non-maximum suppression algorithm, and then use an optical flow tracking algorithm to match and track the screened feature points between consecutive frames to generate static optical flow images, and back-project the 2D feature points on the static optical flow images into 3D feature points to obtain static 3D feature points;
[0051] Use the instance segmentation mask image to distinguish dynamic objects from the static background, assign unique instance labels to the dynamic objects, perform dense feature point detection and screening, generate a dynamic optical flow image using the optical flow tracking algorithm, and perform back-projection of the feature points to obtain dynamic 3D feature points.
[0052] If the VR device has an RGB-D camera or a stereo camera, directly use this camera to capture the RGB image and the depth map If the VR device does not have an RGB-D camera or a stereo camera, it is also possible to enable this VR device to capture the depth map by connecting an external device Note that the superscript τ is the time step.
[0053] If the VR device neither has an RGB-D camera or a stereo camera nor can connect an external device or it is difficult to do so, first input the RGB image into a pre-trained "depth detection model" to estimate the depth map Common pre-trained "depth detection models" include:
[0054] Monodepth2, which predicts the depth map through a monocular image;
[0055] Depth estimation with a Deep Neural Network, which uses a deep neural network to estimate the depth from a monocular image;
[0056] RAFT-Depth, which is a depth estimation method based on the RAFT optical flow algorithm;
[0057] StereoNet, which estimates the depth map through a stereo image pair.
[0058] Input the RGB image into a pre-trained "instance segmentation model", and through "mask operation", the instance segmentation image can be obtained Common pre-trained "instance segmentation models" include:
[0059] Mask R-CNN: A popular instance segmentation network that can perform object detection and segmentation simultaneously.
[0060] DeepLab: A network for semantic segmentation that can also be used for instance segmentation.
[0061] YOLO series: A real-time object detection network that can be used for instance segmentation by combining a segmentation module.
[0062] Panoptic-DeepLab: A network for panoptic segmentation that can handle semantic segmentation and instance segmentation simultaneously.
[0063] The superscript τ in it represents the time step, that is, the time-step, which can be understood as the time interval per frame. Each time step has a corresponding object pose.
[0064] What the "feature extraction" stage outputs is the "static optical flow image" "Dynamic optical flow image" Static feature points and dynamic feature points D 2D,τ , where the subscript 2D represents two-dimensional feature points (i.e., feature points on the image), and the subscript v represents the time step.
[0065] RGB image is mainly used for feature detection and tracking; depth map represents the depth information of each pixel and is used to convert 2D feature points into 3D points; instance segmentation image labels the instance labels of the objects to which each pixel in the image belongs, where background pixels are marked as 0 and dynamic object pixels are marked as 1...n object , where n object is the number of objects in the image; optical flow image or represents the motion between pixels and is used to assist in feature tracking and camera pose estimation.
[0066] "Static feature tracking"
[0067] "Sparse feature point detection": At each time step τ, use a feature detection algorithm to detect sparse static feature points from Among them
[0068] For an indoor environment, use the Shi-Tomasi corner detection algorithm. The Shi-Tomasi corner detection algorithm is a method for identifying corners in an image. Its main purpose is to provide a more stable and reliable corner selection criterion. The core of the Shi-Tomasi algorithm is that it examines the two eigenvalues of the autocorrelation matrix (structure tensor) at each pixel position. For each pixel point, calculate the corresponding autocorrelation matrix (usually obtained by calculating the sum of the squares of the gray-level changes within a window), and then find the two eigenvalues λ1 and λ2 of this matrix. These two eigenvalues represent the intensity of the gray-level changes around this pixel point. The larger eigenvalue corresponds to the most significant change direction, and the smaller eigenvalue corresponds to the second most significant direction. The Shi-Tomasi algorithm only selects the minimum value of the two eigenvalues as the evaluation criterion, that is, R = min(λ1, λ2). Only when R exceeds a certain threshold will the corresponding pixel point be regarded as a corner.
[0069] For the outdoor environment, the ORB (Oriented FAST and Rotated BRIEF) feature detection algorithm is used. ORB uses the FAST (Features from Accelerated Segment Test) algorithm to find feature points. The FAST algorithm determines whether a pixel is a corner point by checking whether a circle of pixels around the pixel forms a local extremum (i.e., the number of consecutive segments that are brighter or darker than the central pixel by a certain threshold). This method has high computational efficiency, but the original FAST corner points lack directional information, which is disadvantageous for achieving rotational invariance in feature matching. To make up for the deficiency of FAST corner points in rotational invariance, ORB calculates the direction of each corner point after detecting the FAST corner points. This is achieved by measuring the direction between the gray centroid in the neighborhood of the corner point and the position of the corner point. The gray centroid refers to the center of gravity based on the pixel intensity distribution in the neighborhood of the corner point. By comparing the position relationship between the centroid and the corner point, a direction vector can be obtained, which gives the ORB feature a certain degree of rotational invariance. After obtaining the key points with directionality, ORB uses a modified version of the BRIEF (Binary Robust Independent Elementary Features) descriptor to describe these key points. The standard BRIEF descriptor generates a binary string as a descriptor by comparing the brightness of a pair of randomly selected pixel points. However, the original BRIEF does not have rotational invariance. Therefore, ORB uses a method called rBRIEF, which takes into account the direction of the key points when constructing the descriptor, thereby enhancing the robustness of the descriptor to rotational changes. Finally, the feature descriptors generated by ORB are used to match the feature points. Since the ORB descriptor is binary, the Hamming distance can be used for efficient matching operations, which is very computationally efficient and suitable for real-time applications.
[0070] "Sparse feature point screening": To improve the robustness of tracking, the Adaptive Non-Maximal Suppression (ANMS) algorithm is used to screen the detected feature points, retain the most informative feature points, and at the same time ensure the uniform distribution of feature points in the image. If the number of tracked feature points is lower than the set "first threshold" η1 (in the experiment, η1 = 1000), the system will re-detect new feature points to supplement.
[0071] "First optical flow tracking algorithm": Through the "first optical flow tracking algorithm", such as the Lucas-Kanade algorithm, the KLT Tracking (Kanade-Lucas-Tomasi Tracking) algorithm, etc., the feature points are tracked between consecutive frames. Perform matching and tracking to generate the correspondence between feature points, and finally generate a "static optical flow image".
[0072] "Inverse projection": Inversely project 2D feature points into 3D feature points.
[0073]
[0074] Among them, represents the 2D pixel measurement coordinates on the image plane, which is a two-dimensional vector containing the row and column coordinates of the pixel. d represents the depth information, that is, the distance from the camera to the observation point. represents the 3D measurement point in the local coordinates, which is a three-dimensional vector. represents the camera pose at time step τ, which is a 4×4 homogeneous transformation matrix. represents the 3D point in the world coordinate system, which is a three-dimensional vector. Π -1 (·) represents the inverse projection function of the camera, that is, converting 2D pixel measurement and depth information into 3D points.
[0075] is to convert to homogeneous coordinate form.
[0076] Obtain 2D image data and depth information from sensors such as RGB-D cameras or stereo cameras, and then through the above formula, restore the position of the 3D point in the local sensor coordinate system from the 2D pixel measurement and depth information.
[0077] Example: Suppose there is an RGB-D camera that captures an image at time step τ, and there is a 2D pixel measurement and the corresponding depth information d = 5 meters, and the camera intrinsic matrix is τ:
[0078] First, use the inverse projection function to convert the 2D pixel measurement and the depth information d into 3D points, that is:
[0079]
[0080] Among them, f is the focal length of the camera.
[0081] Then, convert the 3D point into homogeneous coordinate form:
[0082] Finally, use the camera pose to convert the 3D point into the local sensor coordinates:
[0083]
[0084] where R τ is the rotation matrix, t τ is the translation vector, and are the coordinates of the 3D point in the world coordinate system.
[0085] "Dynamic Feature Tracking"
[0086] "Dense Feature Point Detection": For dynamic objects, the system needs to track the dense feature points D 2t,D on each object. Using the instance segmentation mask image to distinguish dynamic objects from the static background and assign a unique instance label to each dynamic object. If the instance segmentation network does not provide temporally consistent instance labels, the system uses a tracking algorithm (such as ByteTrack) to ensure that the instance labels remain consistent between consecutive frames.
[0087] "Dense Feature Point Screening": Consistent with "Sparse Feature Point Screening", to ensure the robustness of tracking, the ANMS algorithm is used to screen the detected feature points to retain the most informative feature points and evenly distribute them on each dynamic object. If the number of feature points of a certain dynamic object is less than the set "second threshold" η2 (in the experiment, η2 = 800), the system will re-detect new feature points for supplementation.
[0088] "Second Optical Flow Tracking Algorithm": For each dynamic object, the "Second Optical Flow Tracking Algorithm" such as the Farneback algorithm, RAFT, PWC-Net, etc. is used to find the correspondence of feature points between consecutive frames, and finally generate the "dynamic optical flow image".
[0089] "Inverse Projection of Dynamic Feature Points" has the same principle as the "Inverse Projection of Feature Points" described above.
[0090] Obtain the initial camera pose based on the static optical flow image and the static 3D feature points; calculate the motion residuals between consecutive frames on each dynamic object based on the dynamic 3D feature points; obtain the initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals.
[0091] Furthermore, the process of obtaining the initial camera pose based on the static optical flow image and the static 3D feature points includes:
[0092] Construct the correspondence between the 2D static feature points at the current time step and the static 3D feature points at the previous time step, perform a preliminary estimation of the camera pose by minimizing the reprojection error through the PnP algorithm, and verify the estimation result; combine the static optical flow information with the verified estimation result and obtain the initial camera pose through non-linear least squares optimization.
[0093] Furthermore, the motion residuals between consecutive frames on each dynamic object are calculated as follows:
[0094]
[0095] wherein represents the position coordinates of the 3D feature points in the world coordinate system at time step τ, represents the object motion from time step τ - 1 to τ.
[0096] Furthermore, the process of obtaining the initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals includes:
[0097] Construct the correspondence between the 2D dynamic feature points at the current time step and the dynamic 3D feature points at the previous time step, perform a preliminary estimation of the dynamic object motion by minimizing the reprojection error through the PnP algorithm, and verify the estimation result; combine the dynamic optical flow information with the verified estimation result, optimize the preliminarily estimated dynamic object motion through non - linear least squares, and perform a secondary optimization of the optimized dynamic object motion by minimizing the weighted sum of squares of the motion residuals between consecutive frames on each dynamic object to obtain the initial object motion.
[0098] Input and output: The output parameters of the "feature extraction" stage are used as the input parameters of this stage, including the static optical flow image dynamic optical flow image static 3D feature points dynamic 3D feature points D 3D,τ ; the outputs of this stage are the camera trajectory object pose and motion 3D point positions i.e.,
[0099]
[0100] where D represents the time step, represents the set containing all time steps, and the superscript W in the upper left represents the world coordinate system. represents the camera pose at time step τ in the world coordinate system. represents the set of motions of all dynamic objects from time step τ - 1 to τ in the world coordinate system. represents the set of poses of all objects at all time steps at time step τ in the world coordinate system. represents the set of all points at time step τ in the world coordinate system, including dynamic points and static points.
[0101] "Camera pose estimation" + "Optical flow joint optimization": Responsible for estimating the camera pose at each time step τ
[0102] Establish the correspondence between 2D points and 3D points: Use the 2D feature points in the current time step τ and the 3D feature points in the previous time step τ - 1 to establish the correspondence between the two.
[0103] PnP algorithm: Use the Perspective-n-Point (PnP) algorithm to estimate the camera pose The PnP algorithm solves for the camera pose by minimizing the reprojection error:
[0104]
[0105] where, represents the 2D pixel measurement at time step τ (i.e., the pixel coordinates of the feature point), which is a two-dimensional vector, and the superscript i in the upper right indicates the i-th feature point. Π(·) represents the camera projection function that projects a 3D point onto the 2D image plane. represents the camera pose at time step τ. represents the 3D point in the world coordinate system. The residual represents the 2D pixel measurement at time step τ and the difference between the position obtained by projecting the 3D point onto the 2D image plane through the camera pose.
[0106] RANSAC verification: Use the RANSAC (Random Sample Consensus) algorithm to verify the estimation result of the PnP algorithm. RANSAC estimates by randomly selecting subsets and verifies whether these estimates conform to most of the data points. This step can effectively eliminate outliers and improve the robustness of the estimation.
[0107] Optical flow joint optimization: Combine the optical flow information with the estimation result of the PnP algorithm and further improve the camera pose and dynamic object motion estimation through non-linear least squares optimization. Specifically, optimize the following error:
[0108]
[0109] where, f τ-1,τ represents the optical flow vector of the feature point between two frames, represents the set of pixel coordinates of the 2D feature points at time step τ. Π(·) represents the camera projection function. represents the camera pose Inverse matrix of represents the position coordinates of the i-th 3D point in the world coordinate system. After the joint optimization of the PnP algorithm, RANSAC verification, and optical flow, it is also necessary to solve the optical flow optimization problem {θ, f τ-1,τ} to obtain the initial camera pose at the current time step τ
[0110] "Object motion estimation" + "Optical flow optimization": responsible for estimating the motion of dynamic objects at each time step τ
[0111] Establish the correspondence between 2D points and 3D points: Use the 2D feature points in the current time step τ and the 3D feature points in the previous time step τ - 1 to establish the correspondence between the two.
[0112] PnP algorithm: Use the Perspective-n-Point (PnP) algorithm to estimate the motion of dynamic objects The PnP algorithm solves for the object's motion by minimizing the reprojection error:
[0113]
[0114] where represents the 2D pixel measurement at time step τ (i.e., the pixel coordinates of the 2D feature point), which is a two-dimensional vector, and the superscript i in the upper right indicates the i-th feature point. represents the object motion from time step τ - 1 to τ in the world coordinate system. Π(·) represents the projection function of the camera, which projects 3D points onto the 2D image plane. represents the camera pose at time step τ. represents the inverse matrix of represents the 3D point in the world coordinate system at time step τ - 1 The residual represents the difference between the 2D pixel measurement at time step τ
[0115] RANSAC verification: Similarly, use the RANSAC (Random Sample Consensus) algorithm to verify the estimation result of the PnP algorithm. RANSAC estimates by randomly selecting subsets and verifies whether these estimates conform to most data points. This step can effectively eliminate outliers and improve the robustness of the estimation.
[0116] Optical Flow Optimization: Optimize the camera pose by combining optical flow with the motion of dynamic objects and the motion of dynamic objects
[0117] Residual Calculation: Combine the optical flow information with the estimation result of the PnP algorithm
[0118]
[0119] where f τ-1,τ represents the optical flow vector of feature points between two frames, represents the set of pixel coordinates of 2D feature points at time step τ. Π(·) represents the projection function of the camera. represents the camera pose of the inverse matrix. represents the 3D point in the world coordinate system at time step τ - 1 τ-1,τ} to obtain the initial object motion at the current time step τ
[0120] Optimization Problem: Further improve the camera pose and dynamic object motion estimation through non - linear least - squares optimization. That is, optimize τ-1,τ} formula to optimize and to obtain and
[0121]
[0122] where θ represents the system state, including the camera pose and object motion. f τ-1,τ represents the optical flow from time step τ - 1 to τ. ρ h represents the Huber function for robust estimation. represents the optical flow residual, that is, the difference between the optical flow measurement and the optical flow calculated through the system state θ. represents the prior optical flow residual, that is, the difference between the optical flow measurement and the prior optical flow. ∑ f represents the covariance matrix of the optical flow residual, and ∑0 represents the covariance matrix of the prior optical flow residual. N represents the number of optical flow measurements.
[0123] {θ, f τ-1,τ} defines an optimization problem, the goal of which is to find the system state θ and the optical flow f τ-1,τ such that the weighted sum of the squares of all optical flow residuals is minimized. The optical flow residual Robust estimation is performed using the Huber function to reduce the influence of outliers. The optical flow prior residual is used to compare the optical flow measurement with the prior optical flow to improve the accuracy of the estimation.
[0124] "Motion optimization": Responsible for further optimizing the motion estimation of dynamic objects
[0125] Calculate the 3D motion residual: Calculate the motion residual of the 3D points on the dynamic object between consecutive frames, that is, for each 3D point on the dynamic object Calculate its motion residual between time steps τ - 1 and τ:
[0126]
[0127] where represents the position coordinates of the 3D point at time step τ in the world coordinate system. represents the object motion from time step τ - 1 to τ, which is a 4×4 homogeneous transformation matrix. The residual is used to construct the optimization problem, and the position of the 3D point is estimated by minimizing the sum of the squares of all residuals, which is used to measure the position consistency of the 3D point under the object motion.
[0128] Optimization problem: Define an optimization problem, the goal is to minimize the weighted sum of the squares of all 3D motion residuals, and estimate the object motion by achieving this goal The objective function is:
[0129]
[0130] where and represent the sets of dynamic 3D points on the j-th object at time steps τ - 1 and τ respectively. represents the residual of the 3D point position under the object motion. represents the covariance matrix of the 3D point motion residual.
[0131] Use a nonlinear least squares optimization algorithm (such as the Levenberg - Marquardt algorithm) to solve the above optimization problem The Levenberg - Marquardt algorithm is an efficient nonlinear least squares optimization algorithm that can avoid getting stuck in local minima while ensuring the convergence speed.
[0132] Construct a factor graph based on the initial camera pose, the initial object motion, static 3D feature points, and dynamic 3D feature points; perform residual operations based on the factor graph, and obtain the poses and motion trajectories of each dynamic object based on the results of the residual operations, the initial camera pose, and the initial object motion.
[0133] Further, the process of constructing the factor graph includes:
[0134] Use the camera pose, object motion, static 3D feature points, and dynamic 3D feature points as nodes, and use the optical flow information and reprojection error as the constraints of the edges to construct a factor graph.
[0135] Further, the residual operations include 3D point measurement residuals, camera motion residuals, 3D point position residuals, and object motion smoothness residuals.
[0136] Further, construct an objective function based on the results of the residual operations, and obtain the poses and motion trajectories of each dynamic object based on the objective function; the objective function is as follows:
[0137]
[0138] Wherein, represents the residual of the initial camera pose, represents the camera motion residual, represents the 3D point measurement residual, represents the 3D point position residual under object motion, represents the object motion smoothness residual; ∑0, ∑ 3D,R , respectively represent the covariance matrices of the corresponding residuals; ρ h represents the Huber function, represents the set of static 3D points at time step τ, represents the set of dynamic 3D points at time step τ, represents the set of dynamic 3D points on the j-th object at time step τ, represents the set of objects observed at time step τ.
[0139] Further, it is characterized in that it further includes: the current VR device performs contour rendering according to the poses and motion trajectories of each dynamic object, and after cross-verifying the data transmitted by multiple VR devices, transmits the accurate poses and motion trajectories of all moving objects in the scene back to the current VR device to update the contours of the participants.
[0140] Perform multi-person collision detection based on the poses and motion trajectories of each dynamic object.
[0141] "Factor Graph" is a graphical representation method used to describe the relationship between system state variables and measurement data.
[0142] The factor graph can provide a structured and efficient way to represent and solve complex optimization problems, achieving structured representation, efficient solution, modular design, robustness, flexibility, and real-time performance.
[0143] In this system, the system needs to simultaneously estimate the camera pose, object motion, and environmental map, and the relationships between these variables are complex and interdependent. The factor graph makes the optimization problem clearer and easier to solve by representing the relationships between these variables and measurement data as nodes and edges of a graph.
[0144] In this system, the factor graph consists of nodes and edges: the nodes represent system state variables such as the camera pose object motion, and the positions of 3D points
[0145] Residual calculation and joint optimization algorithm: The main purpose is to globally optimize the initial estimate and measurement data to generate a consistent camera trajectory, object motion, and environmental map, and to optimize the system state by minimizing the weighted sum of squares of all residuals to ensure the accuracy and robustness of the estimation results.
[0146] "3D point measurement residual" is used to measure the difference between the 3D measurement point and the 3D point position calculated from the camera pose.
[0147]
[0148] where, represents the 3D measurement point at time step τ and is a three-dimensional vector. represents the inverse matrix of the camera pose at time step τ and is a 4×4 homogeneous transformation inverse matrix. represents the 3D point at time step τ and is a three-dimensional vector. For a static point, when it is first observed, is initialized to For a dynamic point, the dynamic point at each frame at time step τ is also initialized in the same way.
[0149] "Camera motion residual" is used to measure the difference between the camera motion from time step τ - 1 to τ and the camera motion calculated from the camera pose.
[0150]
[0151] where, Denotes the camera motion from time step τ - 1 to τ, which is a 4×4 homogeneous transformation matrix and can be estimated by the inertial measurement unit (IMU). is the inverse matrix of. Denotes the camera pose at time step τ, which is a 4×4 homogeneous transformation matrix, is the inverse matrix of. log(·) represents the matrix logarithm, which is used to convert the transformation matrix into the Lie algebra form, [·] V represents converting a vector in Lie algebra form to the ordinary vector form, log([·] V ) is used to convert the SE(·) in matrix form to the vector form of. The residual is used to construct the optimization problem, and the camera pose and camera motion are estimated by minimizing the sum of the squares of all residuals.
[0152] "3D point position residual" is used to measure the position of the 3D point at time step τ and the difference between the position of the 3D point calculated through object motion at that time.
[0153]
[0154] Among them, represents the homogeneous space coordinates of the 3D point in the world coordinate system at time step τ. Denotes the object motion from time step τ - 1 to τ, which is a 4×4 homogeneous transformation matrix. The residual is used to construct the optimization problem, and the position of the 3D point is estimated by minimizing the sum of the squares of all residuals, which is used to measure the position consistency of the 3D point under object motion.
[0155] "Object motion smoothness residual" is used to measure the difference between the object motion from time step τ - 2 to τ - 1 and the object motion from time step τ - 1 to τ.
[0156]
[0157] Among them, represents the object motion from time step τ - 2 to τ - 1, which is a 4×4 homogeneous transformation matrix. Denotes the object motion from time step τ - 1 to τ, which is also a 4×4 homogeneous transformation matrix. log(·) represents the matrix logarithm, which is used to convert the transformation matrix into the Lie algebra form, [·] VDenotes the conversion of a Lie algebraic form vector to an ordinary vector form, log([·] V ) is used to convert the SE(·) in matrix form to vector form. Is used to construct an optimization problem to estimate object motion by minimizing the sum of the squares of all residuals. The main role of is to enforce constant motion in the world coordinate system and prevent sudden, drastic, and unrealistic changes in object motion between consecutive frames.
[0158] "Joint Optimization Algorithm": Construct the objective function θ MAP , and obtain a globally consistent camera trajectory by minimizing the weighted sum of the squares of all residuals Optimized object pose Optimized object motion Optimized positions of static and dynamic 3D points (i.e., the environmental map).
[0159]
[0160] Among them, Denotes the residual of the initial camera pose, Denotes the camera motion residual, Denotes the residual of 3D point measurement, Denotes the residual of the 3D point position under object motion, Denotes the smoothness residual of object motion. ∑0, ∑ 3D,R , respectively denote the covariance matrices of the corresponding residuals. ρ h Denotes the Huber function for robust estimation. Denotes the set of static 3D points at time step τ, Denotes the set of dynamic 3D points at time step τ, Denotes the set of dynamic 3D points on the j-th object at time step τ. Denotes the set of objects observed at time step τ.
[0161] Sliding window optimization method: Only optimize within the last few time steps to improve efficiency. The window size and overlap can be adjusted according to specific applications to balance computational efficiency and optimization accuracy.
[0162] The core idea of the sliding window optimization method is to limit the optimization process within a fixed-size time window instead of performing global optimization on the entire data sequence. In this way, the system can significantly reduce the computational amount while maintaining high accuracy, thereby improving real-time performance.
[0163] In sliding window optimization, the time window usually contains the most recent several time steps, such as 10 to 20 frames. The data within the window is used to construct the optimization problem, while the data outside the window is ignored or only used to provide prior information. The key to this method lies in how to select the window size and update strategy. The window size is usually determined according to the requirements of the specific application and the computing resources. A larger window can provide a more global optimization effect, but the computational cost will also increase accordingly; while a smaller window focuses more on real-time performance but may sacrifice some global consistency.
[0164] To further improve efficiency, sliding window optimization can be combined with incremental optimization methods. When new measurement data arrives, the system adds the new data to the current window, removes the earliest data, and then performs incremental optimization on the data within the window. This incremental update method avoids re-optimizing the entire window data, thus saving a large amount of computing time.
[0165] This embodiment also provides a multi-person collision detection system for VR large-space immersive tours, including:
[0166] The "client" consists of two stages, namely the "feature extraction" stage and the "algorithm optimization" stage. The former generates feature points and optical flow images, and the latter uses various optimization algorithms to obtain the final object pose and motion trajectory.
[0167] The "server" is used to store static scene data, and performs real-time updates on dynamic objects based on information such as "camera trajectory", "object pose", and "object motion" transmitted back by each VR device. It then distributes the updated dynamic object poses and motion trajectories to different VR devices. Finally, the VR devices render the outlines of these dynamic objects to remind the current participants to avoid collisions with other participants.
[0168] The "server" is connected to each VR device at high speed through the local area network Wifi.
[0169] The "3D point position" is used to construct the "environmental map", that is, the static scene. It is usually only started when the scene is first constructed. Once the server construction is completed, the "3D point position" information will no longer be transmitted to the server in real time. Only when the static scene changes due to reasons such as redecoration or changing the prop position, the server will reconstruct it again. Otherwise, there is no need to transmit this data.
[0170] Information such as "camera trajectory", "object pose", and "object motion" needs to be transmitted to the "server" in real time for each frame. To reduce latency, the current participant's VR device will render the silhouettes of other participants within the field of view based on information such as the current "camera trajectory", "object pose", "object motion", and "3D point position". Then, after the "server" performs cross-verification of the data transmitted through multiple VR devices, information such as the precise poses and motion trajectories of all moving objects in the scene will be sent back to the VR device of the current participant to update the participant's silhouette.
[0171] The "server" will also send the information of the participants inside and outside the group to each VR device, so that each VR device can identify the participants inside and outside the group, and then use different silhouette line colors for participants in different groups to more effectively remind the situation of the current scene participants.
[0172] A VR device for acquiring an RGB image and a depth image, and obtaining an instance segmentation image based on the RGB image;
[0173] A feature extraction module for performing feature extraction based on the instance segmentation image and the depth image to obtain a static optical flow image, static 3D feature points, a dynamic optical flow image, and dynamic 3D feature points;
[0174] A pose and motion estimation module for obtaining an initial camera pose based on the static optical flow image and the static 3D feature points; calculating the motion residuals between consecutive frames on each dynamic object based on the dynamic 3D feature points; obtaining an initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals;
[0175] A factor graph construction module for constructing a factor graph based on the initial camera pose, the initial object motion, static 3D feature points, and dynamic 3D feature points;
[0176] An optimization module for performing residual operations based on the factor graph, and obtaining the poses and motion trajectories of each dynamic object based on the results of the residual operations, the initial camera pose, and the initial object motion;
[0177] A collision detection module for performing multi-person collision detection based on the poses and motion trajectories of each dynamic object.
[0178] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for multi-person collision detection in VR large-space immersive tours, characterized in that, It includes the following steps: Obtain an RGB image and a depth image through a VR device, and obtain an instance segmentation image based on the RGB image; Extract features based on the instance segmentation image and the depth image to obtain a static optical flow image, static 3D feature points, a dynamic optical flow image, and dynamic 3D feature points; Obtain an initial camera pose based on the static optical flow image and the static 3D feature points; calculate the motion residuals between consecutive frames on each dynamic object based on the dynamic 3D feature points; obtain an initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals; Construct a factor graph based on the initial camera pose, the initial object motion, the static 3D feature points, and the dynamic 3D feature points; perform a residual operation based on the factor graph, and obtain the poses and motion trajectories of each dynamic object based on the result of the residual operation, the initial camera pose, and the initial object motion; Perform multi-person collision detection based on the poses and motion trajectories of each dynamic object.
2. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The process of extracting features based on the instance segmentation image and the depth image includes: Select a feature detection algorithm based on the environmental map, perform sparse feature point detection on the instance segmentation image based on the feature detection algorithm, and screen through an adaptive non-maximum suppression algorithm. Then, use an optical flow tracking algorithm to match and track the screened feature points between consecutive frames to generate a static optical flow image, and back-project the 2D feature points on the static optical flow image into 3D feature points to obtain static 3D feature points; Use the instance segmentation mask image to distinguish dynamic objects and static backgrounds, assign a unique instance label to the dynamic objects, perform dense feature point detection and screening, use the optical flow tracking algorithm to generate a dynamic optical flow image, and perform feature point back-projection to obtain dynamic 3D feature points.
3. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The process of obtaining an initial camera pose based on the static optical flow image and the static 3D feature points includes: Construct the correspondence between the 2D static feature points at the current time step and the static 3D feature points at the previous time step, perform a preliminary estimation of the camera pose by minimizing the reprojection error through the PnP algorithm, and verify the estimation result; combine the static optical flow information with the verified estimation result, and obtain the initial camera pose through non-linear least squares optimization.
4. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The motion residual between consecutive frames on each dynamic object is calculated as follows: wherein, represents the position coordinates of the 3D feature points in the world coordinate system at time step τ, represents the object motion from time step τ-1 to τ.
5. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The process of obtaining an initial object motion based on the dynamic optical flow image, the dynamic 3D feature points, and the motion residuals includes: Build the correspondence between the 2D dynamic feature points at the current time step and the dynamic 3D feature points at the previous time step. Perform a preliminary estimation of the dynamic object motion by minimizing the reprojection error using the PnP algorithm, and verify the estimation result. Combine the dynamic optical flow information with the verified estimation result, and optimize the preliminarily estimated dynamic object motion through non-linear least squares. Quadratically optimize the optimized dynamic object motion by minimizing the weighted sum of the squared motion residuals between consecutive frames on each dynamic object to obtain the initial object motion.
6. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The process of constructing the factor graph includes: Taking the camera pose, object motion, static 3D feature points, and dynamic 3D feature points as nodes, and the optical flow information and reprojection error as edge constraints, construct a factor graph.
7. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein The residual operations include 3D point measurement residuals, camera motion residuals, 3D point position residuals, and object motion smoothness residuals.
8. The multi-person collision detection method for VR large-space immersive tour according to claim 1, wherein Construct an objective function based on the residual operation results, and obtain the poses and motion trajectories of each dynamic object based on the objective function; the objective function is as follows: Among them, represents the residual of the initial camera pose, represents the camera motion residual, represents the 3D point measurement residual, represents the 3D point position residual under object motion, represents the object motion smoothness residual; ∑0, ∑ 3D,R , respectively represent the covariance matrices of the corresponding residuals; ρ h represents the Huber function, represents the set of static 3D points at time step τ, represents the set of dynamic 3D points at time step τ, represents the set of dynamic 3D points on the j-th object at time step τ, represents the set of objects observed at time step τ.
9. The method for multi-person collision detection in VR large-space immersive tour according to claim 1, characterized in that, It further includes: The current VR device performs contour rendering according to the poses and motion trajectories of each dynamic object. After cross-verifying the data transmitted by multiple VR devices, the accurate poses and motion trajectories of all moving objects in the scene are transmitted back to the current VR device to update the contours of the participants.
10. A multi-person collision detection system for VR large-space immersive tours, characterized in that, It includes: VR devices, used to obtain RGB images and depth images, and obtain instance segmentation images based on the RGB images; Feature extraction module, used to perform feature extraction based on the instance segmentation images and the depth images to obtain static optical flow images, static 3D feature points, dynamic optical flow images, and dynamic 3D feature points; Pose and motion estimation module, used to obtain the initial camera pose based on the static optical flow images and the static 3D feature points; calculate the motion residuals between consecutive frames on each dynamic object based on the dynamic 3D feature points; obtain the initial object motion based on the dynamic optical flow images, the dynamic 3D feature points, and the motion residuals; Factor graph construction module, used to construct a factor graph based on the initial camera pose, the initial object motion, static 3D feature points, and dynamic 3D feature points; Optimization module, used to perform residual operations based on the factor graph, and obtain the poses and motion trajectories of each dynamic object based on the residual operation results, the initial camera pose, and the initial object motion; Collision detection module, used to perform multi-person collision detection based on the poses and motion trajectories of each dynamic object.
Citation Information
Patent Citations
Virtual reality multi-user interaction method, device and system
CN109671118A
Visual SLAM method based on semantic segmentation of deep learning
CN112132897A
Multi-person VR experience anti-collision control method and device
CN115373512A
Double-row roller bearings with improved sealing performance
KR1020250050511A
Operating system and method for preventing multi-user collision and departure from the virtual reality platform
KR102250870B1