Monocular pedestrian form reconstruction and dynamic scene analysis method based on high-dimensional geometry

Through a method based on high-dimensional geometry, a monocular camera is used to generate a three-dimensional point cloud and combine Transformer and GCN for motion prediction, the problem of high hardware cost and insufficient robustness of a monocular camera in complex environments is solved, and low-cost and high-precision pedestrian detection and behavior prediction is achieved.

CN120339322APending Publication Date: 2025-07-18NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510498495.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when monocular cameras are used for pedestrian detection and behavior prediction in complex environments, there is a problem of high hardware cost and insufficient robustness, especially in dynamic occlusion or multi-target interference scenarios, the accuracy bottleneck is obvious.

Method used

Using a method based on high-dimensional geometry, a single-frame color image is obtained using a monocular camera, a two-dimensional joint key points are extracted through a convolutional neural network, a three-dimensional point cloud data is generated by combining the Rodrigues rotation formula and a skinned multi-body linear model, and a graph convolutional neural network is used to optimize the joint topological relationship, and a Transformer and GCN are fused for motion prediction.

Benefits of technology

Realizing high-precision pedestrian form reconstruction and dynamic scene analysis under low-cost conditions, significantly improving the robustness and accuracy in complex scenarios, and is suitable for intelligent traffic and autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339322A_ABST
    Figure CN120339322A_ABST
Patent Text Reader

Abstract

The invention discloses a monocular pedestrian form reconstruction and dynamic scene analysis method based on high-dimensional geometry. High-precision human body three-dimensional point cloud generation and pedestrian behavior prediction are realized by using a low-cost monocular camera. According to the method, a single-frame color image is used as input, two-dimensional joint point detection, high-dimensional geometric reasoning and deep learning technologies are integrated, and a systematic framework from morphological reconstruction to behavior prediction is constructed. The method comprises joint detection based on OpenPose, three-dimensional point cloud generation of a skin multiple human body linear model (SMPL) and a graph convolutional neural network (GCN), and a motion prediction module fusing Transform and the GCN. Experiments prove that the performance of the method on a public data set is superior to that of the prior art, and particularly the performance is excellent in a dynamic shielding scene. The method has the characteristics of low cost, high precision and strong robustness, and is suitable for the field of intelligent traffic and automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision and intelligent transportation, and particularly relates to a monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry. Background Art

[0002] With the rapid development of autonomous driving technology, pedestrian detection and behavior prediction have become key technologies for improving vehicle safety and system intelligence. As a key link in ensuring vehicle safety and enhancing system intelligence, pedestrian detection and behavior prediction are becoming increasingly important. Traditional methods mostly rely on depth cameras or lidar to obtain high-precision three-dimensional data. Although they have high precision under ideal conditions, their hardware costs are high, and there are performance bottlenecks in complex environments such as dynamic occlusion or multi-target interference. Existing technologies mostly rely on depth cameras or lidar to obtain high-precision point cloud data. Although the effect is good, the hardware cost is high, and the robustness is insufficient in complex environments (such as dynamic occlusion or multi-target interference). Monocular cameras have attracted attention due to their low cost and flexibility. However, due to the lack of depth information, there are accuracy bottlenecks in three-dimensional shape reconstruction and dynamic prediction. The real-time performance and generalization ability of existing technologies in complex scenarios still need to be improved. Summary of the Invention

[0003] Object of the Invention: The object of the present invention is to provide a monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry. It aims to realize the full process from a single-frame image to three-dimensional point cloud generation and high-precision behavior prediction through a low-cost monocular camera, provide new ideas for safety guarantee and interactive perception in intelligent transportation scenarios, significantly reduce hardware dependence, and improve the robustness in complex scenarios. An efficient algorithm framework from human joint segmentation to motion prediction is constructed using a single-frame color image captured by a monocular camera.

[0004] Technical Solution: A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry of the present invention includes the following steps:

[0005] S1. Obtain a single-frame color image through a monocular camera, that is, a single RGB image of a human object in the target scene, and extract two-dimensional joint key points and their associated fields through a convolutional neural network (CNN);

[0006] S2. Based on the Rodrigues rotation formula and the Skinned Multi-Person Linear Model (SMPL), regress the three-dimensional joint positions from the two-dimensional joint key points to generate human point cloud data;

[0007] S3. Based on a graph convolutional neural network (GCN), perform joint segmentation and skeleton topology reconstruction on the human point cloud data to optimize the topological relationship of the joint time series;

[0008] S4. Capture the long-term dependencies and spatial topological relationships in the joint time series through a deep network architecture that combines the ideas of Transformer and GCN, and predict the pedestrian motion trajectory;

[0009] S5. Integrate the human point cloud data in step S2 with the pedestrian motion trajectory in step S4, and output structured data including three-dimensional point cloud coordinates and motion trajectory prediction; apply the output results to intelligent transportation scenario analysis.

[0010] Further, step S1 is specifically: Use a monocular camera to capture a single-frame color image with a resolution of w×h; Generate a two-dimensional confidence map set S and a two-dimensional vector field set L of body part affinity fields (PAFs) through a convolutional neural network (CNN), where the two-dimensional confidence map S includes J confidence maps S = {S1, S2,..., S J}, each part corresponds to a confidence map, where PAFs L includes C two-dimensional vector fields L = {L1, L2,..., L C}, each limb corresponds to a vector field, where PAFs encode the degree of association between joints, and key point detection uses stage-by-stage convolutional inference update, and its calculation process is:

[0011]

[0012] Extract two-dimensional joint key point candidates from the two-dimensional confidence map through non-maximum suppression (NMS) and generate a two-dimensional skeleton structure.

[0013] Further, step S2 is specifically: From the two-dimensional joint key point set S generated in step S1, regress the three-dimensional joint positions J through the Rodrigues rotation formula, where the rotation matrix of each joint is calculated by the following formula:

[0014]

[0015] where is the rotation axis angle of joint j, is the identity matrix; Combine the joint positions J, pose parameters θ, and shape parameters β, and use the SMPL model to generate human point cloud data through the linear blend skinning function W, and the calculation formula is:

[0016]

[0017] where G k ′(θ, J) is the world transformation matrix of joint k; The final position of the point cloud vertex is generated by the following formula:

[0018]

[0019] Wherein:

[0020]

[0021]

[0022] The position of the point cloud vertex is adjusted according to the shape parameter β;

[0023] The joint position is adjusted, and the joint position is adjusted according to the change of the body shape:

[0024]

[0025] Wherein is the joint regression matrix in the training grid.

[0026] Furthermore, step S3 is specifically as follows: input the human body point cloud data generated in step S2 into a graph convolutional neural network (GCN); through the non-Euclidean space modeling ability of GCN, segment the joint points in the point cloud data to generate a joint segmentation result; use GCN to combine the biological prior of the SMPL model to optimize the topological relationship between joint points, reconstruct the human body skeleton structure, and improve the accuracy of three-dimensional shape representation.

[0027] Furthermore, step S4 is specifically as follows: input the three-dimensional joint position sequence generated in step S3 into a hybrid network that fuses Transformer and GCN; through the global attention mechanism of the Transformer model, capture the long-term dependence relationship in the joint time series to generate a time feature representation; through the GCN model, model the spatial topological relationship between joint points in the three-dimensional point cloud data to generate a spatial feature representation; integrate the time feature and the spatial feature to predict the motion trajectory T of the pedestrian pred , and the calculation formula is:

[0028] T pred = Ψ(J, θ, β; Φ)

[0029] Where Ψ is the behavior prediction network.

[0030] Furthermore, step S5 is specifically as follows: integrate the three-dimensional human body shape point cloud data generated in step S2 with the motion trajectory T predicted in step S4 pred ; output structured data including three-dimensional point cloud coordinates and motion trajectory prediction; apply the output result to intelligent transportation scenario analysis, including pedestrian detection, behavior prediction, and autonomous driving obstacle avoidance tasks.

[0031] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored on the memory, and the processor executes the computer program to implement the steps of the method of the present invention.

[0032] The present invention also discloses a computer-readable storage medium, on which computer programs / instructions are stored. When the computer programs / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0033] The present invention also discloses a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the method of the present invention are implemented.

[0034] Advantages: Compared with the prior art, the present invention has the following remarkable advantages:

[0035] 1. Point cloud generation method based on kinematic derivation: Traditional monocular 3D reconstruction mostly relies on end-to-end deep learning and is difficult to fully utilize geometric priors. The present invention combines the Rodrigues matrix transformation formula with kinematic derivation to generate high-precision point clouds through geometric constraints, and uses the SMPL model to achieve accurate joint positions and point cloud generation under low-cost conditions.

[0036] 2. Method of combining GCN and SMPL for applications: By using GCN to model the joint topological relationship of point cloud data and combining the kinematic priors of the SMPL model, the robustness of joint segmentation and skeleton reconstruction in complex dynamic scenarios is optimized, providing reliable input for behavior prediction.

[0037] 3. Dynamic behavior prediction combining the Rodrigues formula: Based on the Rodrigues formula and geometric derivation, a hybrid network of Transformer and GCN is constructed to improve the joint time and space modeling ability and enhance the adaptability of complex motion trajectory prediction.

[0038] 4. Strongly robust dynamic scene perception ability: Experiments on the Human3.6M and CMU Motion Capture datasets show that the present method maintains high precision and robustness in dynamic occlusion and multi-object interference scenarios.

[0039] 5. Practicality and broad application prospects: The present invention realizes the full-process design under low-cost hardware conditions, reduces the dependence on high-cost equipment, and provides an economical and efficient solution for intelligent transportation and autonomous driving.

[0040] 6. The present invention only relies on a monocular camera, significantly reducing the hardware cost, being applicable to mid- and low-end systems. By combining geometric derivation and deep learning, the accuracy of point cloud generation and behavior prediction is better than that of existing methods, performing excellently in complex scenarios and alleviating the performance bottleneck of the lack of depth information in monocular cameras.

[0041] 7. The present invention is applicable to fields such as pedestrian detection in autonomous driving and behavior analysis in intelligent monitoring. By achieving high-precision human motion analysis at low cost, it provides reliable sensing technology support for intelligent transportation systems. Description of the Drawings

[0042] Figure 1 A flowchart of high-dimensional geometry-driven low-cost monocular human body shape reconstruction and dynamic scene analysis;

[0043] Figure 2 The 2D joint points generate a 3D mesh map through the GCN model;

[0044] Figure 3 A spatio-temporal network diagram of human motion prediction based on GCN-Transformer. Detailed Implementation Modes

[0045] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0046] The present invention provides a high-dimensional geometry-driven monocular human body shape reconstruction and dynamic scene analysis method. The implementation process of the present invention will be described in detail below through specific implementation modes. This implementation mode takes pedestrian detection and behavior prediction in an intelligent transportation scene as the background, uses a monocular camera to collect data, and combines a two-dimensional key point detection, three-dimensional shape reconstruction, and dynamic behavior prediction module to verify the effectiveness of the method. All technical details are based on the system framework and experimental verification proposed in this article to ensure the operability and repeatability of the implementation process.

[0047] Example 1: System Framework and Operation Steps

[0048] This example elaborates on the overall framework and operation steps of the system of the present invention. The system takes a single-frame color image captured by a monocular camera as input and generates human body two-dimensional key points, three-dimensional point cloud data, and behavior prediction results. The specific steps are as follows:

[0049] 1. Data Input

[0050] Use a monocular camera to capture a single-frame color image with a resolution of w×h. In specific implementation, the image size is adjusted to 256×256 to adapt to subsequent processing requirements. The input image is in RGB format to ensure that it contains complete information of the human body target and provides basic data for two-dimensional key point detection.

[0051] 2. Two-Dimensional Key Point Detection

[0052] The system performs a feed-forward process on the input image through a convolutional neural network (CNN) to generate two sets of feature representations: a set of two-dimensional confidence maps S and a set of two-dimensional vector fields L of the body part affinity fields (PAFs).

[0053] Two-Dimensional Confidence Map S: It contains J confidence maps S = {S1, S2,..., S J}, and each part corresponds to a confidence map, where The PAFsL consists of C two-dimensional vector fields L = {L1, L2,..., L C}, with each limb corresponding to a vector field, where the PAFs encode the degree of association between joints. The key point detection uses a stage-by-stage convolutional inference update, and its calculation process is as follows:

[0054]

[0055] The key point detection uses a stage-by-stage convolutional inference update. Key point candidates are extracted from the two-dimensional confidence map through non-maximum suppression (NMS), and a two-dimensional skeleton structure is generated. This step is implemented based on the OpenPose algorithm to ensure the accuracy of the detection results and provide reliable input for subsequent three-dimensional inference.

[0056] 3. Three-dimensional joint regression and point cloud generation

[0057] In this step, kinematic derivation and the Rodrigues formula are combined to map the two-dimensional key points to three-dimensional joint positions and generate human point cloud data. The specific process is as follows:

[0058] Joint position regression: The three-dimensional joint positions J are regressed from the set S of two-dimensional key points. The rotation matrix for each joint position is calculated through the Rodrigues formula: where is the axis-angle of rotation for joint j, is the identity matrix; combined with the joint positions J, the pose parameters θ, and the shape parameters β.

[0059] Linear blend skinning: Combining the joint positions J, the pose parameters θ, and the shape parameters β, the human point cloud data is generated through the linear blend skinning function W, and the calculation formula is: where G k ′(θ, J) is the world transformation matrix of joint k.

[0060] Shape and pose adjustment: The final position of the point cloud vertices is generated through the following formula:

[0061]

[0062] where:

[0063]

[0064] The point cloud vertex positions are adjusted according to the shape parameters β.

[0065] Joint position adjustment. The joint positions are adjusted according to the changes in the body shape::

[0066]

[0067] Among them is the joint regression matrix in the training grid.

[0068] Point cloud optimization: Use the SMPL model to generate accurate 3D point cloud joint positions, and perform joint segmentation and skeleton topology reconstruction on the generated point cloud data based on the graph convolutional neural network (GCN). GCN utilizes the non-Euclidean space modeling ability and combines the biological priors of the SMPL model to improve the joint segmentation accuracy and optimize the three-dimensional shape representation.

[0069] 4. Dynamic behavior prediction

[0070] In this step, a spatio-temporal network based on GCN-Transformer is designed to jointly model temporal and spatial features and predict the human motion trajectory. The specific implementation is as follows:

[0071] Time series modeling: Input the three-dimensional joint position sequence into the Transformer model, and capture the long-term dependencies in the time dimension through the global attention mechanism to generate time feature representations.

[0072] Spatial topology modeling: Model the topological relationships between joint points in the point cloud data through GCN to extract local spatial interaction features.

[0073] Trajectory prediction: Integrate temporal and spatial features to predict the pedestrian motion trajectory T pred , and the calculation formula is:

[0074] T pred = Ψ(J, θ, β; Φ).

[0075] Among them, Ψ is the behavior prediction network, Φ is the network parameter, and J, θ, and β are the joint position, pose parameter, and shape parameter respectively. This module shows excellent performance in dynamic scenarios, especially suitable for complex pose changes and partial occlusion situations, significantly improving the prediction accuracy and robustness.

[0076] 5. Result output

[0077] The system integrates the three-dimensional point cloud data and the predicted motion trajectory, and outputs structured results, including three-dimensional joint coordinates and trajectory prediction sequences, for intelligent transportation scenario analysis, such as pedestrian detection and autonomous driving obstacle avoidance.

[0078] Example 2: Experimental verification

[0079] In this example, the effectiveness and robustness of this method are verified through experiments. The experiments are based on multiple public datasets and high-performance computing environments. The specific implementation details are as follows:

[0080] 1. Experimental datasets

[0081] To comprehensively evaluate the method performance, the following three public datasets are selected:

[0082] Human3.6M: It contains 15 types of daily actions (such as walking, running, sitting, etc.) completed by 11 actors in a controlled indoor environment, provides high-resolution RGB images and 3D joint annotations, the training set has approximately 36,000 frames, the background is single, and the annotation accuracy is high.

[0083] CMU Motion Capture: Collected through optical motion capture technology, covering complex actions such as dancing and climbing, provides high-precision 3D joint sequences, and is suitable for verifying the performance in dynamic scenes.

[0084] MPI-INF-3DHP: It contains indoor and outdoor multi-view 3D pose data, has complex backgrounds and lighting changes, and has complete annotations, which is used to evaluate the generality of the method in diverse environments.

[0085] The datasets are selected based on their coverage of multiple scenarios and conditions to ensure the reliability of the experimental results.

[0086] 2. Experimental Setup

[0087] The experiments are completed on a cloud platform server, and the hardware configuration is as follows:

[0088] GPU: NVIDIA Tesla A100 (40GB HBM2);

[0089] CPU: AMD EPYC 7742 (64 cores, 2.25GHz);

[0090] Memory: 256GB;

[0091] Deep learning framework: PyTorch 1.12, supporting CUDA 11.6.

[0092] Training Preprocessing:

[0093] The image size is adjusted to 256×256, using bilinear interpolation;

[0094] Data augmentation includes random horizontal flipping, cropping, brightness and contrast adjustment to improve the model's adaptability;

[0095] Initialize the SMPL model to the standard human pose to ensure reasonable joint distribution.

[0096] Training Parameters:

[0097] Optimizer: Adam;

[0098] Initial learning rate: 1×10 -4 , decaying dynamically during training;

[0099] Batch size: 64;

[0100] Number of training epochs: 300;

[0101] Loss function: A multi-task objective that combines 3D joint position error and keypoint detection loss.

[0102] Evaluation metrics

[0103] To comprehensively and quantitatively evaluate the performance of the proposed method in 3D human body part segmentation, 3D instance segmentation, and human motion prediction tasks, we adopt the following widely recognized metrics:

[0104] Mean Per Joint Position Error (MPJPE): Calculate the average Euclidean distance between the predicted 3D joint positions and the ground truth joint positions, in millimeters (mm), which reflects the overall accuracy of the model's experimental verification.

[0105] PCK@0.5: Calculate the proportion of correctly predicted joints within a 50% joint position error threshold, which mainly reflects the fault tolerance of the model and its robustness in larger scenarios.

[0106] Angle Error: Measure the average error of joint rotation angles, used to evaluate the joint pose reconstruction performance of the model in motion prediction tasks.

[0107] 3D Part Segmentation Average Precision (AP):

[0108] Measure the detection accuracy of the model in 3D human body part segmentation tasks. Specifically, AP P reflects the average segmentation performance at different IoU thresholds, where and correspond to the precision at IoU of 50% and 25% respectively. This metric is mainly used to evaluate the segmentation ability of the model at the part level and accurately reflects the fine-grained characteristics of part division. The formula is as follows:

[0109]

[0110] where IoU is the intersection over union.

[0111] 3D Instance Segmentation Average Precision (AP):

[0112] Used to quantify the 3D segmentation performance of the model at the instance level. AP H represents the average performance at multiple IoU thresholds, and Performance at IoU thresholds of 50% and 25%. This metric mainly reflects the accuracy of the model in segmenting individual instances in complex scenes and is suitable for instance segmentation scenarios with overlap and occlusion. The formula is similar to the AP for part segmentation but is applied at the instance level.

[0113] Relative Mean Per Joint Position Error (RMPJPE, mm) for human motion prediction: Used to measure the accuracy of the model in predicting human motion trajectories in dynamic scenes. R-MPJPE calculates the Euclidean error of the joint positions relative to the ground truth by removing the global translation effect of the root joint, with the unit being millimeters (mm). This metric can accurately evaluate the model's ability to model human kinematic laws and is an important evaluation metric for dynamic scene prediction tasks. The formula is as follows:

[0114]

[0115] where N is the number of frames, J is the number of joints, and root is the position of the root joint.

[0116] Through the comprehensive evaluation of the above metrics, the performance of the method in this paper in fine-grained part segmentation, instance segmentation, and human motion prediction tasks can be comprehensively analyzed, covering the different requirements of static and dynamic scenes.

[0117] 3. Experimental Results

[0118] Overall Performance Comparison

[0119] The experimental results of the method in this paper on the Human3.6M dataset are shown in Table 1 and are compared with the current mainstream methods.

[0120] Table 1: Performance Analysis of 3D Human Joint Segmentation and Motion Prediction Models

[0121]

[0122] As can be seen from Table 1, the method in this paper is superior to the existing methods in multiple evaluation metrics, especially in key performances such as 3D part segmentation (AP P and AP H ) and human motion prediction (R-MPJPE).

[0123] In the 3D part segmentation task, the method in this paper reaches 37.4 in AP P , and reach 94.7 and 99.3 respectively, significantly outperforming other methods. For example, compared with the KPConv+TF-SA network, AP PIt has increased by 8.7%, showing a significant advantage in local feature extraction.

[0124] In the human motion prediction task, the relative position error R-MPJPE of this method reaches 39.7 mm (joint segmentation) and 83.9 mm (motion prediction), both of which are better than the comparison models. For example, compared with the MotionBERT model, the motion prediction error is reduced by 4.8 mm, indicating that this method has stronger capture and prediction capabilities in dynamic scenarios.

[0125] In complex pose scenarios (such as dynamic occlusion and fast actions), the method in this paper combines geometric derivation and deep learning strategies, and utilizes the spatio-temporal joint modeling ability of GCN-Transformer to effectively improve the capture accuracy of joint topology and action trajectories.

[0126] In summary, the model proposed in this paper demonstrates superior performance in multiple metrics. Especially in challenging task scenarios, it can significantly reduce errors and improve the overall accuracy of key point detection and motion prediction.

[0127] Ablation experiments

[0128] To further verify the effectiveness of each module in this method and its contribution to the overall performance, we designed and carried out a series of ablation experiments. Specifically, we removed the key modules in the model one by one, including the Transformer module, the graph convolutional network (GCN) module, and the Rodrigues derivation part, and recorded the performance changes on different evaluation metrics. Table 2 shows the impact of different configurations on the model performance.

[0129] Table 2: Results of ablation experiments (unit: mm)

[0130]

[0131] Removing the Transformer module: After removing the Transformer module, the errors of the model on MPJPE and R-MPJPE increase by 4.4 mm and 3.6 mm respectively, and PCK@0.5 decreases by 3.6%. This shows that the Transformer plays an important role in modeling the long-term dependencies of time series, and can effectively improve the accuracy of human motion prediction through its global attention mechanism.

[0132] Removing the GCN module: After removing the GCN module, the errors of MPJPE and R-MPJPE increase respectively by

[0133] 5.9 mm and 5.4 mm, and PCK@0.5 decreased by 4.1%. This indicates that GCN is crucial for modeling the spatial topological relationships between joints, especially with significant advantages in capturing local motion characteristics in dynamic scenarios.

[0134] Removing Rodrigues derivation: After removing the Rodrigues derivation, the model performance decreased most significantly, where the errors of MPJPE and R-MPJPE increased by 8.8 mm and 7.9 mm respectively, and PCK@0.5 decreased by 5.8%. This is because the Rodrigues derivation can linearize the 3D rotation problem, thus significantly improving the accuracy of joint motion reconstruction. Its absence leads to a large deviation in the model's performance when predicting joint postures.

[0135] Full model: With all modules retained, the method in this paper achieved the best performance in all indicators. Among them, MPJPE and R-MPJPE were 58.4 mm and 45.6 mm respectively, PCK@0.5 reached 88.7%, and the angular error was only 9.8°. This indicates that the synergy of each module in the model is crucial for performance improvement.

[0136] Summary: The ablation experiment results clearly demonstrate the key roles of the Transformer module, GCN module, and Rodrigues derivation in the system. The Transformer module is good at modeling long-term dependencies in time series, the GCN module effectively captures the spatial topological structure of joints, and the Rodrigues derivation provides important mathematical support in 3D motion modeling. The combination of the three enables the method in this paper to perform excellently in human motion prediction and 3D segmentation tasks.

[0137] Example 3: Application scenarios

[0138] This example demonstrates the application of this method in intelligent transportation:

[0139] 1. Autonomous driving: Deployed in the vehicle system, it can detect the three-dimensional postures of pedestrians in real time and predict their motion trajectories, improving the obstacle avoidance ability.

[0140] 2. Intelligent monitoring: Monitor human behaviors in public places to support the recognition of abnormal behaviors. Experiments show that this method can achieve high-precision analysis under low-cost hardware conditions and has wide practicability.

Claims

1. A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry, characterized in that It includes the following steps: S1. Obtain a single-frame color image through a monocular camera, that is, a single RGB image of a human object in the target scene, and extract two-dimensional joint key points and their associated fields through a convolutional neural network (CNN). S2. Based on the Rodrigues rotation formula and the Skinned Multi-Person Linear Model (SMPL), regress the three-dimensional joint positions from the two-dimensional joint key points to generate human point cloud data. S3. Based on the Graph Convolutional Neural Network (GCN), perform joint segmentation and skeleton topology reconstruction on the human point cloud data to optimize the topological relationship of the joint time series. S4. Through a deep network architecture that combines the ideas of Transformer and GCN, capture the long-term dependencies and spatial topological relationships in the joint time series to predict the pedestrian motion trajectory. S5. Integrate the human point cloud data in step S2 with the pedestrian motion trajectory in step S4, and output structured data including three-dimensional point cloud coordinates and motion trajectory prediction; apply the output result to intelligent transportation scenario analysis.

2. The monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry according to claim 1, characterized in that, Step S1 is specifically as follows: Use a monocular camera to capture a single-frame color image with a resolution of w×h; Generate a two-dimensional confidence map set S and a two-dimensional vector field set L of body part affinity fields (PAFs) through a convolutional neural network (CNN), where the two-dimensional confidence map S includes J confidence maps S = {S1, S2,..., S J}, with each part corresponding to a confidence map, where The PAFs L includes C two-dimensional vector fields L = {L1, L2,..., L C}, with each limb corresponding to a vector field, where The PAFs encode the degree of association between joints, and key point detection uses stage-by-stage convolutional inference for update. Its calculation process is as follows: where T P represents the number of iterations in the confidence map prediction stage, and T C represents the number of iterations in the association prediction stage; Extract two-dimensional joint key point candidates from the two-dimensional confidence map through non-maximum suppression (NMS) and generate a two-dimensional skeleton structure.

3. A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry according to claim 1, characterized in that Step S2 is specifically as follows: From the set S of two-dimensional joint key points generated in step S1, regress the three-dimensional joint positions J through the Rodrigues rotation formula, where the rotation matrix of each joint is calculated by the following formula: where is the rotation axis angle of joint j, is the identity matrix; Combined with the joint positions J, pose parameters θ, and shape parameters β, use the SMPL model to generate human point cloud data through the linear blend skinning function W, and the calculation formula is: where G k ′(θ, J) is the world transformation matrix of joint k; the final position of the point cloud vertex is generated by the following formula: Where: The point cloud vertex positions are adjusted according to the shape parameters β; The joint positions are adjusted, and the joint positions are adjusted according to the changes in the body shape: Among them is the joint regression matrix in the training grid.

4. A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry according to claim 1, characterized in that, Step S3 is specifically as follows: Input the human point cloud data generated in step S2 into the Graph Convolutional Neural Network (GCN); through the non-Euclidean space modeling ability of GCN, segment the joint points in the point cloud data to generate joint segmentation results; use GCN to combine the biological priors of the SMPL model to optimize the topological relationship between joint points and reconstruct the human skeleton structure to improve the accuracy of three-dimensional shape representation.

5. A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry according to claim 1, characterized in that, Step S4 is specifically as follows: Input the three-dimensional joint position sequence generated in step S3 into a hybrid network that combines Transformer and GCN; through the global attention mechanism of the Transformer model, capture the long-term dependencies in the joint time series to generate time feature representations; Model the spatial topological relationship between joint points in the three-dimensional point cloud data through the GCN model to generate spatial feature representations; synthesize Time features and spatial features to predict the movement trajectory T of a pedestrian pred , and the calculation formula is as follows: T pred = Ψ(J, θ, β; Φ) Where Ψ is the behavior prediction network.

6. A monocular pedestrian shape reconstruction and dynamic scene analysis method based on high-dimensional geometry according to claim 1, characterized in that, Step S5 specifically includes: integrating the three-dimensional human body shape point cloud data generated in Step S2 with the motion trajectory T predicted in Step S4; outputting structured data including three-dimensional point cloud coordinates and motion trajectory prediction; and applying the output results to intelligent transportation scenario analysis, including pedestrian detection, behavior prediction, and autonomous driving obstacle avoidance tasks. pred Perform integration; output structured data containing three-dimensional point cloud coordinates and motion trajectory prediction; apply the output results to intelligent transportation scenario analysis, including pedestrian detection, behavior prediction, and autonomous driving obstacle avoidance tasks.

7. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method described in claim 1.

8. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in claim 1 are implemented.

9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the method described in claim 1 are implemented.