Monocular video 4D human body and scene real-time three-dimensional reconstruction method and system

Through visual Transformer and improved optimization algorithm combined with adaptive spatiotemporal consistency correction, the deep information extraction and spatiotemporal continuity problems in monocular video 4D reconstruction are solved, and high-precision and low-cost dynamic human body and scene reconstruction is achieved, suitable for augmented reality, virtual reality, intelligent monitoring and motion analysis.

CN120355844AInactive Publication Date: 2025-07-22MIRROR VISION (ZHEJIANG) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510414391.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing monocular video 4D reconstruction technology has shortcomings in deep information extraction, parameter optimization and spatial and temporal continuity correction, making it difficult to achieve high-precision and real-time dynamic human body and scene reconstruction.

Method used

Visual Transformer is used for depth estimation and space-time graph network for human pose estimation, combined with improved whale optimization algorithm and moth flame optimization algorithm for parameter collaborative optimization, and an adaptive space-time consistency correction mechanism is introduced to achieve global and local precise reconstruction.

Benefits of technology

It significantly improves the accuracy, stability and real-time nature of monocular video 4D reconstruction, reduces equipment cost and deployment complexity, and is suitable for fields such as augmented reality, virtual reality, intelligent monitoring and motion analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355844A_ABST
    Figure CN120355844A_ABST
Patent Text Reader

Abstract

The invention discloses a monocular video 4D human body and scene real-time three-dimensional reconstruction method and a monocular video 4D human body and scene real-time three-dimensional reconstruction system. The method comprises the following steps of S1, acquiring a video stream and performing preprocessing; s2, performing monocular depth estimation by a visual Transform, and performing human body posture estimation by a space-time diagram network; s3, the whale optimization algorithm is improved for global search, and the moth flame optimization algorithm is used for local search; s4, jointly updating parameters of the depth estimation network and the human body posture estimation network; s5, performing data fusion to generate a preliminary 4D human body and scene three-dimensional reconstruction model; s6, time sequence modeling and adaptive space-time consistency correction are carried out; and S7, outputting the optimized 4D human body and scene three-dimensional reconstruction model. According to the method, high-precision real-time three-dimensional reconstruction of the 4D human body and the scene is realized through monocular video data, the reconstruction precision, the space-time continuity and the calculation efficiency are improved, and the method is widely applied to the fields of intelligent monitoring, virtual reality, motion analysis and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a method and system for real-time three-dimensional reconstruction of 4D human body and scene from monocular video. Background Art

[0002] In recent years, with the rapid development of computer vision, artificial intelligence, and graphics processing technology, the three-dimensional reconstruction technology based on monocular video has gradually become a research hotspot. Traditional three-dimensional reconstruction methods mainly rely on multi-cameras or depth sensors to collect scene information, but these methods face problems such as high cost, complex deployment, and data synchronization in practical applications. The monocular video reconstruction method has received extensive attention due to its simple equipment and low cost. However, since the monocular sensor itself lacks depth information, how to accurately recover the three-dimensional structure from two-dimensional images has always been a technical problem. In addition, in the dynamic reconstruction of human body and scene, how to achieve 4D reconstruction, that is, to capture information in both time domain and space domain simultaneously, has become an important research direction.

[0003] The existing technologies mainly adopt two major modules, depth estimation and human pose estimation, in the field of monocular video 4D reconstruction. In the depth estimation part, traditional convolutional neural networks (CNNs) are usually used for monocular depth estimation to generate depth maps corresponding to each frame of image, while in the human pose estimation part, graph convolutional networks (GCNs) or other network architectures are mostly used to extract human key points. Although these methods have achieved the reconstruction of scene and human body structures to a certain extent, there are still many deficiencies in practical applications. First, traditional CNNs have problems in capturing global geometric information insufficiently in dealing with complex scenes, and it is difficult to balance between the overall scene structure and the detailed parts. Second, the existing human pose estimation methods are easily affected by factors such as occlusion and illumination changes when dealing with dynamic scenes, resulting in inaccurate extraction of key points, which in turn affects the overall three-dimensional reconstruction effect.

[0004] In addition, in terms of parameter optimization, although traditional global optimization methods such as the whale optimization algorithm have strong global search capabilities, they lack sufficient detailed tuning means in the local search stage; while some local search-based optimization algorithms cannot effectively avoid local extrema and cannot give full play to the complementary advantages of the two optimization strategies. Therefore, how to organically combine global search and local search to form a hybrid optimization strategy that can not only effectively explore the global optimal solution but also make fine adjustments in the local area has become an urgent problem to be solved. At the same time, the parameters of the optimization algorithms in the existing technologies usually adopt fixed values and lack the ability of adaptive regulation, making it difficult to adapt to the challenges brought by different scenes and data changes.

[0005] Temporal modeling and spatio-temporal consistency correction are also weak links in the existing technologies. Most of the existing methods adopt simple temporal modeling techniques, but in the processing of dynamic changes in consecutive frame data, they often cannot effectively ensure the spatio-temporal continuity of the reconstruction results. Due to the large dynamic differences between frames in a video, how to achieve smooth transitions between frames while maintaining dynamic information has become an important factor affecting the 4D reconstruction effect. Traditional spatio-temporal consistency correction methods lack pertinence and are prone to cause breakage or blurring phenomena in the reconstruction model in motion scenes, unable to meet the requirements of real-time interaction and high-precision reconstruction.

[0006] Therefore, how to provide a method and system for real-time three-dimensional reconstruction of 4D human body and scene from monocular video is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] An object of the present invention is to propose a method and system for real-time three-dimensional reconstruction of 4D human body and scene from monocular video. The present invention makes full use of depth vision algorithms, spatio-temporal graph networks, and intelligent optimization techniques. It realizes global depth estimation through a vision Transformer, combines spatio-temporal graph networks for human pose estimation, and uses an improved whale optimization algorithm and a local search algorithm to collaboratively optimize network parameters. At the same time, an adaptive spatio-temporal consistency correction mechanism is introduced to achieve high-precision dynamic reconstruction of consecutive frame data. The present invention has the advantages of high reconstruction accuracy, good system real-time performance, strong robustness, and low-cost deployment.

[0008] The real-time three-dimensional reconstruction method of 4D human body and scene from monocular video according to an embodiment of the present invention includes the following steps:

[0009] S1. Obtain a video stream from a monocular video sensor and preprocess each frame image of the video stream;

[0010] S2. Perform monocular depth estimation on each frame image of the preprocessed video stream by using a depth estimation network based on a vision Transformer to generate a corresponding depth map, and at the same time use a human pose estimation network based on a spatio-temporal graph network for human pose estimation to extract human key point information;

[0011] S3. Use an improved whale optimization algorithm to globally search and update the key parameters in the depth estimation network, and use a moth-flame optimization algorithm to locally search and optimize the key parameters in the human pose estimation network;

[0012] S4. Collaboratively integrate the global search results and local search results obtained in step S3, and collaboratively optimize the depth estimation network and the human pose estimation network through a joint update mechanism;

[0013] S5. Perform multi-level data fusion on the depth map and human key point information to generate a preliminary 4D human and scene three-dimensional reconstruction model;

[0014] S6. Use temporal modeling technology to perform dynamic reconstruction and spatio-temporal consistency correction on the continuous frame data in the preliminary 4D human and scene three-dimensional reconstruction model;

[0015] S7. Output the three-dimensional reconstruction model of the 4D human and scene after overall parameter optimization and spatio-temporal correction.

[0016] Optionally, the S2 specifically includes:

[0017] S21. Input each frame image I in the preprocessed video stream t into the multi-scale feature extraction module based on Vision Transformer to obtain the multi-scale feature representation F t ;

[0018] S22. Based on the multi-scale feature representation, use the Vision Transformer encoder to perform depth estimation and generate the depth map D t ;

[0019] S23. Input each frame image I in the preprocessed video stream t into the pose estimation module based on the spatio-temporal graph network at the same time, and use spatio-temporal relationship modeling to extract the human local feature L t ;

[0020] S24. Process the human local feature L t through the graph convolution and spatio-temporal dynamic fusion module to generate the human key point information K t ;

[0021] S25. Use the multi-head attention mechanism to perform interactive fusion on the depth map D t and the human key point information K t to generate the joint feature J t .

[0022] Optionally, the S3 specifically includes:

[0023] S31. Set the key parameters of the depth estimation network as W enc and the fitness function F enc (W enc ), and initialize the parameter candidate solution set N represents the size of the candidate solution set;

[0024] S32. Use the improved whale optimization algorithm to perform global search and update on the parameter candidate solution set ;

[0025] S33. Set the key parameters of the human pose estimation network as W fusion and the fitness function F fusion (W fusion ), and initialize the parameter candidate solution set M represents the number of candidate solutions for the human pose estimation network parameters;

[0026] S34. Use the moth - flame optimization algorithm to perform local search optimization on the parameter candidate solution set and update each candidate solution:

[0027]

[0028] where W flame represents the current local optimal parameter, γ and δ are control coefficients, r j is a random number between 0 and 1, is the updated candidate solution, and exp is the exponential function;

[0029] S35. According to the evaluation results of the fitness function F enc and F fusion , respectively select the optimal parameters and from the updated candidate solutions as the final optimized parameters of the depth estimation network and the human pose estimation network.

[0030] Optionally, the S32 specifically includes:

[0031] S321. Set the key parameters of the depth estimation network as W enc , initialize the candidate solution set where N represents the total number of candidate solutions, and the initial velocity v i (0) of each candidate solution, and initialize the global optimal parameter W best ;

[0032] S322. Set the iteration count t = 1 and the maximum iteration count T max , and define the adaptive contraction coefficient where a0 is the preset initial contraction coefficient;

[0033] S323. Calculate the Euclidean distance of each candidate solution best from the global optimal solution W

[0034] S324. Introduce the acceleration update based on the damped resonance model to simulate the accelerated pursuit behavior adopted by whales during predation. When the candidate solution approaches the global optimal solution, it has both spiral randomness and a direct pursuit effect similar to damped acceleration. Calculate the velocity of candidate solution i:

[0035]

[0036] where μ is the damping factor, κ is the acceleration constant, v i (t + 1) represents the velocity update value of candidate solution i at the (t + 1)-th iteration, and v i (t) represents the velocity of candidate solution i at the t-th iteration. Calculate the position of the candidate solution according to the updated velocity:

[0037]

[0038] where represents the new position of candidate solution i after the acceleration update using the damped resonance model;

[0039] S325. Combine the spiral update and the damped resonance update, and use weighted fusion to update each candidate solution as follows:

[0040]

[0041] where ω is the fusion weight, β is the sine update control coefficient, is the updated candidate solution, and sin is the sine function;

[0042] S326. Evaluate the updated candidate solution set according to the fitness function F enc , and select the candidate solution that makes F enc reach the optimal value as the new global optimal parameter W best , update the iteration count t ← t + 1, and terminate the update until t ≥ T max , and output the optimized depth estimation network parameters

[0043] S327. Use the chaotic mapping mechanism to adaptively regulate the fusion weight ω and the contraction coefficient a(t):

[0044] Set the initial chaotic variable C(0), and use the Logistic chaotic mapping:

[0045] C(t + 1) = μ C ·C(t)·(1 - C(t));

[0046] Among them, C(t + 1) represents the chaotic variable at the next iteration t + 1, C(t) represents the chaotic variable at the iteration number t, and μ C represents the parameter that controls the chaotic behavior;

[0047] The current chaotic variable C(t) is used to update the fusion weight and the contraction coefficient:

[0048] ω(t) = ω min +(ω max -ω min )·C(t);

[0049] a(t) = a0·C(t);

[0050] Among them, ω min and ω max are respectively the minimum and maximum values of the fusion weight, a0 is the initial contraction coefficient, and ω(t) represents the fusion weight at the iteration number t;

[0051] In each iteration, the updated fusion weight and contraction coefficient are used to replace the original parameters, and are used to update each candidate solution in S325 for update.

[0052] Optionally, the S4 specifically includes:

[0053] S41. The candidate parameter sets of the depth estimation network and the human pose estimation network after being updated by the S3 step are and where N represents the total number of candidate solutions of the depth estimation network parameters, and M represents the number of candidate solutions of the human pose estimation network parameters;

[0054] S42. Define the joint fitness function F joint , and evaluate the overall performance of the parameter updates of the depth estimation network and the human pose estimation network:

[0055]

[0056] Among them, λ is the trade-off factor, is the fitness function of the depth estimation network parameters :

[0057]

[0058] Among them, N s represents the total number of samples used for evaluation, I j represents the j-th input image, represents the predicted depth generated by the depth estimation network based on the vision Transformer using the parameter for the image I j Dj Denote the image I j The corresponding true depth data;

[0059] Among them, Is the fitness function of the human pose estimation network parameters :

[0060]

[0061] Among them, N k Denote the total number of image samples used for evaluation, I l Denote the l-th input image, Denote the human pose estimation network based on the spatio-temporal graph network using the parameters For the image I l The predicted human key point information generated, K l Denote the image I l The corresponding true human key point information;

[0062] S43. For each pair in the candidate parameter set Calculate the joint fitness value

[0063]

[0064] S44. According to the calculated joint fitness value, select the parameter pair that makes the joint fitness function F joint Reach the optimal value

[0065] S45. Adopt a joint update mechanism to use the selected parameter pair to simultaneously update the depth estimation network based on the Vision Transformer and the human pose estimation network based on the spatio-temporal graph network to complete the collaborative optimization of the two networks.

[0066] Optionally, the S6 specifically includes:

[0067] S61. Set the continuous frame data sequence in the preliminary 4D human and scene three-dimensional reconstruction model as Among them, R t Denote the reconstruction result of the t-th frame, T denotes the total number of frames;

[0068] S62. Use the continuous frame data sequence as the input and perform dynamic reconstruction through the temporal modeling network f TM To generate the temporal dynamic reconstruction result

[0069] S63. Introduce adaptive spatio-temporal consistency correction to the dynamic reconstruction result of each frame and calculate the temporal adaptation coefficient λ t :

[0070]

[0071] Among them, represents the temporal dynamic reconstruction result of the t-th frame, represents the temporal dynamic reconstruction result of the (t - 1)-th frame, and δ is a preset bias parameter. represents the Euclidean distance between consecutive frames, and exp is the exponential function. Adaptive weighting is used to correct the temporal dynamic reconstruction result:

[0072]

[0073] Among them, represents the reconstructed result after correction of the previous frame, represents the reconstructed result after correction of the t-th frame;

[0074] S64. Construct a spatio-temporal consistency metric function E temp represents the difference between consecutive frames:

[0075]

[0076] S65. According to the reconstructed result after correction Output the final 4D human body and scene three-dimensional reconstruction model after temporal modeling and spatio-temporal consistency correction.

[0077] According to the monocular video 4D human body and scene real-time three-dimensional reconstruction system of the embodiment of the present invention, it includes the following modules:

[0078] Video acquisition and preprocessing module, used to obtain a video stream from a monocular video sensor and preprocess the video stream;

[0079] Depth estimation module, used to perform monocular depth estimation and generate a depth map;

[0080] Human body pose estimation module, used to perform human body pose estimation and extract key point information;

[0081] Parameter optimization module, used to globally search and update the depth estimation network parameters, locally search and optimize the human body pose estimation network parameters, and perform collaborative optimization through joint update;

[0082] Data fusion module, used to perform multi-level fusion of the depth map and human body key point information according to preset rules to generate a preliminary 4D human body and scene three-dimensional reconstruction model;

[0083] Temporal modeling and correction module, used to perform dynamic reconstruction and spatio-temporal consistency correction on consecutive frame data by using temporal modeling technology;

[0084] The output display and interaction module is used to output the optimized 4D human body and scene 3D reconstruction model, and provide real-time interaction and display functions.

[0085] The beneficial effects of the present invention are:

[0086] The present invention realizes 4D real-time 3D reconstruction of human bodies and scenes in monocular videos by comprehensively using visual transformers, spatiotemporal graph networks, and improved hybrid optimization strategies, fundamentally overcoming the defects of the prior art in depth information extraction, parameter optimization, and spatiotemporal continuity correction. Using a depth estimation module based on visual transformers, the system can effectively capture global geometric information, while the human posture estimation module based on spatiotemporal graph networks accurately extracts dynamic human key point information. By comprehensively evaluating the two parts of the parameters through a joint fitness function, and using a joint update mechanism to coordinately optimize the parameters of the depth estimation and human posture estimation networks, the present invention realizes dual accurate reconstruction of scene structure and human body dynamics.

[0087] In terms of parameter optimization, the present invention innovatively introduces an improved whale optimization algorithm, which realizes adaptive update of candidate solutions through the organic fusion of global search and local search, combined with damped resonance model and chaotic mapping mechanism, and significantly improves the robustness and accuracy of depth estimation network parameters in complex scenes. At the same time, in the link of time series modeling and time-space consistency correction, the system performs weighted fusion of continuous frame data through an adaptive weight mechanism, so that the dynamic reconstruction results can achieve smooth transition between frames while ensuring the true presentation of motion information, eliminating the reconstruction breakage or blurring caused by excessive differences between frames in traditional methods.

[0088] In general, the present invention not only significantly improves the accuracy, stability and real-time performance of monocular video 4D human body and scene reconstruction, but also has obvious advantages in terms of equipment cost, deployment complexity and data processing efficiency. The application of this system will provide reliable, efficient and low-cost solutions for the fields of augmented reality, virtual reality, intelligent monitoring, motion analysis and digital human modeling, and effectively promote the intelligent and popular development of related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0090] Figure 1 This is a flow chart of the method for real-time 3D reconstruction of 4D human body and scene from monocular video proposed by the present invention;

[0091] Figure 2Schematic diagram of the structure of the monocular video 4D human body and scene real-time three-dimensional reconstruction system proposed by the present invention. Detailed implementation manners

[0092] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0093] Refer to Figure 1 , the monocular video 4D human body and scene real-time three-dimensional reconstruction method, including the following steps:

[0094] S1. Obtain a video stream from a monocular video sensor and preprocess each frame image of the video stream;

[0095] S2. Use a depth estimation network based on Vision Transformer to perform monocular depth estimation on each frame image of the preprocessed video stream, generate a corresponding depth map, and at the same time use a human pose estimation network based on a spatio-temporal graph network to perform human pose estimation and extract human key point information;

[0096] S3. Use an improved whale optimization algorithm to globally search and update the key parameters in the depth estimation network, and use a moth-flame optimization algorithm to locally search and optimize the key parameters in the human pose estimation network;

[0097] S4. Collaboratively integrate the global search results and local search results obtained in step S3, and co-optimize the depth estimation network and the human pose estimation network through a joint update mechanism;

[0098] S5. Perform multi-level data fusion on the depth map and the human key point information to generate a preliminary 4D human body and scene three-dimensional reconstruction model;

[0099] S6. Use temporal modeling technology to perform dynamic reconstruction and spatio-temporal consistency correction on the continuous frame data in the preliminary 4D human body and scene three-dimensional reconstruction model;

[0100] S7. Output the three-dimensional reconstruction model of the 4D human body and scene after overall parameter optimization and spatio-temporal correction.

[0101] In this embodiment, the S2 specifically includes:

[0102] S21. Input each frame image I t in the preprocessed video stream into a multi-scale feature extraction module based on Vision Transformer to obtain a multi-scale feature representation F t ;

[0103] S22. Based on multi-scale feature representation, use a vision Transformer encoder for depth estimation to generate a depth map D t ;

[0104] S23. Input each frame image I in the preprocessed video stream t into the pose estimation module based on the spatio-temporal graph network at the same time, and use spatio-temporal relationship modeling to extract the local human body features L t ;

[0105] S24. Process the local human body features L t through the graph convolution and spatio-temporal dynamic fusion module to generate the human body key point information K t ;

[0106] S25. Use the multi-head attention mechanism to interact and fuse the depth map D t and the human body key point information K t to generate the joint feature J t .

[0107] In this embodiment, the specific content of S3 includes:

[0108] S31. Set the key parameters of the depth estimation network as W enc and the fitness function F enc (W enc ), and initialize the parameter candidate solution set N represents the size of the candidate solution set;

[0109] S32. Use the improved whale optimization algorithm to globally search and update the parameter candidate solution set ;

[0110] S33. Set the key parameters of the human pose estimation network as W fusion and the fitness function F fusion (W fusion ), and initialize the parameter candidate solution set M represents the number of candidate solutions for the human pose estimation network parameters;

[0111] S34. Use the moth-flame optimization algorithm to locally search and optimize the parameter candidate solution set and update each candidate solution:

[0112]

[0113] where W flame represents the current local optimal parameter, γ and δ are control coefficients, r j is a random number between 0 and 1, is the updated candidate solution, and exp is the exponential function;

[0114] S35. According to the evaluation result of the fitness function F enc and F fusion select the optimal parameters from the updated candidate solutions respectively and as the final optimized parameters of the depth estimation network and the human pose estimation network.

[0115] In this embodiment, the S32 specifically includes:

[0116] S321. Set the key parameters of the depth estimation network as W enc , initialize the candidate solution set where N represents the total number of candidate solutions, and the initial velocity v i (0) corresponding to each candidate solution, and initialize the global optimal parameter W best ;

[0117] S322. Set the iteration count t = 1 and the maximum iteration count T max , and define the adaptive contraction coefficient where a0 is the preset initial contraction coefficient;

[0118] S323. Calculate the Euclidean distance from each candidate solution best to the global optimal solution W

[0119] S324. Introduce the acceleration update based on the damped harmonic model to simulate the accelerated pursuit behavior adopted by whales during predation. When the candidate solution approaches the global optimal solution, it has both spiral randomness and a direct pursuit effect similar to damped acceleration. Calculate the velocity of candidate solution i:

[0120]

[0121] where μ is the damping factor, κ is the acceleration constant, v i (t + 1) represents the velocity update value of candidate solution i at the (t + 1)-th iteration, and v i (t) represents the velocity of candidate solution i at the t-th iteration. Calculate the position of the candidate solution according to the updated velocity:

[0122]

[0123] where represents the new position of candidate solution i after adopting the acceleration update of the damped harmonic model;

[0124] S325. Combine the spiral update and the damped harmonic update, and use weighted fusion for each candidate solution Perform an update:

[0125]

[0126] where ω is the fusion weight and β is the sine update control coefficient. is the updated candidate solution, and sin is the sine function;

[0127] S326. Evaluate the updated candidate solution set according to the fitness function F enc and select the candidate solution that makes F enc reach the optimal value as the new global optimal parameter W best , update the iteration count t←t + 1 until t≥T max and then terminate the update, and output the optimized depth estimation network parameters

[0128] S327. Adopt a chaotic mapping mechanism to adaptively regulate the fusion weight ω and the contraction coefficient a(t):

[0129] Set the initial chaotic variable C(0) and adopt the Logistic chaotic mapping:

[0130] C(t + 1) = μ C · C(t) · (1 - C(t));

[0131] where C(t + 1) represents the chaotic variable at the next iteration t + 1, C(t) represents the chaotic variable at the iteration count t, and μ C represents the parameter controlling the chaotic behavior;

[0132] Use the current chaotic variable C(t) to update the fusion weight and the contraction coefficient:

[0133] ω(t) = ω min + (ω max - ω min ) · C(t);

[0134] a(t) = a0 · C(t);

[0135] where ω min and ω max are the minimum and maximum values of the fusion weight respectively, a0 is the initial contraction coefficient, and ω(t) represents the fusion weight at the iteration count t;

[0136] In each iteration, use the updated fusion weight and contraction coefficient to replace the original parameters and use them to update each candidate solution in S325.

[0137] In this embodiment, the specific content of S4 includes:

[0138] The candidate parameter sets after the depth estimation network and the human pose estimation network are updated respectively in step S3 are and where N represents the total number of candidate solutions for the parameters of the depth estimation network, and M represents the number of candidate solutions for the parameters of the human pose estimation network;

[0139] S42. Define the joint fitness function F joint , and evaluate the overall performance of the parameter updates of the depth estimation network and the human pose estimation network:

[0140]

[0141] where λ is a trade-off factor, is the fitness function for the parameters of the depth estimation network:

[0142]

[0143] where N s represents the total number of samples used for evaluation, I j represents the j-th input image, represents the predicted depth generated by the depth estimation network based on the Vision Transformer using the parameters for the image I j , and D j represents the corresponding true depth data of the image I j ;

[0144] where is the fitness function for the parameters of the human pose estimation network:

[0145]

[0146] where N k represents the total number of image samples used for evaluation, I l represents the l-th input image, represents the predicted human key point information generated by the human pose estimation network based on the spatio-temporal graph network using the parameters for the image I l , and K l represents the corresponding true human key point information of the image I l ;

[0147] S43. For each pair of in the candidate parameter set, calculate the joint fitness value

[0148]

[0149] S44. Select the parameter pair that makes the joint fitness function F reach the optimal value from the candidate parameter set according to the calculated joint fitness value joint parameter pair

[0150] S45. Adopt a joint update mechanism to use the selected parameter pair to simultaneously update the depth estimation network based on the Vision Transformer and the human pose estimation network based on the spatio-temporal graph network, and complete the collaborative optimization of the two networks.

[0151] In this embodiment, the S6 specifically includes:

[0152] S61. Set the continuous frame data sequence in the preliminary 4D human and scene three-dimensional reconstruction model as where R t represents the reconstruction result of the t-th frame, and T represents the total number of frames;

[0153] S62. Use the continuous frame data sequence as the input, and perform dynamic reconstruction through the temporal modeling network f TM to generate the temporal dynamic reconstruction result

[0154] S63. Introduce adaptive spatio-temporal consistency correction for the dynamic reconstruction result of each frame, and calculate the temporal adaptation coefficient λ t :

[0155]

[0156] where, represents the temporal dynamic reconstruction result of the t-th frame, represents the temporal dynamic reconstruction result of the (t - 1)-th frame, δ is a preset bias parameter, represents the Euclidean distance between consecutive frames, exp is the exponential function, and the temporal dynamic reconstruction result is corrected by adaptive weighting:

[0157]

[0158] where, represents the corrected reconstruction result of the previous frame, represents the corrected reconstruction result of the t-th frame;

[0159] S64. Construct a spatio-temporal consistency metric function E temp to represent the difference between consecutive frames:

[0160]

[0161] S65. According to the corrected reconstruction result Output the 4D human body and scene three-dimensional reconstruction model after temporal modeling and spatio-temporal consistency correction.

[0162] Reference Figure 2 , a monocular video 4D human body and scene real-time three-dimensional reconstruction system, including the following modules:

[0163] Video acquisition and preprocessing module, used to obtain the video stream from the monocular video sensor and preprocess the video stream;

[0164] Depth estimation module, used to perform monocular depth estimation and generate a depth map;

[0165] Human body pose estimation module, used to perform human body pose estimation and extract key point information;

[0166] Parameter optimization module, used to globally search and update the depth estimation network parameters, locally search and optimize the human body pose estimation network parameters, and perform collaborative optimization through joint updates;

[0167] Data fusion module, used to perform multi-level fusion of the depth map and human body key point information according to preset rules to generate a preliminary 4D human body and scene three-dimensional reconstruction model;

[0168] Temporal modeling and correction module, used to perform dynamic reconstruction and spatio-temporal consistency correction on continuous frame data using temporal modeling technology;

[0169] Output display and interaction module, used to output the optimized 4D human body and scene three-dimensional reconstruction model and provide real-time interaction and display functions.

[0170] Example 1:

[0171] To verify the feasibility of the present invention in implementation, the present invention is applied to a large commercial complex. Existing technologies often rely on multi-cameras or depth sensors to collect data in indoor complex environments. However, these solutions are difficult to meet the requirements of real-time monitoring and dynamic analysis due to the high cost of equipment, complex deployment, and large data synchronization difficulties. At the same time, traditional methods for monocular depth estimation based on convolutional neural networks (CNNs) have deficiencies in capturing global geometric information, and for pose estimation methods based on graph convolutional networks (GCNs), when encountering interferences such as occlusion and illumination changes, the accuracy of human body key point extraction is severely affected, resulting in unsatisfactory reconstruction effects.

[0172] In this embodiment, the system first uses the video acquisition and preprocessing module to obtain a real-time video stream from a monocular video sensor, and performs denoising, enhancement, and normalization processing on each frame of the image to ensure the quality of the input image. Subsequently, the depth estimation module uses a network based on Vision Transformer to perform monocular depth estimation on the preprocessed image, generating an accurate depth map; while the human pose estimation module adopts a design based on a spatio-temporal graph network to quickly and accurately extract the key point information of the human body in the video frame. The data outputs of these two modules lay a solid foundation for subsequent data fusion and temporal modeling.

[0173] In terms of parameter optimization, the present invention introduces an improved hybrid optimization strategy. The improved whale optimization algorithm is respectively used to globally search and update the parameters of the depth estimation network, and the moth-flame optimization algorithm is used to locally search and optimize the parameters of the human pose estimation network. Then, through the joint fitness function and joint update mechanism, the collaborative optimization of the two sets of parameters is realized, so that the system can not only achieve the optimal exploration of global parameters as a whole, but also perform fine-tuning in the local area. After this optimization process, the system can maintain a high-precision reconstruction effect in different complex scenarios, significantly reducing the reconstruction error caused by improper parameter settings.

[0174] In the temporal modeling and spatio-temporal consistency correction stage, the system uses the temporal modeling network to perform dynamic reconstruction on the continuous frame data in the preliminary 4D reconstruction model, and introduces an adaptive spatio-temporal consistency correction mechanism. This mechanism calculates the Euclidean distance between consecutive frames and performs weighted fusion of the reconstruction results of the current frame and the previous frame according to the adaptive weight, so as to achieve smooth transition between frames while retaining the dynamic change information, greatly improving the continuity and stability of the reconstruction model. The entire system not only demonstrates extremely high robustness in complex dynamic environments, but also realizes real-time data processing, ensuring that the system response delay is extremely low, and is applicable to various scenarios such as security monitoring, motion behavior analysis, and intelligent interaction.

[0175] Table 1 Performance comparison data of the 4D reconstruction system

[0176]

[0177] As can be seen from the data in Table 1, the system of the present invention is significantly superior to traditional monocular reconstruction methods and multi-camera sensor systems in key indicators such as average reconstruction error, frame rate, real-time response delay, spatio-temporal consistency, and user satisfaction. First of all, the average reconstruction error is only 4.8 mm, compared with 8.0 mm of the traditional monocular system and 6.5 mm of the multi-camera system, showing the excellent performance of this system in capturing global geometric information and detail reproduction. This improvement benefits from the adoption of Vision Transformer combined with the improved hybrid optimization strategy, enabling the depth estimation and human pose estimation modules to be collaboratively optimized, effectively reducing the error.

[0178] Meanwhile, the frame rate of this system reaches 32 frames per second, far higher than 25 frames per second of traditional monocular methods and 28 frames per second of multi-camera systems, ensuring the smoothness and real-time nature of the images during monitoring and dynamic analysis. The real-time response latency is controlled within 48 milliseconds, while those of other systems are 70 milliseconds and 65 milliseconds respectively, indicating that the innovation in parameter optimization and timing modeling of this invention effectively improves the response speed.

[0179] In addition, the spatio-temporal consistency index shows that the mean square error of this system is 0.0021, far lower than 0.0040 of traditional monocular methods and 0.0032 of multi-camera systems. This is mainly attributed to the adaptive spatio-temporal consistency correction mechanism, which automatically adjusts the weights according to the dynamic changes between consecutive frames, achieves smooth transitions between frames, and ensures continuity and stability in dynamic scenes. Finally, the results of the user satisfaction survey also show that this system has obtained a satisfaction rate of 93%, significantly higher than 80% - 85% of traditional methods, fully demonstrating the superior experience of this invention in practical applications.

[0180] Generally speaking, the system of this invention realizes the 4D real-time reconstruction effect with low error, high frame rate, low latency, and high stability by virtue of the innovative depth vision algorithm, improved hybrid optimization strategy, and adaptive spatio-temporal consistency correction technology. These advantages not only greatly improve the reconstruction accuracy of the system in complex dynamic scenes but also meet the actual application requirements such as real-time monitoring and dynamic analysis, fully proving the application prospects and market competitiveness of this invention in related fields.

[0181] The above are only the preferred specific embodiments of this invention, but the protection scope of this invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by this invention, according to the technical solution and inventive concept of this invention, makes equivalent substitutions or changes, and all should be covered within the protection scope of this invention.

Claims

1. A monocular video 4D human body and scene real-time three-dimensional reconstruction method and system, characterized in that It includes the following steps: S1. Obtain a video stream from a monocular video sensor and preprocess each frame image of the video stream; S2. Use a depth estimation network based on Vision Transformer to perform monocular depth estimation on each frame image of the preprocessed video stream, generate corresponding depth maps, and at the same time use a human pose estimation network based on a spatio-temporal graph network to perform human pose estimation and extract human key point information; S3. Use an improved whale optimization algorithm to globally search and update the key parameters in the depth estimation network, and use a moth-flame optimization algorithm to locally search and optimize the key parameters in the human pose estimation network; S4. Collaboratively integrate the global search results and local search results obtained in step S3, and co-optimize the depth estimation network and the human pose estimation network through a joint update mechanism; S5. Perform multi-level data fusion on the depth map and human key point information to generate a preliminary 4D human and scene three-dimensional reconstruction model; S6. Use temporal modeling technology to perform dynamic reconstruction and spatio-temporal consistency correction on the continuous frame data in the preliminary 4D human and scene three-dimensional reconstruction model; S7. Output a 4D human and scene three-dimensional reconstruction model after overall parameter optimization and spatio-temporal correction.

2. The monocular video 4D human body and scene real-time three-dimensional reconstruction method according to claim 1, wherein, The specific content of S2 includes: S21, each frame image I in the preprocessed video stream t Input to the multi-scale feature extraction module based on visual Transformer to obtain the multi-scale feature representation F t ; S22. Based on multi-scale feature representation, use a vision Transformer encoder for depth estimation to generate a depth map D t ; S23. Input each frame image I in the preprocessed video stream t into the pose estimation module based on the spatio-temporal graph network simultaneously, and extract the local human body features L by using spatio-temporal relationship modeling t ; S24. Process the local human feature L t through the graph convolution and spatio-temporal dynamic fusion module to generate the human key point information K t ; S25. Use the multi-head attention mechanism to interact and fuse the depth map D t with the human key point information K t to generate the joint feature J t .

3. The monocular video 4D human body and scene real-time three-dimensional reconstruction method according to claim 1, characterized in that, The specific content of S3 includes: S31. Set the key parameters of the depth estimation network to W enc and the fitness function F enc (W enc ), and initialize the parameter candidate solution set N represents the size of the candidate solution set; S32. Use the improved whale optimization algorithm to globally search and update the set of candidate parameter solutions ; S33. Set the key parameters of the human pose estimation network to W fusion and the fitness function F fusion (W fusion ), and initialize the set of candidate solutions for the parameters M represents the number of candidate solutions for the parameters of the human pose estimation network; S34. Use the moth - flame optimization algorithm to perform local search optimization on the set of parameter candidate solutions and update each candidate solution: Among them, W flame represents the current local optimal parameter, γ and δ are control coefficients, and r j is a random number between 0 and 1, is the updated candidate solution, and exp is the exponential function; S35. According to the evaluation results of the fitness function F enc and F fusion respectively select the optimal parameters and from the updated candidate solutions as the final optimized parameters of the depth estimation network and the human pose estimation network.

4. The monocular video 4D human body and scene real-time three-dimensional reconstruction method according to claim 3, wherein, The specific content of S32 includes: S321. Set the key parameters of the depth estimation network to W enc , and initialize the candidate solution set where N represents the total number of candidate solutions and the initial velocity v corresponding to each candidate solution i (0), and initialize the global optimal parameter W best ; S322. Set the iteration count \(t = 1\) and the maximum number of iterations \(T\). max , and define the adaptive shrinkage coefficient where \(a_0\) is the preset initial shrinkage coefficient. S323. For each candidate solution calculate the Euclidean distance from the global optimal solution W best ​ S324. Introduce acceleration update based on a damped resonance model to simulate the accelerated pursuit behavior adopted by whales during predation. When the candidate solution is close to the global optimal solution, it has both spiral randomness and a direct pursuit effect similar to damped acceleration. Calculate the velocity of candidate solution i: where μ is the damping factor, κ is the acceleration constant, and v i (t + 1) represents the updated velocity value of candidate solution i at the (t + 1)-th iteration, and v i (t) represents the velocity of candidate solution i at the t-th iteration. The position of the candidate solution is calculated based on the updated velocity: Among them, represents the new position of candidate solution i after the acceleration update using the damped harmonic model; S325. Combine spiral update and damped harmonic update, and use weighted fusion to update each candidate solution as follows: where ω is the fusion weight and β is the sine update control coefficient, is the updated candidate solution, and sin is the sine function; S326. For the updated candidate solution set Evaluate according to the fitness function F enc and select the candidate solution that makes F enc reach the optimal value as the new global optimal parameter W best , update the iteration count t←t + 1 until t≥T max and then terminate the update and output the optimized depth estimation network parameters S327. Use a chaotic mapping mechanism to adaptively regulate the fusion weight ω and the contraction coefficient a(t): Set the initial chaotic variable C(0) and use the Logistic chaotic mapping: C(t + 1) = μ C ·C(t)·(1 - C(t)); Among them, C(t + 1) represents the chaotic variable at the next iteration t + 1, C(t) represents the chaotic variable at the iteration number t, and μ C represents the parameter that controls the chaotic behavior; Use the current chaotic variable C(t) to update the fusion weight and the contraction coefficient: ω(t) = ω min + (ω max - ω min )·C(t); a(t) = a0·C(t); Among them, ω min and ω max are the minimum and maximum values of the fusion weight respectively, a0 is the initial contraction coefficient, and ω(t) represents the fusion weight at the iteration number t; In each iteration, the updated fusion weights and contraction coefficients are used to replace the original parameters for updating each candidate solution in S325. Update is performed.

5. The monocular video 4D human body and scene real-time three-dimensional reconstruction method according to claim 1, wherein, The specific content of S4 includes: The candidate parameter sets after the S3 step update for the depth estimation network and the human pose estimation network are and where N represents the total number of candidate solutions for the depth estimation network parameters, and M represents the number of candidate solutions for the human pose estimation network parameters; S42. Define the combined fitness function F joint , and evaluate the overall performance of the parameter updates of the depth estimation network and the human pose estimation network: where λ is a trade-off factor, is the fitness function of the deep estimation network parameters : Among them, N s represents the total number of samples for evaluation, I j represents the j-th input image, represents that the depth estimation network based on Vision Transformer utilizes the parameter to generate the predicted depth for the image I j D, j represents the true depth data corresponding to the image I j ; Among them, is the fitness function of the human body pose estimation network parameters : Among them, N k represents the total number of image samples for evaluation, I l represents the l-th input image, represents the parameters used by the human pose estimation network based on the spatio-temporal graph network to generate the predicted human key point information for image I l K, l represents the true human key point information corresponding to image I l ; S43. For each pair in the candidate parameter set calculate the joint fitness value S44. According to the calculated combined fitness value, select the parameter pair from the candidate parameter set that makes the combined fitness function F joint reach the optimal value S45. Use a joint update mechanism to use the selected parameter pairs to simultaneously update the depth estimation network based on Vision Transformer and the human pose estimation network based on the spatio-temporal graph network to complete the co-optimization of the two networks.

6. The monocular video 4D human body and scene real-time three-dimensional reconstruction method according to claim 1, wherein, The specific content of S6 includes: S61. Set the continuous frame data sequence in the preliminary 4D human body and scene three-dimensional reconstruction model as where R t represents the reconstruction result of the t-th frame, and T represents the total number of frames; S62. Use the continuous frame data sequence as input and perform dynamic reconstruction through the temporal modeling network f TM to generate a temporal dynamic reconstruction result S63. Introduce adaptive spatio-temporal consistency correction to the dynamic reconstruction result of each frame, and calculate the temporal adaptation coefficient λ t :[[-END]] Among them, represents the temporal dynamic reconstruction result of the t-th frame, represents the temporal dynamic reconstruction result of the (t - 1)-th frame, and δ is a preset bias parameter. represents the Euclidean distance between consecutive frames, exp is the exponential function, and adaptive weighting is used to correct the temporal dynamic reconstruction result: Among them, represents the reconstructed result after correction of the previous frame, represents the reconstructed result after correction of the t-th frame; S64. Construct the spatio-temporal consistency metric function E temp Indicates the difference between consecutive frames: S65. According to the reconstructed result after calibration Output the final 4D human body and scene 3D reconstruction model after temporal modeling and spatio-temporal consistency calibration.

7. Monocular video 4D human body and scene real-time three-dimensional reconstruction system, the monocular video 4D human body and scene real-time three-dimensional reconstruction method according to any one of claims 1 to 6, characterized in that, It includes the following modules: A video acquisition and preprocessing module, which is used to obtain a video stream from a monocular video sensor and preprocess the video stream; A depth estimation module, which is used to perform monocular depth estimation and generate depth maps; A human pose estimation module, which is used to perform human pose estimation and extract key point information; A parameter optimization module, which is used to globally search and update the parameters of the depth estimation network, locally search and optimize the parameters of the human pose estimation network, and perform co-optimization through joint updates; A data fusion module, which is used to perform multi-level fusion of the depth map and human key point information according to preset rules to generate a preliminary 4D human and scene three-dimensional reconstruction model; A temporal modeling and correction module, which is used to use temporal modeling technology to perform dynamic reconstruction and spatio-temporal consistency correction on continuous frame data; Output display and interaction module, which is used to output the optimized 4D human body and scene 3D reconstruction model and provide real-time interaction and display functions.