Method for estimating non-confocal non-vision-field three-dimensional human body posture
By combining multi-view depth sensing and Transformer architecture with multi-scale time-frequency domain analysis, the problem of insufficient accuracy in 3D human pose estimation in non-confocal non-view imaging is solved, and high-precision pose reconstruction under extreme conditions is achieved.
Patent Information
- Application Number
- CN202510986387.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing non-view-of-sight imaging techniques lack an effective end-to-end deep learning framework under non-confocal conditions, making it difficult to achieve high-precision 3D human pose estimation, especially under conditions of extremely low signal-to-noise ratio and data compression, resulting in insufficient reconstruction accuracy.
Multi-view depth sensing devices are used to capture three-dimensional geometric information of the human body, constructing a depth map dataset. The Transformer architecture is used to model joint relationships and temporal correlations in a high-dimensional feature space. Photon data features are extracted through multi-scale time-frequency domain analysis and an improved non-local module. A composite loss function optimization model is established, and an ICCD detector is used to improve signal acquisition efficiency. Attitude estimation is optimized through optimal control theory.
Under conditions of strong noise and data compression, millimeter-level accuracy in 3D human pose reconstruction was achieved, significantly improving the accuracy and robustness of pose estimation in non-view-of-sight scenarios.
Smart Images

Figure CN120873681A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of computational photography, deep learning, and non-view imaging technology, and particularly relates to a method for non-confocal, non-view three-dimensional human pose estimation. Background Technology
[0002] Traditional optical imaging techniques typically require the observed target to be within the direct line of sight of the imaging system, a fundamental requirement that significantly limits their application scenarios. The emergence of Non-Line-of-Sight Imaging (NLOS) technology, however, fundamentally overcomes this physical limitation, enabling the detection and imaging of objects outside the line of sight through indirect optical path reconstruction. This revolutionary technological advancement not only significantly expands the applicable scope of imaging systems but also brings groundbreaking application possibilities to several cutting-edge fields such as autonomous driving, medical imaging, emergency rescue, and advanced manufacturing.
[0003] While non-line-of-sight (NLOS) imaging technology overcomes the line-of-sight limitations of traditional optical imaging, enabling the visual detection of hidden scenes, the raw transient measurement data it acquires possesses several inherent characteristics, posing significant technical challenges to 3D human motion posture reconstruction. Specifically, current NLOS imaging systems suffer from severe energy attenuation due to multiple photon scattering, resulting in a substantial reduction in effective signal strength. When a signal propagates to an interfacial surface, due to the diffuse reflection characteristics of the interface itself, the signal energy scatters and diffuses in a hemispherical pattern in all directions, with only a very small portion being captured by the imaging device. This leads to the receiver acquiring extremely weak signals from hidden scenes in NLOS perception systems, requiring related signal processing algorithms to achieve scene reconstruction and other perception functions in an environment with extremely low signal-to-noise ratios. Simultaneously, the geometric feature information of concealed targets is also severely attenuated during diffuse reflection. Reconstructing the original 3D structural information of the target from the multi-scattered signal constitutes a significant technical challenge in this field. Furthermore, the long signal acquisition time and low efficiency are also major challenges, necessitating the exploration of array-based detection methods to improve data acquisition efficiency.
[0004] In recent years, non-view-of-sight (NVO) imaging technology has gradually become a research hotspot in the field of computer vision. Achieving high-precision, high-resolution reconstruction of hidden scenes has become a key issue in this area. Existing research shows that NVOCC reconstruction algorithms can be mainly divided into two directions: one is analytical methods based entirely on mathematical and physical models, and the other adopts a model-driven and data-driven approach, integrating traditional algorithms with deep learning techniques to improve reconstruction performance. The O'Toole team innovatively proposed the "Light Cone Transform (LCT)" algorithm based on the NLOS optical forward model. This technique effectively achieves fast, high-resolution reconstruction in long-distance, high-noise environments by performing temporal nonlinear resampling and frequency-domain Wiener filtering on the time-resolved photon detection histogram. Compared to traditional methods, the LCT algorithm significantly improves imaging distance and accuracy while maintaining conventional hardware configurations, and reduces computational complexity and laser power requirements. Isogawa et al. proposed a deep learning-based NLOS human pose estimation method. This method uses multi-frame photon transient data acquired by a SPAD sensor as input. First, a feature extraction network encodes the raw data into deep feature vectors. Then, an LSTM network is used to dynamically fuse temporal features, ultimately outputting a sequence of human poses in the hidden scene. Notably, this research innovatively integrates a P2PSFNet structure and a ResNet18 module into the traditional LCT algorithm framework. By introducing trainable parameters, it significantly improves the algorithm's adaptability to input data and its robustness to noise. Current research on non-view-of-sight imaging mainly focuses on confocal system configurations. However, in non-confocal scenarios, the traditional Light Cone Transform (LCT) algorithm has significant limitations, lacking an effective end-to-end deep learning framework to address the crucial issue of 3D human pose estimation under non-confocal conditions. Summary of the Invention
[0005] Objective of the Invention: The technical problem to be solved by the present invention is to address the shortcomings of the prior art by providing a method for non-confocal, non-viewpoint 3D human pose estimation, comprising the following steps:
[0006] Step 1: Use multi-view depth sensing devices to capture the three-dimensional geometric information of the human body and construct a depth map dataset of the hidden human body scene for simulating transient photon data.
[0007] Step 2: Simulate the intermediate surface photon histogram corresponding to each depth map based on the intermediate surface light field distribution model under non-confocal conditions. This is used to simulate the photon spatiotemporal distribution measurement results obtained by the real non-confocal non-line-of-view imaging system.
[0008] Step 3: In the spatial dimension, design a Transformer architecture. In a high-dimensional feature space of 32×32×1024, adopt a hierarchical position encoding strategy to comprehensively model the human joint relationships within each frame and the temporal correlation between frames.
[0009] Step 4: Establish a multi-scale time-domain and frequency-domain joint analysis framework. Extract spectral features of photon data in parallel at three different scales: 4×4×4, 8×8×8, and 16×16×16 through 3D discrete cosine transform operation. Finally, perform weighted fusion to achieve cross-scale feature complementarity in the frequency domain. In addition, introduce an improved non-local module to capture the long-distance dependence of transient photon data in the frequency domain and achieve global feature extraction. Finally, fuse the time-domain and frequency-domain features through a cross-attention mechanism to output transient features.
[0010] Step 5: To optimize the model training process, a composite loss function is established, including pose accuracy loss based on joint coordinate error, discriminative loss based on classification supervision, and spatiotemporal smoothness loss. By dynamically weighting and fusing the loss term, the model training process is optimized and the accuracy and robustness of human pose estimation in non-visual scenes are improved.
[0011] Step 6: Use a humanoid strategy conditioned on transient features to control the humanoid in the physics simulator MuJoCo and generate a pose sequence. The method has been verified to quantitatively and accurately estimate the 3D human pose time series. Even in challenging scenarios with strong noise and data compression (sampling rate reduced to 20%), it can still maintain millimeter-level pose reconstruction accuracy.
[0012] In step 1, in terms of hardware configuration, four Kinect for Windows v2 depth sensing devices are deployed to collect human body data synchronously in a surround layout to ensure coverage without blind spots.
[0013] Each depth sensing device maintains synchronized data acquisition, with depth map accuracy down to the millimeter level, recording the distance information from objects to the depth sensing devices.
[0014] Step 2 includes:
[0015] In a non-line-of-sight imaging system based on the principle of active detection, the imaging device is placed on one side of the obstruction, and the target object to be detected is located on the other side of the obstruction. This area is usually referred to as the invisible area or hidden space.
[0016] The non-line-of-sight imaging system includes an active light source capable of emitting precisely controlled light signals and a high-sensitivity photoelectric signal receiving device.
[0017] The active light source uses a dedicated laser emitter with picosecond-level pulse width;
[0018] The high-sensitivity photoelectric signal receiving device uses a high-speed image intensified camera (ICCD) with array detection capability as the detection device to improve the acquisition speed and efficiency.
[0019] The picosecond laser is positioned at point P in space. source (x source ,y source ,z source At position x, source y source z source Let x, y, and z represent the x, y, and z coordinates of a picosecond laser source, respectively. An idealized narrowband laser pulse sequence (in practical engineering implementations, the pulse width is typically on the order of picoseconds) is emitted at a fixed frequency towards a point on the intermediate surface. The initial energy of the light pulse at the time of emission is denoted as I. source During a single optical pulse emission cycle, the single optical pulse emitted by the laser travels a distance of r. s After flight, it reaches the irradiation point P on the intermediate surface. i (x i ,y i ,z i ), where x i y i z i Let x, y, and z represent the coordinates of the irradiated point on the intermediate surface, respectively, from P source To P i In the optical signal transmission path, the laser pulse always maintains its straight-line propagation characteristics without diffuse reflection. Furthermore, based on the simplification of the theoretical model, various energy loss factors that may exist in the transmission medium are temporarily ignored. Therefore, when the optical pulse reaches the illumination point P... i At that time, the energy I of the light pulse i Still I source :
[0020] I i =I source
[0021] At this time, the light pulse travels a spatial distance r. s for:
[0022]
[0023] When the light pulse reaches the irradiation point P on the intermediate surface i (x i ,y i ,z i After that, in the theoretical modeling process, the following settings are made: First, the energy absorption effect of the intermediate surface on the incident light signal is approximately negligible; second, the intermediate surface has idealized Lambertian reflection characteristics, which can achieve isotropic uniform scattering.
[0024] Based on the given conditions, the energy of the incident light pulse will be at the collision point P. i (x i ,yi ,z i Centered on ), it exhibits an isotropic and uniform diffusion distribution pattern in three-dimensional space, with a portion of the optical signal energy passing through a distance r. hid Reach the spatial coordinates P where the micro-surface of the target object is located within the hidden scene. hid (x hid ,y hid ,z hid ), where x hid y hid z hid Represent the x, y, and z coordinates of the illumination point on the target object within the hidden scene, and the energy I of the diffusely reflected light signal. hid The decay will be:
[0025]
[0026] When the light signal propagates to the surface of the hidden object, some of the energy is absorbed by the object at point P. hid The reflectance at that location is Alb(P) hid ), calculate the residual energy intensity I′ of the light signal after absorption by the micro-surface of the hidden object. hid :
[0027] I' hid =I hid Alb(P hid )
[0028] When the light signal reaches the illumination point P on the surface of the hidden object hid At that time, it will be in P hid The light scatters uniformly outwards from the center, with some of the scattered light signals traveling a distance r. v Finally, it reaches point P on the intermediate surface. v (x v ,y v ,z v During this process, the energy of the light signal will further attenuate, and the intensity I of the light signal will decrease. v Represented as:
[0029]
[0030] When the light signal propagates to point P on the intermediate surface v At that time, it will be based on point P. v Diffuse reflection occurs at the center and the light is scattered uniformly in all directions. A portion of the scattered light signal will be located at spatial coordinate P. iccd (x iccd ,y iccd ,z iccd The ICCD detector with array detection capability at position x captures the image, where x is located. iccd yiccd z iccd Let x, y, and z represent the x, y, and z coordinates of the illumination point on the target object within the hidden scene, respectively. At this time, the intensity I of the light signal received by the ICCD detector... iccd Represented as:
[0031]
[0032] During the entire process, the photon underwent three diffuse reflections, and the light signal emitted by the laser was I source Eventually decays to I iccd Received by the detector, the following formula is derived:
[0033]
[0034] In practical non-line-of-sight imaging detection systems, the ICCD detectors employed possess three key technological advantages: First, their time resolution is exceptionally high, with typical time accuracy reaching the picosecond (10⁻¹² seconds) level; second, the detector exhibits ultra-high optical signal sensitivity, effectively detecting extremely weak light signals under the quantum limit, specifically achieving precise detection and counting at the single-photon level; and third, it possesses array detection capabilities, enhancing acquisition speed and efficiency. Under ideal operating conditions, the ICCD detector can transmit its real-time acquired photon event information to the downstream signal processing system. This data essentially constitutes a time-dependent photon statistical distribution, specifically represented by a nanosecond or picosecond time unit as the x-axis and the number of photons detected per unit time as the y-axis, thus forming a single-dimensional photon counting histogram distribution with time resolution characteristics. hid (t), the simulation calculation method is as follows:
[0035] hist hid (t)=I iccd ·δ(t-τ hid )
[0036] τ hid =(r s +r hid +r v +r iccd ) / c
[0037] Where t = 0, 1, 2, 3, ..., T⁻¹; δ(·) is the Dirac function, used to characterize the idealized instantaneous response; T is the time length of the transient photon histogram; τ hidIt fully describes the flight time of an optical signal throughout the entire transmission link, that is, the flight time consumed by the optical signal in the complete process from the initial emission of the pulsed laser, the first reflection process to the intermediate surface, the second reflection process through the surface of the hidden target object, the third interaction with the intermediate reflecting surface, to the final reception.
[0038] In step 3, the spatial dimension adopts the Transformer architecture, and a spatial Transformer module is established to encode the local spatial correlation between 3D joints in a single frame of data. The spatial self-attention layer considers the position information of 3D joints and returns the latent feature representation of the current input transient photon data.
[0039] Step 4 includes: establishing a multi-scale frequency domain module for processing high-dimensional transient photon signals;
[0040] The multi-scale frequency domain module receives transient photon data of size 32×32×1024 as input. It extracts multi-granularity features through hierarchical frequency domain transformation, extracting spectral features of the photon data in parallel at three different scales: 4×4×4, 8×8×8, and 16×16×16 using 3D discrete cosine transform. Finally, these features are weighted and fused to achieve cross-scale feature complementarity in the frequency domain. For a size of N... x ×N y ×N z The input tensor f(x,y,z) is N, where N is N. x N y and N z Let x, y, and z be the magnitudes of the input 3D data, respectively. The specific calculation process of the 3D-DCT 3D Discrete Cosine Transform F(u,v,w) is as follows:
[0041]
[0042] The normalization coefficient C(k) satisfies:
[0043]
[0044] Where the intermediate parameter N∈{N x N y N z}, where k represents the coordinate position of the current frequency component in the frequency domain.
[0045] Step 4 also includes: introducing an improved non-local module to capture the long-range dependency of transient photon data in the frequency domain, achieving global feature extraction, where i represents the index of the output position, j represents all possible positions to be enumerated, and w(x i ,y i) is a function for calculating the correlation factors of i and j, g is a representation of the input signal, the response is normalized by C(x), and the non-local operation can take all positions into account to obtain long-range dependencies:
[0046]
[0047] Step 5 includes:
[0048] The attitude accuracy loss based on joint coordinate error is measured by the Mean Per-JointPositionError (MPJPE), and the corresponding loss function is L. MPJPE The direct constraint model calculates the Euclidean distance between the predicted joint coordinates and the true values. For each frame in the input sequence, the pose accuracy loss based on joint coordinate error first calculates the L2 norm distance between the predicted and true coordinates of all relevant nodes (the dataset used in this invention contains 25 standard human keypoints). Then, the average error of all joints is taken. The mathematical expression is:
[0049]
[0050] Where G represents the number of joints. p represents the predicted 3D joint coordinates. i Represents the actual three-dimensional joint coordinates;
[0051] To further improve the model's discriminative ability, a discriminative loss Li that incorporates classification supervision was introduced during training. Classify For the input human posture data, additional posture category labels (such as walking, jumping, sitting, waving, etc.) were added, and L was used to annotate the data. Classify This serves as a supervisory signal to constrain the network's optimization process. The loss function, by jointly optimizing the pose regression and classification tasks, forces the network to learn more discriminative pose feature representations while accurately predicting joint coordinates. Classify The mathematical expression is:
[0052]
[0053] Where c k Encoding for the true pose category. This represents the probability distribution of the categories predicted by the network.
[0054] To effectively utilize the temporal correlation information in video sequences, this invention proposes to improve pose estimation accuracy by modeling the motion continuity between video frames. Specifically, a spatiotemporal smoothness loss function L is designed. optThis function, by constraining the smoothness of pose changes between adjacent frames, not only significantly improves the temporal consistency of 3D pose estimation in video sequences, but also effectively alleviates the performance degradation caused by ignoring inter-frame correlation.
[0055]
[0056] in, Indicates the predicted pose in frame t. Predicted pose in frame t-1 The difference vector, Let be the vector corresponding to the actual pose change between frame t and frame (t-1), where The pose is the actual pose in frame t.
[0057] Step 6 includes:
[0058] We choose to model human motion as the outcome of optimal control governed by a reward function, and use a Markov decision process to establish a transient image sequence τ. 1:T To a pose sequence p 1:T The mapping of input state s t This includes steps 3 and 4, from the transient sequence τ 1:T Extracted time-domain and frequency-domain fusion features and the current human dynamic state z t The environment generates the next state s through physical simulation. t+1 It iterates through training based on the degree of consistency between the human body's three-dimensional posture and its real form.
[0059] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.
[0060] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.
[0061] Beneficial Results: This method exhibits the following performance advantages: First, in the task of 3D human pose time series reconstruction, this method achieves quantitative and accurate estimation, with reconstruction errors controlled at the millimeter level; second, even under extremely challenging test environments—including strong noise interference and severe data compression (sampling rate only 20% of the original data)—this method can still maintain millimeter-level reconstruction accuracy. This performance significantly outperforms other existing non-viewpoint pose estimation methods, demonstrating the strong robustness and reliability of this method in practical applications. Attached Figure Description
[0062] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0063] Figure 1 This is a flowchart of the method of the present invention.
[0064] Figure 2 This is a schematic diagram of the imaging system.
[0065] Figure 3 This is the result of the present invention in achieving three-dimensional human pose estimation. Detailed Implementation
[0066] like Figure 1 As shown, this embodiment of the invention provides a method for non-confocal, non-viewpoint 3D human pose estimation, comprising the following steps:
[0067] Step 1: Acquire 3D geometric information of the human body through a multi-view depth sensing system to construct a dataset containing depth maps of hidden scenes. In terms of hardware configuration, four high-precision depth sensing devices are spatially deployed in a circular array with 90-degree intervals to form a comprehensive data acquisition system, ensuring 360-degree coverage of the human target without blind spots. Secondly, a precise time synchronization module (synchronization accuracy <1ms) ensures strict alignment of the data acquisition timing among the multiple devices.
[0068] Step 2: Based on the physical characteristics of the non-confocal imaging system, a light field transmission modeling method is proposed to simulate the multiple scattering process of photons at the intermediate surface. The input depth map is converted into a high-fidelity time-resolved photon histogram to simulate the photon spatiotemporal distribution measurement results obtained by the real non-confocal non-line-of-view imaging system.
[0069] In non-line-of-sight imaging systems employing active detection technology, the optical detection device is placed on one side of an obstacle, while the target to be imaged is on the other side; this area is often referred to as the hidden space. The key components of this system include two core units: an active illumination device capable of generating precisely modulated light signals, and an optical signal detection module with ultra-high sensitivity. More specifically, the system uses a laser generator capable of emitting picosecond-level ultra-narrow pulses as the illumination source, while an ICCD capable of detecting single-photon-level signals and possessing array detection capabilities is used as the detection device at the signal receiving end. Assume the picosecond laser is located at spatial coordinate point P. source (x source ,y source ,z souce At a certain point on the intermediate surface, an idealized narrowband laser pulse sequence (with a pulse width on the picosecond scale in practical engineering applications) is emitted at a constant frequency. The initial emitted light pulse energy is denoted as I. sourceDuring a single pulse emission cycle, the laser pulse travels a straight-line distance r. s The irradiation point P on the intermediate surface is then reached. i (x i ,y i ,z i During this transmission process, the light pulse maintains ideal straight-line propagation characteristics and does not produce any diffuse reflection effects. To simplify the theoretical analysis model, various energy attenuation mechanisms of the light signal in the transmission medium are temporarily disregarded. Therefore, when the light pulse reaches the illumination point P... i At that time, its energy I i Still I source :
[0070] I i =I source
[0071] In addition, the spatial distance r that the light pulse travels at this time s
[0072]
[0073] When the light pulse reaches the irradiation point P on the intermediate surface i (x i ,y i ,z i Afterwards, when establishing the theoretical model, the following basic assumptions were introduced: First, the energy absorption effect of the intermediate surface on the incident light signal can be approximated as a negligible secondary factor; second, the interface has ideal Lambertian reflection characteristics, enabling a completely isotropic and uniform scattering distribution. Based on the above assumptions, the incident light pulse collides at point P on the medium surface... i (x i ,y i ,z i An ideal spherical wavefront radiation field is formed at a distance r, which has strictly isotropic characteristics. A portion of the optical signal energy passes through this field. hid Reach the spatial coordinates P where the micro-surface of the target object is located within the hidden scene. hid (x hid ,y hid ,z hid The energy of the light signal after diffuse reflection, I hid Will decay to
[0074]
[0075] When a light signal propagates to the surface of a hidden object, some of the energy is absorbed by the object. The depth of this scene at point P is obtained from a synthetic dataset. hid Reflectance Alb(P) at this location hidThen, the residual energy intensity I′ of the light signal after absorption by the microsurface can be calculated. hid ,
[0076] I' hid =I hid Alb(P hid )
[0077] When the light signal reaches the surface P of the hidden object hid At that time, it will be in P hid It scatters uniformly outwards from the center. A portion of the scattered light signal travels a distance r... v Finally, it reaches a point P on the intermediate surface. v (x v ,y v ,z v During this process, the energy of the optical signal will further attenuate, and its intensity can be expressed as:
[0078]
[0079] When the light signal propagates to point P on the intermediate surface v At that time, diffuse reflection will occur around that point, and the light will be scattered uniformly in all directions. A portion of the scattered light signal will be located at spatial coordinate P. iccd (x iccd ,y iccd ,z iccd The ICCD detector with array detection capability captures the signal at point ( ). At this time, the intensity I of the optical signal received by the ICCD is... iccd Represented as:
[0080]
[0081] In practical detection systems for non-line-of-sight imaging, the ICCD, as the core detector, exhibits three significant technical characteristics: First, the device possesses excellent time-resolution performance, with time measurement accuracy typically reaching the picosecond level; second, the detector boasts extremely high optical signal detection sensitivity, even capable of accurately detecting extremely weak optical signals approaching the quantum limit, specifically demonstrating its ability to accurately identify and record individual photon events; and third, it has array-based detection capabilities, enhancing acquisition speed and efficiency. Under optimal operating conditions, the ICCD detector can transmit real-time acquired photon event data to the subsequent processing system. This data essentially presents time-dependent photon statistical characteristics, that is, based on picosecond time intervals, counting the number of photons detected within each time unit, ultimately forming a time-resolved one-dimensional photon count distribution histogram (hist). hid (t). Its simulation calculation method is as follows:
[0082] hist hid (t)=Ispad ·δ(t-τ hid ), t=0,1,2,3,......T-1
[0083] τ hid =(r s +r hid +r v +r iccd ) / c
[0084] In this mathematical model, δ(·) is the Dirac function, used to describe the ideal instantaneous response characteristics of the system; parameter T represents the total time span of the transient photon histogram; τ hid This fully depicts the flight time history of the optical signal throughout the entire transmission path, specifically as follows: the accumulated time delay from the initial emission of the pulsed laser source, through the first reflection by the intermediate reflective surface, to the second reflection on the surface of the hidden target object, then back to the intermediate surface for the third interaction, and finally to the reception by the detector.
[0085] Step 3, as follows Figure 2 As shown, in terms of spatial dimension modeling, a spatial feature encoding module based on the Transformer structure is designed to extract local correlation features between 3D joints in a single frame of data. The spatial self-attention mechanism in this module generates a deep feature representation of the current frame by analyzing the 3D coordinate information of the joints.
[0086] Step 4, the specific steps are as follows:
[0087] like Figure 2 As shown, a multi-scale frequency domain feature extraction architecture is proposed for efficient processing of high-dimensional transient photon signals. This module takes 32×32×1024-dimensional transient photon data as input and employs a hierarchical frequency domain analysis method for multi-granularity feature learning. Specifically, the system simultaneously performs three-dimensional discrete cosine transform at three different resolution scales: 4×4×4, 8×8×8, and 16×16×16, extracting spectral feature representations at each scale. Finally, an adaptive weight fusion mechanism integrates the multi-scale frequency domain features, achieving complementary advantages across scales. For a size of N... x ×N y ×N z The specific calculation process of the 3D-DCT transformation F(u,v,w) of the input tensor f(x,y,z) is as follows:
[0088]
[0089] The normalization coefficient C(k) satisfies:
[0090]
[0091] Furthermore, an enhanced non-local computation module is proposed to model the global correlation of transient photon signals in the frequency domain, thereby achieving global contextual modeling of photon distribution characteristics. Here, i is the index of the output position, and j represents all possible positions to be enumerated. f is the function for calculating the correlation factor between i and j, and g is the representation of the input signal. The response is normalized using C(x). Therefore, the non-local computation mechanism can establish global correlation modeling between positions, enabling the learning of feature dependencies across long distances.
[0092]
[0093] Step 5, the specific steps are as follows:
[0094] To optimize the model training process, a composite loss function is proposed, which includes pose accuracy loss based on joint coordinate error, discriminative loss fused with classification supervision, and spatiotemporal smoothness loss. By dynamically weighting and fusing the above loss terms, the model training process is optimized and the accuracy and robustness of human pose estimation in non-visual scenes are improved.
[0095] The attitude accuracy loss based on joint coordinate error is specifically measured by the Mean Per-Joint Position Error (MPJPE), and its corresponding loss function is L. MPJPE The model directly constrains the Euclidean distance between the predicted joint coordinates and the ground truth. Specifically, for each frame in the input sequence, MPJPE Loss first calculates the L2 norm distance between the predicted and ground truth coordinates of all key points (the dataset used in this study contains 25 standard human key points), and then averages the errors of all joints. The mathematical expression is as follows:
[0096]
[0097] To further improve the model's discriminative ability, a discriminative loss L based on classification supervision was introduced during training. Classify Specifically, for the input human pose data, additional pose category labels (such as walking, jumping, sitting, waving, etc.) are added, and these labels are used as supervisory signals to constrain the network's optimization process. This loss function, by jointly optimizing the pose regression and classification tasks, forces the network to learn more discriminative pose feature representations while accurately predicting joint coordinates.
[0098] To effectively utilize the temporal correlation information in video sequences, this method proposes to improve pose estimation accuracy by modeling the motion continuity between video frames. Specifically, a temporally aware optimization objective function L is designed. optThis function, by constraining the smoothness of pose changes between adjacent frames, not only significantly improves the temporal consistency of 3D pose estimation in video sequences, but also effectively alleviates the performance degradation caused by ignoring inter-frame correlation.
[0099]
[0100] in, Let represent the difference vector between the predicted pose of frame t and frame t-1. This is the vector corresponding to the actual pose change between frame t and frame t-1.
[0101] Step 6, specifically the steps are as follows:
[0102] Human motion is modeled using the framework of optimal control theory, treating it as a dynamic system optimization process driven by a reward function. Based on Markov decision processes (MDPs), a model is constructed from the transient image sequence τ. 1:T To human pose sequence p 1:T The mapping relationship. Input state s t This includes extracting features from the transient sequence τ through the feature extraction network in steps 3 and 4. 1:T Extracted time-domain and frequency-domain fusion features and the current human dynamic state z t The environment generates the next state s through physical simulation. t+1 During training, parameter tuning is achieved by continuously reducing the deviation between the predicted 3D pose and the actual motion capture data.
[0103] like Figure 3 As shown, the 3D pose estimation results of this invention are presented intuitively through a top-to-bottom comparison. The top image is the input depth map, i.e., the raw depth data collected by the depth sensor, which uses grayscale gradients to represent the distance information between objects and the camera device in the scene; the bottom image is the corresponding 3D pose estimation visualization result, in which the spatial pose of the target object is reconstructed by connecting the skeletal joints.
[0104] This invention provides a method for non-confocal, non-viewpoint 3D human pose estimation. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A method for non-confocal, non-viewpoint 3D human pose estimation, characterized in that, Includes the following steps: Step 1: Use multi-view depth sensing devices to capture the three-dimensional geometric information of the human body and construct a depth map dataset of the hidden human body scene for simulating transient photon data. Step 2: Simulate the intermediate surface photon histogram corresponding to each depth map based on the intermediate surface light field distribution model under non-confocal conditions. This is used to simulate the photon spatiotemporal distribution measurement results obtained by the real non-confocal non-line-of-view imaging system. Step 3: In the spatial dimension, design a Transformer architecture and adopt a hierarchical positional encoding strategy in the high-dimensional feature space to comprehensively model the human joint relationships within each frame and the temporal correlations between frames. Step 4: Establish a multi-scale time-domain and frequency-domain joint analysis framework. Extract spectral features of photon data in parallel at different scales through 3D discrete cosine transform operation. Finally, perform weighted fusion to achieve cross-scale feature complementarity in the frequency domain. In addition, introduce an improved non-local module to capture the long-distance dependence of transient photon data in the frequency domain and achieve global feature extraction. Finally, fuse the time-domain and frequency-domain features through a cross-attention mechanism to output transient features. Step 5: To optimize the model training process, a composite loss function is established, including pose accuracy loss based on joint coordinate error, discriminative loss based on classification supervision, and spatiotemporal smoothness loss. By dynamically weighting and fusing the loss term, the model training process is optimized and the accuracy and robustness of human pose estimation in non-visual scenes are improved. Step 6: Use a humanoid strategy conditioned on transient features to control the humanoids in the physics simulator MuJoCo and generate pose sequences.
2. The method according to claim 1, characterized in that, In step 1, in terms of hardware configuration, four depth sensing devices are deployed to collect human body data synchronously in a surround layout. Each depth sensing device maintains synchronized data acquisition, with depth map accuracy down to the millimeter level, recording the distance information from objects to the depth sensing devices.
3. The method according to claim 2, characterized in that, Step 2 includes: The non-line-of-sight imaging system includes an active light source capable of emitting precisely controlled light signals and a high-sensitivity photoelectric signal receiving device. The active light source uses a dedicated laser emitter with picosecond-level pulse width; The high-sensitivity photoelectric signal receiving device uses a high-speed image intensifier camera ICCD with array detection capability as the detection device. The picosecond laser is positioned at point P in space. source (x source ,y source ,z source At position ) where x source y source z source Let I represent the x, y, and z coordinates of a picosecond laser source. An idealized narrowband laser pulse sequence is emitted at a fixed frequency towards a point on an intermediate surface. The initial energy of the light pulse at the time of emission is denoted as I. source During a single optical pulse emission cycle, the single optical pulse emitted by the laser travels a distance of r. s After flight, it reaches the irradiation point P on the intermediate surface. i (x i ,y i ,z i ), where x i y i z i Let x, y, and z represent the coordinates of the irradiated point on the intermediate surface, respectively, from P source To P i In the optical signal transmission path, the laser pulse always maintains its straight-line propagation characteristics and does not undergo diffuse reflection. When the optical pulse reaches the illumination point P... i At that time, the energy I of the light pulse i Still I source : I i =I source At this time, the light pulse travels a spatial distance r. s for: When the light pulse reaches the irradiation point P on the intermediate surface i (x i ,y i ,z i After that, in the theoretical modeling process, the following settings are made: First, the energy absorption effect of the intermediate surface on the incident light signal is approximately negligible; second, the intermediate surface has idealized Lambertian reflection characteristics, which can achieve isotropic uniform scattering. Based on the given conditions, the energy of the incident light pulse will be at the collision point P. i (x i ,y i ,z i Centered on ), it exhibits an isotropic and uniform diffusion distribution pattern in three-dimensional space, with a portion of the optical signal energy passing through a distance r. hid Reach the spatial coordinates P where the micro-surface of the target object is located within the hidden scene. hid (x hid ,y hid ,z hid ), where x hid y hid z hid Represent the x, y, and z coordinates of the illumination point on the target object within the hidden scene, and the energy I of the diffusely reflected light signal. hid The decay will be: When the light signal propagates to the surface of the hidden object, some of the energy is absorbed by the object at point P. hid The reflectance at that location is Alb(P) hid ), calculate the residual energy intensity I′ of the light signal after absorption by the micro-surface of the hidden object. hid : I′ hid =I hid White(P hid ) When the light signal reaches the illumination point P on the surface of the hidden object hid At that time, it will be in P hid The light scatters uniformly outwards from the center, with some of the scattered light signals traveling a distance r. v Finally, it reaches point P on the intermediate surface. v (x v ,y v ,z v The intensity of the light signal I v Represented as: When the light signal propagates to point P on the intermediate surface v At that time, it will be based on point P. v Diffuse reflection occurs at the center and the light is scattered uniformly in all directions. A portion of the scattered light signal will be located at spatial coordinate P. iccd (x iccd ,y iccd ,z iccd The ICCD detector with array detection capability at position x captures the image, where x is located. iccd y iccd z iccd Let x, y, and z represent the x, y, and z coordinates of the illumination point on the target object within the hidden scene, respectively. At this time, the intensity I of the light signal received by the ICCD detector... iccd Represented as: During the entire process, the photon underwent three diffuse reflections, and the light signal emitted by the laser was I source Eventually decays to I iccd Received by the detector, the following formula is derived: Using nanosecond or picosecond time units as the x-axis and the number of photons detected per unit time as the y-axis, a single-dimensional photon counting histogram distribution with time resolution is formed. hid (t), the simulation calculation method is as follows: hist hid (t)=I iccd ·δ(t-τ hid ) τ hid =(r s +r hid +r v +r iccd ) / c Where t = 0, 1, 2, 3, ..., T⁻¹; δ(·) is the Dirac function; T is the time length of the transient photon histogram; τ hid It fully describes the flight time of an optical signal throughout the entire transmission link, that is, the flight time consumed by the optical signal in the complete process from the initial emission of the pulsed laser, the first reflection process to the intermediate surface, the second reflection process through the surface of the hidden target object, the third interaction with the intermediate reflecting surface, to the final reception.
4. The method according to claim 3, characterized in that, In step 3, the spatial dimension adopts the Transformer architecture, and a spatial Transformer module is established to encode the local spatial correlation between 3D joints in a single frame of data. The spatial self-attention layer considers the position information of 3D joints and returns the latent feature representation of the current input transient photon data.
5. The method according to claim 4, characterized in that, Step 4 includes: establishing a multi-scale frequency domain module for processing high-dimensional transient photon signals; The multi-scale frequency domain module receives transient photon data of size 32×32×1024 as input. It extracts multi-granularity features through hierarchical frequency domain transformation, extracting spectral features of the photon data in parallel at three different scales: 4×4×4, 8×8×8, and 16×16×16 using 3D discrete cosine transform. Finally, these features are weighted and fused to achieve cross-scale feature complementarity in the frequency domain. For a size of N... x ×N y ×N z The input tensor f(x,y,z) is N, where N is N. x N y and N z Let x, y, and z be the magnitudes of the input 3D data, respectively. The specific calculation process of the 3D-DCT 3D Discrete Cosine Transform F(u,v,w) is as follows: The normalization coefficient C(k) satisfies: Where the intermediate parameter N∈{N x N y N z }, where k represents the coordinate position of the current frequency component in the frequency domain.
6. The method according to claim 5, characterized in that, Step 4 also includes: introducing an improved non-local module to capture the long-range dependency of transient photon data in the frequency domain, achieving global feature extraction, where i represents the index of the output position, j represents all possible positions to be enumerated, and w(x i ,y i ) is a function for calculating the correlation factors of i and j, g is used as a representation of the input signal, and the response is normalized through C(x) to obtain the long-range dependency:
7. The method according to claim 6, characterized in that, Step 5 includes: The attitude accuracy loss based on joint coordinate error is measured by the average position error per joint, corresponding to the loss function L. MPJPE The direct constraint model calculates the Euclidean distance between the predicted joint coordinates and the true values. For each frame in the input sequence, the pose accuracy loss based on joint coordinate error is first calculated by the L2 norm distance between the predicted and true 3D coordinates of all joints, and then the average error of all joints is taken. The mathematical expression is: Where G represents the number of joints. p represents the predicted 3D joint coordinates. i Represents the actual three-dimensional joint coordinates; A discriminative loss L, which incorporates classification supervision, was introduced during training. Classify For the input human posture data, additional posture category labels were added, and L was assigned to the corresponding labels. Classify As the optimization process of the supervised signal constrained network, L Classify The mathematical expression is: Where c k Encoding for the true pose category. This represents the probability distribution of the categories predicted by the network. Design a spatiotemporal smoothness loss function L opt : in, Indicates the predicted pose in frame t. Predicted pose in frame t-1 The difference vector, Let be the vector corresponding to the actual pose change between frame t and frame (t-1), where The pose is the actual pose in frame t.
8. The method according to claim 7, characterized in that, Step 6 includes: We choose to model human motion as the outcome of optimal control governed by a reward function, and use a Markov decision process to establish a transient image sequence τ. 1:T To a pose sequence p 1:T The mapping of input state s t This includes steps 3 and 4, from the transient sequence τ 1:T Extracted time-domain and frequency-domain fusion features and the current human dynamic state z t The environment generates the next state s through physical simulation. t+1 It iterates through training based on the degree of consistency between the human body's three-dimensional posture and its real form.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 8.