A method and system for multi-view 3D astronaut pose estimation in spacecraft

By employing a multi-view 3D astronaut posture estimation method, utilizing human segment models and dynamic centroid positioning, combined with fine-grained voxel subspace and multi-plane adaptive weighted fusion, the problems of poor adaptability and insufficient occlusion robustness of 3D posture estimation within spacecraft cabins are solved, achieving high-precision and stable 3D human posture output.

CN122492810APending Publication Date: 2026-07-31INNOVATION ACAD FOR MICROSATELLITES OF CAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INNOVATION ACAD FOR MICROSATELLITES OF CAS
Filing Date
2026-04-14
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing spacecraft cabin attitude perception technologies suffer from poor adaptability, insufficient robustness in obstructed scenarios, low accuracy in 3D attitude estimation, and limited real-time performance, making it difficult to meet the high-precision, stable, and real-time monitoring requirements in microgravity environments.

Method used

A multi-view 3D astronaut pose estimation method is adopted. By acquiring synchronous multi-view human image sequences, a human segment model is constructed, human centroid encoding and dynamic centroid estimation are performed, and high-precision 3D human pose results are output by combining fine-grained voxel subspace and multi-plane adaptive weighted fusion.

Benefits of technology

It improves the accuracy and stability of 3D human pose estimation in microgravity environments, enhances robustness under occlusion and complex working conditions, reduces voxel discretization error, and has good engineering deployment capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492810A_ABST
    Figure CN122492810A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for multi-view 3D astronaut posture estimation in spacecraft. The method involves: acquiring 3D voxel feature volumes containing joint confidence information; performing human centroid encoding and dynamic centroid estimation based on a human segment model; forming a human body space description based on the 3D voxel feature volumes; fusing the dynamic human centroid localization results with the human body space description to obtain accurate human body positioning that adaptively adjusts with changes in human posture; constructing a fine-grained voxel subspace and extracting scale-aware features based on the accurate human body positioning; projecting the fine-grained voxel subspace and scale-aware features onto multiple orthogonal planes, regressing 2D heatmaps of human joints on each orthogonal plane to obtain joint prediction results for each orthogonal plane; and learning the weights of different planes in the 3D reconstruction process based on the confidence of the prediction results, adaptively weighting and fusing the joint coordinates to obtain the 3D human posture result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent sensing and human-computer interaction technology within spacecraft cabins, specifically to a multi-view, three-dimensional astronaut attitude estimation method and system for spacecraft. It also relates to a corresponding computer device and computer-readable storage medium. Background Technology

[0002] Currently, although several technical solutions exist for personnel attitude perception and three-dimensional spatial perception within spacecraft cabins, there are still significant shortcomings and limitations, including the following main problems: 1. Reliance on contact-based sensing devices, resulting in poor adaptability. Some existing attitude sensing solutions rely on wearable inertial measurement units (IMUs), pressure sensors, or magnetic induction devices to obtain human attitude information. These solutions significantly interfere with and restrict astronaut behavior, making it difficult to meet the requirements for free movement of astronauts in the microgravity environment inside the spacecraft cabin. At the same time, calibration and drift issues of wearable devices also lead to long-term instability in attitude estimation accuracy.

[0003] 2. Traditional vision solutions rely solely on single-viewpoint or low-dimensional features, resulting in insufficient accuracy in 3D reconstruction. Existing human pose estimation schemes based on monocular cameras often only provide two-dimensional skeleton data, failing to reliably obtain three-dimensional spatial coordinates. While some multi-camera layout schemes achieve 3D reconstruction, they often rely on low-precision geometric calibration or simple triangulation algorithms. Consequently, accuracy drops sharply under conditions of occlusion, rapid movement, and complex pose changes, failing to meet the requirements for high-precision 3D positioning and attitude description in aerospace applications.

[0004] 3. Poor adaptability to complex environments. Existing visual perception systems exhibit significant performance degradation in complex scenarios such as changing lighting, cluttered backgrounds, and severe occlusion. In particular, under conditions of confined space, dense objects, and frequent occlusion in microgravity cabins, serious phenomena such as false detections, missed detections, and misconnected skeletons often occur, making it impossible to provide stable and reliable attitude data.

[0005] 4. Lack of effective utilization of multi-view data consistency and geometric constraints. Most existing solutions lack systematic view geometry modeling and optimization strategies, simply fusing 2D detection results from multiple cameras, ignoring spatial consistency between multiple views, camera calibration errors, and joint optimization for 3D pose reconstruction, resulting in significant errors in 3D estimation.

[0006] 5. Real-time performance and computational efficiency cannot meet the real-time online monitoring requirements of the space station. Current accurate 3D attitude reconstruction schemes have high computational complexity, requiring numerous post-processing steps or offline optimization. This makes it difficult to achieve real-time, high frame rate operation on resource-constrained aerospace computers or edge computing devices, reducing the system's practicality and deployability.

[0007] 6. Lack of unified modeling for local details and overall three-dimensional relationships of the human body. Traditional solutions often focus on the treatment of limb joints, but lack comprehensive modeling of the overall spatial geometry of the human body, joint motion constraints, and dynamic consistency. This results in three-dimensional human body models inferred from multiple perspectives having skeletal structures that do not conform to the laws of human kinematics and are unstable.

[0008] 7. Difficulty in supporting further intelligent analysis and behavior understanding. Existing attitude perception methods mainly output skeleton points or simple position information, lacking the ability to extract and correlate higher-level features such as action categories, task semantics, and behavioral intentions. Therefore, they cannot meet the high-level intelligent analysis requirements in aerospace missions, such as automatic behavior recognition, emergency alarms, and collaborative mission evaluation. Summary of the Invention

[0009] To address the aforementioned shortcomings in existing technologies, this invention provides a multi-view, three-dimensional astronaut attitude estimation method and system for spacecraft. It also provides a corresponding computer device and computer-readable storage medium.

[0010] According to one aspect of the present invention, a multi-view three-dimensional astronaut pose estimation method for spacecraft is provided, comprising: Acquire a synchronized multi-view human image sequence, extract two-dimensional heat maps of each joint point of the human body based on each human image, and convert them into three-dimensional voxel feature volumes containing joint point confidence information to form an initial three-dimensional feature voxel space. A human segment model is constructed, and human centroid encoding and dynamic centroid estimation are performed based on the human segment model to obtain the dynamic human centroid localization result. Based on the three-dimensional voxel feature, the spatial occupancy parameters of the human body are regressed on the projection plane to form an initial description of the human body's spatial occupancy. By fusing the dynamic human centroid positioning result with the initial human body space description, an accurate human body space that adaptively adjusts with changes in human posture is obtained, resulting in an optimized human body space. Based on the optimized human body occupancy, a fine-grained voxel subspace is constructed and human joint scale perception features are extracted. The fine-grained voxel subspace and human joint scale perception features are projected onto multiple orthogonal planes. Two-dimensional heat maps of human joints are regressed on each orthogonal plane, and the coordinates of the joints on each plane are calculated based on the centroid of the heat map to obtain the joint prediction results for each orthogonal plane. Based on the confidence level of the joint prediction results of each orthogonal plane, the weights of different planes in the 3D reconstruction process are learned, and the joint coordinates of multiple planes are adaptively weighted and fused to output the final 3D human pose result.

[0011] Preferably, the step of acquiring a synchronized multi-view human image sequence, extracting two-dimensional heatmaps of each joint point of the human body based on each human image, and converting them into a three-dimensional voxel feature volume containing joint point confidence information to form an initial three-dimensional feature voxel space includes: By using multiple fixed-view image acquisition devices, a synchronous multi-view human image sequence is obtained, and the intrinsic and extrinsic parameters of each image acquisition device are calibrated to establish a unified world coordinate system. For the images acquired by each image acquisition device, two-dimensional heat map information of each joint point of the human body is extracted to obtain the confidence distribution of two-dimensional joint points from multiple perspectives. Based on the calibration parameters, the two-dimensional heatmaps of each joint point under each viewpoint are back-projected to a unified three-dimensional space to construct a three-dimensional voxel feature body containing joint point confidence information, thus forming an initial three-dimensional feature voxel space.

[0012] Preferably, the step of constructing a human segment model, and performing human centroid encoding and dynamic centroid estimation based on the human segment model to obtain dynamic human centroid localization results includes: Based on the topological relationships of human joints, the human body is divided into multiple segments; Assign a corresponding mass coefficient and proximal and distal parameters to each body segment to construct a human segment model; Based on the aforementioned human segment model, the overall center of mass of the human body is calculated using the segment torque synthesis method. The overall centroid of the human body is represented as a probability distribution in three-dimensional space, and a human centroid code is generated. Using the human centroid encoding as a supervisory signal, the dynamic centroid position of the human body is regressed through a three-dimensional convolutional network, the dynamic centroid is estimated, and the dynamic human centroid localization result is obtained.

[0013] Preferably, the step of regressing the spatial occupancy parameters of the human body on the projection plane based on the three-dimensional voxel feature volume to form an initial human body spatial description includes: Based on the three-dimensional voxel feature, the spatial occupancy parameters of the human body are regressed on the projection plane, including: planar position, spatial size and height information, to form an initial human body occupancy space description based on the root node.

[0014] Preferably, the step of fusing the dynamic human centroid positioning result with the initial human body space description to obtain an accurate human body position that adaptively adjusts with changes in human posture includes: The dynamic human centroid localization result is fused with the initial human occupancy space result based on the root node. Through a non-maximum suppression strategy, an accurate human occupancy that adaptively adjusts with changes in human posture is obtained.

[0015] Preferably, the step of constructing a fine-grained voxel subspace and extracting human joint scale-sensing features based on the optimized human body occupancy includes: Based on the optimized human body space, the fine-grained voxel subspace corresponding to the human body is obtained by cropping the initial three-dimensional feature voxel space. Human joint feature extraction is performed on the fine-grained voxel subspace; A scale perception mechanism that integrates spatial attention and channel attention is introduced to adaptively weight and fuse different scale features to obtain human joint scale perception features.

[0016] According to a second aspect of the present invention, a multi-view three-dimensional astronaut attitude estimation system for spacecraft is provided, comprising: Image acquisition module, which is used to acquire synchronized multi-view human image sequences; The 3D voxel feature body construction module extracts two-dimensional heat maps of each joint point of the human body based on each human body image and converts them into three-dimensional voxel feature bodies containing joint point confidence information to form an initial three-dimensional feature voxel space. The module for human centroid coding and dynamic centroid estimation is used to construct a human segment model, perform human centroid coding and dynamic centroid estimation based on the human segment model, and obtain the dynamic human centroid localization result. The human body space description module, based on the three-dimensional voxel feature, regresses the spatial occupancy parameters of the human body on the projection plane to form an initial human body space description. The human body space optimization module is used to fuse the dynamic human body centroid positioning result with the initial human body space description to obtain an accurate human body space that is adaptively adjusted with changes in human body posture, thus obtaining an optimized human body space. A fine-grained voxel subspace construction module is used to construct a fine-grained voxel subspace and extract human joint scale perception features based on the optimized human body occupancy. The joint prediction module projects the fine-grained voxel subspace and human joint scale perception features onto multiple orthogonal planes, regresses two-dimensional heat maps of human joints on each orthogonal plane, and calculates the coordinates of the joints on each plane based on the centroid of the heat map to obtain the joint prediction results for each orthogonal plane. The 3D human pose output module is used to learn the weights of different planes in the 3D reconstruction process based on the confidence of the prediction results of the joints of each orthogonal plane, and to adaptively weight and fuse the joint coordinates of multiple planes to output the final 3D human pose result.

[0017] Preferably, the image acquisition module includes: multiple image acquisition devices with fixed viewing angles, and the multiple image acquisition devices are deployed inside the spacecraft cabin.

[0018] According to a third aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, can be used to perform the method described in any one of the above inventions.

[0019] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, can be used to perform the method described in any one of the above inventions.

[0020] By adopting the above technical solution, the present invention has at least one of the following beneficial effects compared with the prior art: This invention improves the accuracy and stability of 3D human posture estimation in microgravity environments. By introducing a human segment model and a dynamic centroid positioning mechanism, the 3D posture estimation results better conform to human physical characteristics, significantly reducing posture drift and positioning errors.

[0021] Significantly enhances robustness under occlusion and complex conditions. This invention employs a multi-view voxel fusion and orthogonal plane adaptive weighting strategy, enabling stable output of complete human pose even when some viewpoints are missing or occluded.

[0022] This invention reduces voxel discretization errors and improves joint positioning accuracy. Through refined human body occupancy constraints and local voxel reconstruction, it reduces the involvement of irrelevant space in the calculation, thereby improving joint regression accuracy.

[0023] It possesses excellent engineering deployment and scalability capabilities. This invention adopts a non-contact visual perception solution, eliminating the need for astronauts to wear additional equipment. It is easily integrated into existing spacecraft cabin vision systems and can be extended to other human posture perception scenarios in constrained and complex environments. Attached Figure Description

[0024] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the multi-view three-dimensional astronaut attitude estimation method for spacecraft in a preferred embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram illustrating the hierarchical dependencies between steps in a preferred embodiment of the present invention.

[0026] Figure 3This is a flowchart illustrating a hierarchical workflow combining global localization and local fine estimation in a preferred embodiment of the present invention.

[0027] Figure 4 This is a flowchart illustrating the process of obtaining the centroid from a human segment model in a preferred embodiment of the present invention.

[0028] Figure 5 This is a flowchart of the dual-benchmark fusion process combining dynamic centroid and root node in a preferred embodiment of the present invention.

[0029] Figure 6 This is a diagram of a refined spatial working architecture that combines voxel subspace clipping and scale perception in a preferred embodiment of the present invention.

[0030] Figure 7 This is a diagram illustrating the working architecture of a multi-plane orthogonal projection structure with natural information complementarity, as shown in a preferred embodiment of the present invention.

[0031] Figure 8 This is a schematic diagram of the constituent modules of a multi-view three-dimensional astronaut attitude estimation system for spacecraft in a preferred embodiment of the present invention. Detailed Implementation

[0032] The embodiments of the present invention are described in detail below: These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.

[0033] Existing technologies for human attitude perception and 3D spatial awareness within spacecraft cabins typically suffer from the following technical problems: reliance on contact-based sensing devices, resulting in poor adaptability; traditional visual solutions based solely on single-viewpoint or low-dimensional features, leading to insufficient 3D reconstruction accuracy; weak adaptability to complex environments; lack of effective utilization of multi-view data consistency and geometric constraints; inability to meet the real-time and computational efficiency requirements for real-time online monitoring of space stations; lack of unified modeling of local human details and whole-body 3D relationships; and difficulty in supporting further intelligent analysis and behavioral understanding. Therefore, existing human attitude perception technologies within spacecraft cabins exhibit poor adaptability to microgravity environments, insufficient robustness to occlusion scenarios, low 3D attitude estimation accuracy, and limited real-time performance.

[0034] To address the aforementioned issues, one embodiment of the present invention provides a multi-view three-dimensional astronaut posture estimation method for spacecraft. Based on human segment model constraints, it implements multi-view voxel-based three-dimensional human posture estimation, achieving high-precision, stable, and non-contact perception of astronauts' three-dimensional posture and spatial positioning in a microgravity environment.

[0035] Specifically, such as Figure 1 As shown, the multi-view three-dimensional astronaut pose estimation method for spacecraft provided in this embodiment may include: M1 acquires a synchronized multi-view human image sequence, extracts two-dimensional heat maps of each joint point of the human body based on each human image, and converts them into three-dimensional voxel feature bodies containing joint point confidence information to form an initial three-dimensional feature voxel space. M2, construct a human segment model, and perform human centroid coding and dynamic centroid estimation based on the human segment model to obtain the dynamic human centroid localization result; M3, based on three-dimensional voxel features, regresses the spatial occupancy parameters of the human body on the projection plane to form an initial description of the human body's spatial occupancy. M4 integrates the dynamic human centroid localization results with the initial human body space description to obtain an accurate human body space that adapts to changes in human posture, resulting in an optimized human body space. M5 constructs a fine-grained voxel subspace based on the optimized human body occupancy and extracts human body joint scale perception features. M6 projects fine-grained voxel subspace and human joint scale perception features onto multiple orthogonal planes, regresses two-dimensional heatmaps of human joints on each orthogonal plane, and calculates the coordinates of joints on each plane based on the centroid of the heatmap to obtain the joint prediction results for each orthogonal plane. M7 learns the weights of different planes in the 3D reconstruction process based on the confidence of the prediction results of the joints of each orthogonal plane, performs adaptive weighted fusion of the joint coordinates of multiple planes, and outputs the final 3D human pose result.

[0036] In some preferred embodiments, M1, which acquires a synchronized multi-view human image sequence, extracts two-dimensional heatmaps of each joint point of the human body based on each human image, and converts them into a three-dimensional voxel feature body containing joint point confidence information to form an initial three-dimensional feature voxel space, may further include: M11 uses multiple fixed-view image acquisition devices to acquire synchronous multi-view human image sequences, and performs intrinsic and extrinsic parameter calibration on each image acquisition device to establish a unified world coordinate system. M12 extracts two-dimensional heat map information of each joint point of the human body from the images acquired by each image acquisition device, and obtains the confidence distribution of two-dimensional joint points from multiple perspectives. M13, based on calibration parameters, backprojects the two-dimensional heatmaps of each joint point from each viewpoint to a unified three-dimensional space, constructs a three-dimensional voxel feature body containing joint point confidence information, and forms an initial three-dimensional feature voxel space.

[0037] In some preferred embodiments, the above-mentioned M2, which constructs a human segment model, performs human centroid encoding and dynamic centroid estimation based on the human segment model to obtain dynamic human centroid localization results, may further include: M21 divides the human body into multiple segments based on the topological relationship of human joints; M22 assigns a corresponding mass coefficient and proximal and distal parameters to each body segment to construct a human segment model; M23, based on the human segment model, uses the segment torque synthesis method to calculate the overall center of mass of the human body; M24 represents the overall centroid of the human body as a probability distribution in three-dimensional space, generating a human centroid code. M25 uses the human centroid encoding as a supervision signal, regresses the dynamic centroid position of the human body through a three-dimensional convolutional network, estimates the dynamic centroid, and obtains the dynamic human centroid localization result.

[0038] In some preferred embodiments, M24, which represents the overall centroid of the human body as a probability distribution in three-dimensional space and generates a human centroid code, may further include: The density of different parts of the human body can be regarded as the same. The human body is studied in segments, and the continuous human body model is transformed into a continuous set of segments composed of different body segments. For the mass of each body segment, the continuous body segment is converted into a discrete particle system composed of the joints at the proximal and distal ends of the body segment; Based on the conclusions of statistical experiments on human biomechanics, weights are assigned to the particles of each body segment to solve for the centroid of the discrete particle system.

[0039] In some preferred embodiments, the M25 described above, which uses the human centroid encoding as a supervision signal and regresses the dynamic centroid position of the human body through a three-dimensional convolutional network, may further include: During training, human body segments are aggregated using M24 to establish a supervisory signal H. XYZ ; Considering the systematic errors in the calibration process and the random errors in the inference process, the human centroid is represented as a Gaussian distribution P with mean S and variance σ; By minimizing the true distribution of the three-dimensional human centroid (supervisory signal H) XYZ The distance loss between the network's predicted centroid and the actual centroid result is used to train a 3D convolutional neural network. The trained 3D convolutional neural network is used to regress the dynamic centroid position of the human body.

[0040] In some preferred embodiments, the above-mentioned M3, based on three-dimensional voxel features, regresses the spatial occupancy parameters of the human body on the projection plane to form an initial human body occupancy space description, and may further include: Based on three-dimensional voxel features, the joint offset loss of each plane is minimized by training a neural network, thereby regressing the spatial occupancy parameters of the human body on the projection plane, including: planar position, spatial size and height information, forming an initial human body occupancy spatial description based on the root node.

[0041] In some preferred embodiments, M4, which fuses the dynamic human centroid positioning result with the initial human body space description to obtain an accurate human body position that adaptively adjusts with changes in human posture, may further include: By fusing the dynamic human centroid localization results with the initial human occupancy space results based on the root node, and using a non-maximum suppression strategy, a precise human occupancy that adaptively adjusts with changes in human posture is obtained.

[0042] Furthermore, through a non-maximum suppression strategy, a precise human body occupancy space that adaptively adjusts with changes in human posture is obtained, including: M41 determines the size of the human body's horizontal space by calculating the maximum distance from the center point (including the root node and the dynamic human body's center of mass) to each joint. M42 calculates the human body space based on the dynamic human body centroid positioning result and the human body space based on the root node, respectively. M43 performs non-maximum suppression processing on the calculation results of the two to obtain the accurate human body occupancy.

[0043] In some preferred embodiments, the M5 described above, which constructs a fine-grained voxel subspace and extracts human joint scale perception features based on the optimized human body occupancy, may further include: M51, based on the optimized human body occupancy, cuts out the fine-grained voxel subspace (three-dimensional voxel subspace) corresponding to the human body in the initial three-dimensional feature voxel space. M52 performs multi-scale human joint feature extraction in fine-grained voxel subspace; M53 introduces a scale perception mechanism that integrates spatial attention and channel attention mechanisms. It adaptively weights and fuses human joint features at different scales to obtain the final human joint scale perception features.

[0044] In some preferred embodiments, the above-mentioned M51, which, based on the optimized human body occupancy, trims the fine-grained voxel subspace corresponding to the human body in the initial three-dimensional feature voxel space, may further include: M511, in the horizontal direction, the human body's occupied space is the maximum distance from the calculation center point (including the root node and the dynamic human body's centroid) to each joint, thus obtaining the size of the human body's horizontal occupied space. In M512, a height supervision signal H is constructed using the height value of the root node in the vertical direction. Z By performing nonmaximum suppression on the confidence level of the Gaussian distribution of the root node horizontal plane, candidate positions for height pruning are obtained. M513, by minimizing the height supervision signal H Z A one-dimensional convolutional neural network is trained using the distance difference between the height predicted by the network and the height predicted by the network. M514 uses a trained one-dimensional convolutional neural network to regress the height features corresponding to the candidate positions obtained by height clipping, and obtains the final height features. M515, based on the size of the horizontal space occupied by the human body and its final height characteristics, cuts out the fine-grained voxel subspace corresponding to the human body in the initial three-dimensional feature voxel space.

[0045] In some preferred embodiments, the above-mentioned M53 introduces a scale perception mechanism that integrates spatial attention and channel attention mechanisms to adaptively weight and fuse human joint features at different scales, and may further include: M531 preserves features at different scales from different network layers during neural network processing. M532, through channel attention and spatial attention mechanisms, enables the attention mechanism to guide feature extraction at all abstraction levels, enhances the network's ability to focus on key regions and channels of the task, and achieves adaptive weighted fusion of features at different scales.

[0046] In some preferred embodiments, M7, based on the confidence level of the joint point prediction results of each orthogonal plane, learns the weights of different planes in the 3D reconstruction process, and adaptively weights and fuses the joint point coordinates of multiple planes. It may further include: M71 decomposes the fine-grained voxel subspace into three planes, xy, xz, and yz, according to the spatial orthogonal plane. M72 projects the scale-sensing features of human joint points in fine-grained voxel subspace onto three orthogonal planes via orthogonal projection. M73 uses the true value of three-dimensional joints as the supervision signal and designs a two-dimensional neural network by minimizing the distance difference between the cross-plane fusion weight W and the product of the orthogonal plane joints and the supervision signal. The final 3D pose estimation result is calculated by employing a pairwise Softmax normalization strategy for the weights W. The following detailed description, in conjunction with a preferred embodiment, further illustrates each step of the technical solution provided in the above embodiment of the present invention.

[0047] The preferred embodiment of the multi-view three-dimensional astronaut pose estimation method for spacecraft includes the following steps: Step S1: Multi-view image acquisition and camera parameter calibration.

[0048] Multiple fixed-view image acquisition devices were deployed inside the spacecraft cabin to acquire synchronous multi-view human image sequences; the intrinsic and extrinsic parameters of each camera were calibrated to establish a unified world coordinate system, providing a geometric basis for subsequent multi-view information fusion.

[0049] This step is the basic input step, ensuring that multi-view data are consistent within the same spatial coordinate system.

[0050] Step S2: Two-dimensional human joint detection.

[0051] For each camera image, two-dimensional heat map information of each joint of the human body is extracted to obtain the confidence distribution of two-dimensional joints from multiple perspectives.

[0052] In some preferred embodiments, a two-dimensional human posture detection network can be used to extract two-dimensional heat map information of each joint of the human body, or other networks with corresponding algorithm functions can be used to extract two-dimensional heat map information of each joint of the human body.

[0053] Step S3: Multi-view feature backprojection and 3D feature voxel space construction.

[0054] Based on the camera calibration parameters, the two-dimensional joint heatmaps from each viewpoint are back-projected onto a unified three-dimensional space to construct a three-dimensional voxel feature body containing joint confidence information, thus forming an initial three-dimensional feature voxel space.

[0055] This step converts two-dimensional information into a three-dimensional voxel representation, which is the foundation for multi-view fusion.

[0056] Step S4: Human centroid encoding and dynamic centroid estimation based on human segment model. This step is the first core improvement step of the present invention.

[0057] S41, based on the topological relationship of human joints, divide the human body into multiple segments; S42 assigns a corresponding quality coefficient and proximal and distal parameters to each body segment; S43, based on the human segment model, uses the segment torque synthesis method to calculate the overall center of mass of the human body; S44 represents the human centroid as a probability distribution in three-dimensional space, generating the human centroid code; S45 uses the human body centroid encoding as a supervisory signal and regresses the dynamic centroid position of the human body through a three-dimensional convolutional network.

[0058] Existing technologies typically use fixed root nodes as the human body positioning reference, which is prone to significant offset errors in microgravity environments. This step introduces human centroid encoding constrained by a human segment model, utilizing the physical prior of human mass distribution to achieve more stable and reliable human body positioning. This step significantly reduces the uncertainty in human body positioning caused by attitude levitation in microgravity environments.

[0059] Step S5: Human body space occupancy area regression.

[0060] Based on three-dimensional voxel features, the spatial occupancy parameters of the human body are regressed on the projection plane, including planar position, spatial size and height information, to form an initial description of the human body's spatial occupancy.

[0061] This step is used to determine the approximate range of motion of the human body in three-dimensional space, providing constraints for subsequent fine joint estimation.

[0062] Step S6: Human body space optimization based on dynamic centroid and root node fusion. This step is the second core improvement step of the present invention.

[0063] The dynamic human centroid localization result obtained in step S4 is fused with the human body occupancy space result based on the root node in step S5. Through confidence comparison and non-maximum suppression strategy, the accurate human body occupancy is obtained by adaptively adjusting with changes in human posture.

[0064] Existing methods rely solely on a single root node or static bounding box for human spatial localization. This step achieves adaptive correction of human spatial occupancy through joint constraints of dynamic centroid and static root node; effectively reducing the joint search space and minimizing quantization errors caused by voxel partitioning.

[0065] Step S7: Construction of fine-grained voxel subspace and extraction of scale-aware features. This step is the third core improvement step of the present invention.

[0066] S71, based on the optimized human body occupancy, the fine-grained voxel subspace corresponding to the human body is obtained by cropping in the initial three-dimensional feature voxel space; S72, multi-scale feature extraction of voxel subspace; S73 introduces a scale-aware mechanism that integrates attention mechanisms to adaptively weight and fuse features at different scales.

[0067] Human body dimensions vary significantly among different individuals and in different postures. This step enhances the model's adaptability to scale changes through a scale perception and attention fusion mechanism, avoiding a decrease in detection accuracy due to differences in body size or posture.

[0068] Step S8: Multi-plane orthogonal projection and joint point heatmap prediction.

[0069] The fine-grained voxel subspace is projected onto multiple orthogonal planes, and two-dimensional heat maps of human joints are regressed on each orthogonal plane. The coordinates of the joints on each plane are calculated based on the centroid of the heat map.

[0070] This step, through multi-plane prediction, helps reduce quantization error in a single direction and improves the accuracy of joint positioning.

[0071] Step S9: Multi-plane adaptive weighted fusion to generate three-dimensional human pose. This step is the fourth core improvement step of the present invention.

[0072] Based on the confidence scores of the joint prediction results of each orthogonal plane, the weights of different planes in the 3D reconstruction process are learned, and the joint coordinates of multiple planes are adaptively weighted and fused to output the final 3D human pose result.

[0073] This step avoids the errors caused by simple averaging or fixed weight fusion in traditional methods; through adaptive weight learning, it improves robustness in cases of occlusion or missing viewpoints.

[0074] Steps S1 to S3 constitute M1 in this embodiment of the invention; steps S4 to S9 constitute M2 to M7 in this embodiment of the invention, respectively. Through the above steps, this invention achieves: dynamic human centroid localization based on human physical priors; adaptive spatial occupancy constraints to reduce voxel quantization errors; multi-scale attention fusion joint estimation to adapt to changes in human scale; and highly robust 3D human pose output even in complex, occluded, and microgravity environments.

[0075] Based on the technical solutions in the above preferred embodiments, the following provides a more detailed explanation of the connection relationships between the various method steps from a structural perspective, the overall working principle, the design principles of the core steps, and the correspondence between structure, principle, and effect.

[0076] From the perspective of structure and information flow, the method provided by the above embodiments of the present invention forms a hierarchical processing structure that is "from coarse to fine, from global to local, and from physical prior constraints to data-driven optimization". The steps are not executed in isolation, but there is a clear data dependency and logical progression relationship.

[0077] 1. Construction Relationship of Multi-View Input to Unified 3D Structure: Steps S1-S3 constitute the basic 3D structure construction layer of this method. Among them, S1 provides the geometric relationship between multi-view synchronized images and the camera; S2 generates the 2D joint probability distribution in each view; S3 uses calibration parameters to uniformly map the 2D probability information to the same 3D feature voxel space. Structurally, the 3D feature voxel space of S3 is the common geometric carrier for all subsequent steps. Subsequent centroid estimation, spatial clipping, and joint regression are all completed in this unified voxel coordinate system.

[0078] 2. Decoupling between the global human body localization structure and the local joint estimation structure: Steps S4-S6 form a global human body localization and occupancy modeling substructure independent of joint regression; steps S7-S9 are local high-precision joint estimation substructures executed under the constraints of the above structure. This structural "localization first, estimation later" decoupling is an important feature that distinguishes this invention from traditional end-to-end voxel regression methods.

[0079] 3. Hierarchical constraints between the human body's center of mass, occupancy space, and voxel subspace: Structurally, the core steps exhibit the following explicit hierarchical dependencies, such as... Figure 2 As shown, this hierarchical dependency ensures that the error does not spread across the entire space, but is instead constrained and converged step by step.

[0080] The following section focuses on explaining the physical and geometric principles underlying the four core improvement steps: S4, S6, S7, and S9.

[0081] (a) Step S4, the principle of human centroid coding based on human segment model.

[0082] The position of the human body's center of mass in three-dimensional space is essentially determined by the mass distribution and spatial position of its various segments. The human body segment model describes the human body through the following structural constraints: the human body is decomposed into several rigid segments; each segment has a relatively stable mass ratio; the overall center of mass is the weighted sum of the centers of mass of each segment. This relationship originates from classical rigid body mechanics and is independent of the specific posture, orientation, and direction of gravity of the human body.

[0083] Traditional 3D pose estimation methods often use the "pelvic point / root node" as a spatial reference. This implicitly assumes a fixed direction of gravity and a predominantly standing posture. However, in microgravity environments, this assumption does not hold, and the root node becomes unstable. The advantages of the human body's center of mass are: it does not rely on the direction of gravity or contact constraints; it possesses natural invariance to posture rotation and tumbling; and it is determined by the overall structure, making it insensitive to local joint errors, which are less likely to cause drastic shifts.

[0084] Therefore, by using the centroid derived from the human segment model as the three-dimensional positioning anchor point, i.e. the global spatial anchor point, the problem of unstable positioning reference under microgravity environment is eliminated from the perspective of physical structure, thus improving the stability and consistency of spatial positioning.

[0085] (ii) Step S6, the principle of human body space optimization by merging dynamic centroid and root node.

[0086] The geometric principles underlying this step are as follows: the root node is located from the skeletal topology, reflecting the geometric center of the human joints; the centroid is located from the mass distribution, reflecting the physical center of the human body; ideally, both should be located within the same human structure, but their deviation directions differ under noise or occlusion conditions.

[0087] From a structural perspective, root node errors mostly originate from local joint detection failures, while centroid errors are primarily caused by incomplete overall posture or missing body segments. By performing confidence-weighted fusion of these two methods, the resulting human body positioning space more closely approximates the overall geometric envelope of the real human body. This allows for structural compensation by the other method when one is unreliable, resulting in a human body positioning space that better conforms to the overall geometric boundaries of the real human body. By performing confidence-weighted fusion of dynamic centroid positioning results and root node positioning results, spatial offsets caused by the failure of a single positioning reference can be effectively suppressed, thereby obtaining a more stable and reasonable three-dimensional human body positioning range.

[0088] This step, starting from geometric complementarity, constructs a more robust spatial constraint than a single positioning reference, reducing the uncertainty in the human body space trimming stage.

[0089] (III) Step S7, the principle of fine-grained volumetric subspace and scale perception mechanism.

[0090] The spatial sampling principle underlying this step is as follows: Voxelization is essentially a spatial discrete sampling process, and its error magnitude is closely related to the voxel size and the proportion of effective information in the total voxel count. In the full-space voxel count, the human body occupies only a very small proportion, and irrelevant voxels introduce a large amount of noise and quantization error. This leads to the dilution of effective information, amplification of quantization error, and noise interference in the learning process. This step significantly improves the "effective voxel density" by using a cropped human body subspace. By first determining the precise occupancy space of the human body and then constructing a fine-grained voxel subspace, the effective voxel density is significantly improved, reducing the interference of irrelevant background voxels on feature learning and improving spatial representation accuracy at the same voxel resolution. Simultaneously, the human body exhibits significant scale variations under different individuals and postures. This step introduces a scale-aware mechanism to structurally adaptively adjust the receptive field for different human body shapes and extended posture layers, avoiding "oversampling" or "undersampling" at a fixed scale.

[0091] This step, based on the discrete sampling structure and scale adaptation principle, reduces the sources of quantization error and the impact of voxel quantization error on joint estimation accuracy at the voxel sampling structure level.

[0092] (iv) Step S9, the principle of multi-plane orthogonal projection and adaptive weighted fusion.

[0093] The geometric complementarity principle underlying this step is the spatial projection complementarity principle: When 3D spatial information is projected onto a 2D plane, information compression is inevitable. Different orthogonal planes retain different types of information in joint localization. Some joints are clearly visible on one plane, but their information may be degraded on another plane due to occlusion or overlap, thus exhibiting different advantages. Specifically, the XY plane is sensitive to horizontally unfolded joints, while the XZ / YZ plane is more sensitive to height and front / back structures. A single plane inevitably lacks information in certain directions, failing to fully and stably represent the 3D joint position. Occlusion is usually directional, and the visibility of different joints varies on different planes. This step predicts joints on multiple orthogonal planes separately and introduces an adaptive weighting mechanism to learn plane weights. It dynamically allocates contributions based on the confidence level of each plane's prediction results, automatically suppressing planes significantly affected by occlusion or projection degradation, strengthening the dominant role of the most reliable information direction in 3D reconstruction, and thus dynamically selecting the most reliable viewpoint.

[0094] This step is based on the principle of spatial information complementarity and statistical weighted fusion, which reduces occlusion and directionality errors at the spatial projection structure level, effectively reducing the impact of occlusion and projection directionality on 3D pose estimation.

[0095] This invention constructs a stable, interpretable, and error-controllable three-dimensional human posture estimation process at the methodological structure level by using physical structure priors (human segments → centroid), geometric structure constraints (occupying space → voxel clipping), and complementary multi-scale and multi-planar structures. This fundamentally improves the applicability and robustness in microgravity, occlusion, and complex cabin environments.

[0096] This invention uses the physical structure of the human body as a global stability constraint, restricts the search space with geometric envelope, controls the range of error propagation, and reduces the uncertainty of discretization and occlusion by multi-scale and multi-plane complementarity. It can ensure the stability, accuracy and interpretability of the three-dimensional human pose estimation results from the principle level in complex environments such as microgravity, frequent occlusion and space constraints.

[0097] This invention constructs a three-dimensional human posture estimation scheme suitable for complex microgravity environments by introducing prior knowledge of human physical structure, a hierarchical spatial constraint mechanism, and a multi-plane information complementarity structure into the methodological structure. Compared with existing technologies, this scheme has significant advantages in terms of structural rationality, estimation stability, environmental adaptability, and engineering deployability.

[0098] This invention employs a hierarchical workflow of "global localization + local fine estimation" to avoid global error propagation. This workflow is as follows: Figure 3As shown. Existing methods mostly employ end-to-end regression in a single voxel space. Once the overall human body positioning deviates, the error propagates throughout the entire space. This invention structurally divides the method into two parts: a global human body positioning structure (centroid + occupancy space) and a local joint fine estimation structure. This hierarchical structure constrains the global error before it enters joint estimation, and joint regression is performed only within the "credible space," significantly reducing the system's sensitivity to single-step errors.

[0099] This invention introduces a workflow constraint of "human segment model → centroid," resulting in a more stable positioning reference. This workflow is as follows: Figure 4 As shown. This invention no longer relies on a single joint as a spatial anchor point, but instead constructs the human body's center of mass through a physical structural link of human segments-mass-spatial position. The advantage of this structural approach is that the positioning reference comes from the overall structure rather than local features, making it insensitive to the loss or occlusion of local joints, and particularly suitable for scenarios without a fixed attitude reference in microgravity environments.

[0100] This invention improves the reliability of the human body spatial envelope through a workflow diagram that integrates a dynamic centroid and a root node. The workflow is as follows: Figure 5 As shown in the diagram, this workflow uses the root node to reflect the topological center of the skeleton and the centroid to reflect the physical center of the human body. The errors of these two nodes originate from different sources and have different directions. Through structural fusion, the other node can compensate for any unreliability of one, preventing overall displacement of the human body's spatial footprint and improving the accuracy of subsequent voxel clipping.

[0101] This invention employs a refined spatial working architecture of "voxel subspace pruning + scale awareness". This working architecture is as follows: Figure 6 As shown. Figure 6 In this model, if two people are detected in the global voxel space, the human sub-region is detected using two methods: root node detection and dynamic centroid detection. The smaller sub-region detected by either method is the final fine-grained voxel space. The advantages of this structure are that the human body occupies only a small proportion of the total space, and through structural pruning, computation is concentrated in the effective area of ​​the human body. Simultaneously, the introduction of a scale-aware structure automatically adapts to different body shapes and postures, avoiding the accuracy loss caused by a fixed voxel scale, thus achieving higher information density, lower voxel quantization error, and better computational efficiency.

[0102] This invention employs a multi-plane orthogonal projection structure, which possesses a naturally complementary information-rich working architecture. This working architecture is as follows: Figure 7 As shown in the diagram. The advantage of this structure is that it preserves geometric information from different directions on the same plane, and occlusion is often directional. This structure allows the entire method to operate independently of a single projection direction, dynamically selecting the most reliable information source and significantly improving stability in occluded scenarios.

[0103] Through the above-described structural design, the present invention exhibits the following advantages at the functional level: I. Advantages at the structural level 1. Employing a hierarchical processing structure reduces the risk of global error propagation. This invention structurally divides the overall method into a global human body localization layer and a local joint fine estimation layer. First, it determines the reliable range of the human body in three-dimensional space using the human body's centroid and its occupied space, and then performs joint regression within this range. This structure avoids the problem of error propagation throughout the entire space caused by overall localization deviation in traditional end-to-end voxel regression methods.

[0104] 2. Introduce human segment models to construct stable human spatial reference structures. By segmenting the human body into multiple physically meaningful segments using a segmental model and constructing the body's center of mass based on the mass distribution of these segments, the overall human body positioning no longer relies on a single joint node. This structural improvement ensures that the positioning benchmark originates from the overall physical structure of the human body, making it insensitive to local joint occlusion or detection errors, and significantly improving positioning stability.

[0105] 3. The human body spatial structure fused with dual positioning references is more reliable. This invention constructs a human body spatial structure with dual-reference constraints by integrating the human body's center of mass positioning results with the skeleton's root node positioning results. This structure fully utilizes the complementary characteristics of the physical center and the topological center, maintaining the rationality of the human body's spatial envelope even when either positioning result is unreliable, thus improving the overall structural robustness.

[0106] 4. Voxel subspace clipping structure improves spatial representation accuracy By cropping the global voxel space to obtain a fine-grained voxel subspace containing only the human body region, the proportion of effective information in the voxel space is significantly increased. This structure reduces the participation of irrelevant background voxels in the calculation, thereby reducing voxel discretization error at the structural level and improving the accuracy of subsequent joint estimation.

[0107] 5. Multi-plane orthogonal projection structures possess the advantage of complementary information. This invention employs a multi-orthogonal plane projection structure, where different planes retain spatial information of the human body in different directions, and these information are then fused using adaptive weights. This structure avoids the problem of missing information in a single projection direction, enabling the system to automatically select the plane with the most reliable information for 3D reconstruction.

[0108] II. Advantages at the Functional Level 1. Significantly improves the stability of 3D human pose estimation in microgravity environments. This invention does not rely on gravity direction or fixed attitude assumptions, and can adapt to the free floating, rolling and complex movement states of astronauts in microgravity environment, and achieve stable three-dimensional human posture output.

[0109] 2. It has a stronger ability to adapt to occlusion and loss of view. By employing multi-view voxel fusion, multi-plane joint regression, and adaptive weighting mechanisms, this invention can maintain high attitude estimation reliability even when some viewpoints are occluded or information is missing.

[0110] 3. Improve the accuracy of 3D positioning and joint estimation By constraining the human body's spatial footprint and constructing fine-grained voxel subspaces, this invention effectively reduces the joint search range, lowers voxel quantization errors, and improves the accuracy of three-dimensional joint positioning.

[0111] 4. Non-contact sensing methods are suitable for long-term on-orbit applications. This invention achieves human posture estimation based on visual perception, without requiring astronauts to wear any sensing devices, and does not affect their normal operations. It is suitable for long-term deployment in the spacecraft cabin environment.

[0112] 5. The method has a clear structure and good engineering scalability. Each processing step in this invention has an independent structure and clear logic. It can flexibly replace or expand the two-dimensional detection network, voxel construction method and fusion strategy, and is easy to integrate with the in-cabin intelligent system and on-orbit service system.

[0113] Based on the same inventive concept, one embodiment of the present invention also provides a multi-view three-dimensional astronaut attitude estimation system for spacecraft.

[0114] Specifically, such as Figure 8 As shown, the multi-view three-dimensional astronaut attitude estimation system for spacecraft provided in this embodiment may include: Image acquisition module, which is used to acquire synchronized multi-view human image sequences; The 3D voxel feature body construction module extracts two-dimensional heat maps of each joint point of the human body based on each human body image and converts them into three-dimensional voxel feature bodies containing joint point confidence information to form an initial three-dimensional feature voxel space. The Human Centroid Coding and Dynamic Centroid Estimation Module is used to construct a human segment model, perform human centroid coding and dynamic centroid estimation based on the human segment model, and obtain the dynamic human centroid localization result. The human body space description module is based on three-dimensional voxel features. It regresses the spatial space occupancy parameters of the human body on the projection plane to form an initial human body space description. The human body space optimization module is used to fuse the dynamic human body centroid positioning results with the initial human body space description to obtain an accurate human body space that adaptively adjusts with changes in human posture, resulting in an optimized human body space. A fine-grained voxel subspace construction module is used to construct a fine-grained voxel subspace and extract scale-aware features based on the optimized human body occupancy. The joint prediction module projects a fine-grained voxel subspace onto multiple orthogonal planes, regresses a two-dimensional heat map of human joints on each orthogonal plane, and calculates the coordinates of the joints on each plane based on the centroid of the heat map, thus obtaining the joint prediction results for each orthogonal plane. The 3D human pose output module is used to learn the weights of different planes in the 3D reconstruction process based on the confidence of the prediction results of the joints of each orthogonal plane, and to adaptively weight and fuse the joint coordinates of multiple planes to output the final 3D human pose result.

[0115] In some preferred embodiments, the image acquisition module may further include: multiple image acquisition devices with fixed viewing angles, which are deployed inside the spacecraft cabin.

[0116] It should be noted that the steps in the method provided by the present invention can be implemented using corresponding modules, devices, units, etc. in the system. Those skilled in the art can refer to the technical solution of the method to realize the composition of the system. That is, the embodiments in the method can be understood as preferred examples for building the system, and will not be elaborated here.

[0117] An embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can be used to perform any of the methods described in the above embodiments of the present invention.

[0118] Optionally, the memory is used to store programs; the memory may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs and functional modules that implement the above methods), computer instructions, etc., and the aforementioned computer programs and computer instructions can be partitioned and stored in one or more memories. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by the processor.

[0119] A processor is used to execute computer programs stored in memory to implement the various steps of the methods or various modules of the systems involved in the above embodiments. For details, please refer to the relevant descriptions in the preceding method and system embodiments.

[0120] The processor and memory can be separate structures or integrated structures. When the processor and memory are separate structures, they can be coupled together via a bus.

[0121] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to perform the method of any of the above embodiments of the present invention.

[0122] Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a user device. Of course, the processor and storage medium can also exist as discrete components in a communication device.

[0123] The multi-view three-dimensional astronaut posture estimation method and system for spacecraft provided in the above embodiments of the present invention adopt astronaut state perception and space occupancy measurement technology based on multi-view computer vision and three-dimensional human posture estimation. It can be widely applied to astronaut posture perception and intelligent management in manned spacecraft, space stations and their on-orbit service systems, and is applicable to the following specific application scenarios and fields: 1. Intelligent management of manned spacecraft and space station cabins. This invention can be deployed inside enclosed, confined, and microgravity cabins such as space stations, manned spacecraft, and cargo spacecraft (such as the "Qingzhou" cargo spacecraft). It uses multi-view visual sensors to perform real-time, continuous, and non-contact perception of the astronauts' three-dimensional human posture and spatial positioning, and is used for monitoring personnel activities, recording mission processes, and assessing safety status.

[0124] 2. Human-computer interaction and human-computer collaborative operation for astronauts. The high-precision three-dimensional human posture information obtained by this invention can be used as the core input signal of the human-computer interaction system, supporting natural interaction between astronauts and the in-cabin computer system, intelligent operating terminal and service robot, and improving the efficiency and reliability of astronauts performing complex operations in microgravity environment.

[0125] 3. Space Station Intelligent Robots and On-Orbit Services. This invention can provide astronauts with real-time attitude and spatial position priors for in-cabin service robots, collaborative robotic arms, and on-orbit maintenance systems, avoiding the risk of human-machine collisions, supporting safe human-machine collaborative operation and autonomous decision-making, and is applicable to the construction of future intelligent space stations and on-orbit service systems.

[0126] 4. Astronaut health monitoring and training assessment. This invention can be used for astronaut body posture monitoring, motion analysis, and behavior assessment during their time in orbit, and can transmit relevant data back to Earth for astronaut ground training simulation, microgravity posture adaptability analysis, and long-term health status assessment.

[0127] 5. 3D Human Body Perception and Intelligent Monitoring in Special Environments. Besides aerospace applications, the technical solution of this invention can also be extended to other complex and restricted environments, such as submarines, deep-sea compartments, and nuclear facility sections, where non-contact, high-precision 3D human body posture perception is required. It has strong engineering versatility and promotional value.

[0128] The application effects of the technical solutions provided in the above embodiments of the present invention will be further explained in detail below with reference to specific application examples.

[0129] Specific application example 1 This specific application example illustrates a scenario of daily work monitoring within the space station module. Multiple visual acquisition devices, fixedly installed on the module walls and roof, simultaneously image the astronauts from multiple perspectives. In the microgravity environment, the astronauts are in a free-floating state, and their body postures frequently rotate and roll. This example first constructs a unified three-dimensional feature voxel space based on the multi-view images. It then estimates the overall center of mass of the human body using a segmental model and uses this center of mass as a global spatial positioning reference. Combined with skeletal root node information, it determines the spatial occupancy area of ​​the human body within the module. Based on this, the system constructs a fine-grained voxel subspace only within the effective occupancy space of the human body and performs multi-planar joint regression. Even if the astronaut's arms or torso are obscured by equipment within the module, it can still stably output complete and continuous three-dimensional human posture results, thereby achieving accurate recording and analysis of the astronauts' daily work movements.

[0130] Specific application example 2 This specific application example illustrates a scenario of astronauts collaborating with in-cabin service robots. When astronauts perform equipment maintenance or material handling tasks inside the cabin, their body posture and movement trajectory exhibit significant directional uncertainty and spatial intersection characteristics. This example uses the fusion of the human body's center of mass and root node to obtain a reliable human spatial envelope, and updates the human body's occupied space in real time. This information is provided to the in-cabin service robot as a reference for obstacle avoidance and collaboration. In this state, the invention can continuously output stable three-dimensional human posture information, enabling the service robot to accurately determine the astronaut's spatial position and movement trend even when the astronaut is rapidly turning, drifting, or extending their limbs, thereby avoiding the risk of human-robot collisions and improving the safety of collaborative operations.

[0131] Specific application example 3 This specific application example illustrates an astronaut attitude perception scenario under complex occlusion conditions. During in-cabin operations, parts of an astronaut's body may be obscured by in-cabin equipment, floating objects, or other astronauts, resulting in a significant loss of information about human joint points from a single viewpoint. This example utilizes multi-view voxel fusion and a multi-plane orthogonal projection structure to map effective information from different directions to the same voxel space, and uses an adaptive weighted fusion mechanism to suppress projection planes significantly affected by occlusion. In this state, even if one or more camera views are obstructed for extended periods, this invention can still recover the complete human posture by relying on other viewpoints and planar information, achieving continuous perception of human behavior under occlusion conditions.

[0132] Specific application example 4 This specific application example illustrates the scenario of astronaut in-orbit training and health monitoring. When astronauts perform stretching, twisting, and limb coordination training exercises inside the cabin, their body dimensions and postures change significantly. This example utilizes a human body space constraint and scale perception mechanism to automatically adjust the size of the voxel subspace and the feature receptive field, ensuring that the human body of different sizes and with varying ranges of motion can be accurately modeled at an appropriate spatial resolution. In this state, the invention can continuously output high-precision three-dimensional joint trajectories, providing a reliable data foundation for astronaut motion assessment, training effect analysis, and long-term health monitoring.

[0133] Specific application example 5 This specific application example illustrates the scenario of identifying and monitoring emergency situations within the cabin. When an astronaut experiences abnormal posture changes inside the cabin, such as unplanned drift, unbalanced rotation, or abnormal stagnation, this example performs real-time analysis of the human body's spatial position and joint motion status based on a continuous three-dimensional posture sequence. Because this invention uses the human body's center of mass as a stable spatial reference, it can maintain overall positioning continuity even in cases of rapid changes in astronaut posture or instability in local joint detection, thus providing reliable input for abnormal situation identification and safety alarms.

[0134] The multi-view three-dimensional astronaut attitude estimation method and system for spacecraft provided in the above embodiments of the present invention solve the following technical problems and achieve outstanding technical effects: 1. Solving the problem of unstable human positioning reference in microgravity environments. Traditional 3D attitude estimation methods generally use root nodes or fixed joints as human positioning references, which lead to significant positioning errors in microgravity environments due to large changes in neutral body position and attitude floating. This invention introduces a human centroid encoding method based on a human segment parameter model. By utilizing the physical prior of human mass distribution, it achieves accurate estimation of the dynamic centroid of the human body, thereby improving the stability and accuracy of human spatial positioning.

[0135] 2. Addressing the issue of large spatial quantization errors in voxel-based 3D pose estimation. Existing voxel-based 3D pose estimation methods inevitably introduce quantization errors during voxel space partitioning, especially when the human body's occupied area is too large or pose changes drastically. This error is amplified and propagates to the joint regression stage. This invention first regresses the human body's centroid and precise human body occupied space, and then performs joint estimation in a local fine-grained voxel space, effectively reducing the search space and significantly decreasing the error accumulation caused by voxel discretization.

[0136] 3. Solving the problem of unstable attitude inference caused by multi-view occlusion and incomplete viewpoints. For complex working conditions such as confined space inside spacecraft cabins, floating objects, and easily obstructed camera views, this invention adopts a multi-view voxel fusion strategy. It maps information from multiple camera viewpoints to a unified three-dimensional feature voxel space and, through multi-plane orthogonal projection and adaptive weighted fusion mechanisms, achieves reliable inference of human joint information in partially occluded scenarios, significantly improving the robustness of attitude estimation.

[0137] 4. Addressing the issue of decreased model accuracy due to changes in human body scale. After introducing the human center of mass localization mechanism, variations in individual astronauts and their postures can cause differences in human body scale, affecting joint detection accuracy. To address this, this invention designs a scale perception module that integrates an attention mechanism. Through adaptive fusion of multi-scale features, it effectively improves the model's adaptability to changes in human body scale, ensuring stable detection performance under different body shapes and postures.

[0138] 5. Solving the problem of balancing accuracy and computational efficiency in real-time in-cabin applications. This invention adopts a modular network structure design to decouple human positioning and attitude estimation tasks, and combines a hierarchical design with voxel spatial resolution to reduce overall computational complexity while ensuring the accuracy of three-dimensional attitude estimation, thus meeting the dual requirements of real-time performance and reliability in the resource-constrained environment of a spacecraft cabin.

[0139] Any matters not covered in the above embodiments of the present invention are well-known in the art.

[0140] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A multi-view three-dimensional astronaut pose estimation method for a spacecraft, the method comprising: include: Acquire a synchronized multi-view human image sequence, extract two-dimensional heat maps of each joint point of the human body based on each human image, and convert them into three-dimensional voxel feature volumes containing joint point confidence information to form an initial three-dimensional feature voxel space. A human segment model is constructed, and human centroid encoding and dynamic centroid estimation are performed based on the human segment model to obtain the dynamic human centroid localization result. Based on the three-dimensional voxel feature, the spatial occupancy parameters of the human body are regressed on the projection plane to form an initial description of the human body's spatial occupancy. By fusing the dynamic human centroid positioning result with the initial human body space description, an accurate human body space that adaptively adjusts with changes in human posture is obtained, resulting in an optimized human body space. Based on the optimized human body occupancy, a fine-grained voxel subspace is constructed and human joint scale perception features are extracted. The fine-grained voxel subspace and human joint scale perception features are projected onto multiple orthogonal planes. Two-dimensional heat maps of human joints are regressed on each orthogonal plane, and the coordinates of the joints on each plane are calculated based on the centroid of the heat map to obtain the joint prediction results for each orthogonal plane. Based on the confidence level of the joint prediction results of each orthogonal plane, the weights of different planes in the 3D reconstruction process are learned, and the joint coordinates of multiple planes are adaptively weighted and fused to output the final 3D human pose result.

2. The multi-view three-dimensional astronaut pose estimation method for a spacecraft of claim 1, wherein, The process of acquiring a synchronized multi-view human image sequence involves extracting two-dimensional heatmaps of each joint point from each human image and converting them into a three-dimensional voxel feature volume containing joint point confidence information, forming an initial three-dimensional feature voxel space, including: By using multiple fixed-view image acquisition devices, a synchronous multi-view human image sequence is obtained, and the intrinsic and extrinsic parameters of each image acquisition device are calibrated to establish a unified world coordinate system. For the images acquired by each image acquisition device, two-dimensional heat map information of each joint point of the human body is extracted to obtain the confidence distribution of two-dimensional joint points from multiple perspectives. Based on the calibration parameters, the two-dimensional heatmaps of each joint point under each viewpoint are back-projected to a unified three-dimensional space to construct a three-dimensional voxel feature body containing joint point confidence information, thus forming an initial three-dimensional feature voxel space.

3. The multi-view three-dimensional astronaut pose estimation method for a spacecraft of claim 1, wherein, The construction of the human segment model, and the subsequent human centroid encoding and dynamic centroid estimation based on the human segment model to obtain the dynamic human centroid localization result, include: Based on the topological relationships of human joints, the human body is divided into multiple segments; Assign a corresponding mass coefficient and proximal and distal parameters to each body segment to construct a human segment model; Based on the aforementioned human segment model, the overall center of mass of the human body is calculated using the segment torque synthesis method. The overall centroid of the human body is represented as a probability distribution in three-dimensional space, and a human centroid code is generated. Using the human centroid encoding as a supervisory signal, the dynamic centroid position of the human body is regressed through a three-dimensional convolutional network, the dynamic centroid is estimated, and the dynamic human centroid localization result is obtained.

4. The multi-view three-dimensional astronaut pose estimation method for a spacecraft of claim 1, wherein, The process of regressing the spatial occupancy parameters of the human body on the projection plane based on the three-dimensional voxel feature volume to form an initial human body spatial description includes: Based on the three-dimensional voxel feature, the spatial occupancy parameters of the human body are regressed on the projection plane, including: planar position, spatial size and height information, to form an initial human body occupancy space description based on the root node.

5. The multi-view three-dimensional astronaut attitude estimation method for spacecraft according to claim 1, characterized in that, The step of fusing the dynamic human centroid positioning result with the initial human body space description to obtain an accurate human body position that adaptively adjusts with changes in human posture includes: The dynamic human centroid localization result is fused with the initial human occupancy space result based on the root node. Through a non-maximum suppression strategy, an accurate human occupancy that adaptively adjusts with changes in human posture is obtained.

6. The multi-view three-dimensional astronaut attitude estimation method for spacecraft according to claim 1, characterized in that, The step of constructing a fine-grained voxel subspace and extracting human joint scale-sensing features based on the optimized human body occupancy includes: Based on the optimized human body occupancy, the fine-grained voxel subspace corresponding to the human body is obtained by cropping in the initial three-dimensional feature voxel space. Human joint feature extraction is performed on the fine-grained voxel subspace; A scale perception mechanism that integrates spatial attention and channel attention is introduced to adaptively weight and fuse human joint features at different scales to obtain human joint scale perception features.

7. A multi-view three-dimensional astronaut attitude estimation system for spacecraft, characterized in that, include: Image acquisition module, which is used to acquire synchronized multi-view human image sequences; The 3D voxel feature body construction module extracts two-dimensional heat maps of each joint point of the human body based on each human body image and converts them into three-dimensional voxel feature bodies containing joint point confidence information to form an initial three-dimensional feature voxel space. The module for human centroid coding and dynamic centroid estimation is used to construct a human segment model, perform human centroid coding and dynamic centroid estimation based on the human segment model, and obtain the dynamic human centroid localization result. The human body space description module, based on the three-dimensional voxel feature, regresses the spatial occupancy parameters of the human body on the projection plane to form an initial human body space description. The human body space optimization module is used to fuse the dynamic human body centroid positioning result with the initial human body space description to obtain an accurate human body space that is adaptively adjusted with changes in human body posture, thus obtaining an optimized human body space. A fine-grained voxel subspace construction module is used to construct a fine-grained voxel subspace and extract human joint scale perception features based on the optimized human body occupancy. The joint prediction module projects the fine-grained voxel subspace and human joint scale perception features onto multiple orthogonal planes, regresses two-dimensional heat maps of human joints on each orthogonal plane, and calculates the coordinates of the joints on each plane based on the centroid of the heat map to obtain the joint prediction results for each orthogonal plane. The 3D human pose output module is used to learn the weights of different planes in the 3D reconstruction process based on the confidence of the prediction results of the joints of each orthogonal plane, and to adaptively weight and fuse the joint coordinates of multiple planes to output the final 3D human pose result.

8. The multi-view three-dimensional astronaut attitude estimation system for spacecraft according to claim 7, characterized in that, The image acquisition module includes: multiple image acquisition devices with fixed viewing angles, and the multiple image acquisition devices are deployed inside the spacecraft cabin.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to perform the method of any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program can be used to perform the method of any one of claims 1-6.