Somatosensory interaction sailboat manipulation simulation method and system based on three-dimensional human body posture estimation
Through the improved three-dimensional human posture estimation method, combined with lightweight Transformer and graph convolution model, the problems of high computing cost and insufficient real-time performance in the prior art are solved, and efficient somatosensory interactive sailboat manipulation simulation is achieved.
Patent Information
- Application Number
- CN202510419871.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The prior art has high computational cost in three-dimensional human posture estimation based on deep learning, difficult to meet the needs of real-time sailing maneuvering simulation, and has high deployment complexity, which limits the popularization of somatosensory interaction technology.
The lightweight Transformer model PoseformerV2 is adopted, combining the graph convolution model of space-time expansion and the selective state space model of hierarchical joint enhancement. Through the three-dimensional human posture estimation method of graph-guided state space enhancement, it realizes efficient extraction of local and global features, which is suitable for somatosensory interactive sailboat manipulation simulation.
On the premise of ensuring real-time performance, the accuracy of three-dimensional human posture estimation and the applicability of the model are improved, and are suitable for somatosensory interactive sailboat maneuver simulation tasks, reducing calculation costs and deployment complexity.
Smart Images

Figure CN120428850A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sailboat maneuvering simulation, and in particular relates to a somatosensory interactive sailboat maneuvering simulation method and system based on three-dimensional human body posture estimation. Background Art
[0002] Sailboat maneuvering simulations break the constraints of external factors like seasons, weather, and geographic environment, becoming an effective tool for promoting sailing. Current research can be categorized based on interaction methods: keyboard-and-mouse-based sailboat maneuvering simulations, physical-object-based sailboat maneuvering simulations, and VR headset-based sailboat maneuvering simulations. For example, prior art patents for sailboat simulations incorporate sailboat kinematic models to create interactive desktop-level keyboard-and-mouse-based sailboat simulation systems. However, these approaches are limited to keyboard and mouse interaction, making it difficult to build a memorized understanding of sailboat maneuvers through simulated training. Simulation systems based on physical simulators often offer superior maneuvering performance. For example, prior art proposes an external rudder device for sailboat driving simulation training, providing more accurate rudder angle information for sailboat maneuvering simulations. Prior art also proposes an OP-level sailboat maneuvering simulation platform and its control method, offering land-based simulations for beginners. However, these methods require specialized equipment and facilities, resulting in complex deployment. VR headset-based sailboat simulations maintain a high level of immersion while reducing space requirements. Prior art uses VR technology and provides users with an adjustable lighting system, but the high cost of VR headsets makes them difficult to promote. Somatosensory interaction technology, with its advantages of simple deployment, low cost, and natural interaction, can effectively overcome the shortcomings of the aforementioned methods and provide a new development direction for sailing maneuvering simulation. The core technology of somatosensory interaction is high-quality and rapid estimation of 3D pose. Early research in this area used wearable sensors. Existing technologies have proposed somatosensory interaction methods based on wearable smart devices, achieving high estimation accuracy but limiting the user's freedom of movement and reducing the user experience. Another area of research has used depth cameras such as Kinect to acquire 3D pose. Existing technologies have designed somatosensory interactive virtual rehabilitation training methods and systems based on Kinect to guide users through designated rehabilitation exercises. While these methods have demonstrated the effectiveness of Kinect devices in somatosensory interaction tasks, their deployment costs remain relatively high. Therefore, these solutions are not conducive to the promotion and popularization of sailing. With the development of deep learning-based 3D human pose estimation methods, the 3D coordinate positions and angles of human joints can be estimated from 2D images. This opens the door to lower-cost sailing simulations based on standard cameras, but it also presents certain challenges.
[0003] Current deep learning-based 3D human pose estimation tasks can be categorized into single-stage and two-stage approaches based on their approach. Single-stage approaches typically design an end-to-end network to directly extract features from images and predict 3D joint information. This approach offers higher information utilization. However, these network designs are often complex, particularly when predicting volumetric heatmaps or regressing 3D coordinates. This requires extremely high computational cost, making it difficult to incorporate temporal information and leading to significant temporal consistency issues. Two-stage approaches first utilize 2D pose estimation methods to extract 2D joint information from the input image or video; then, 3D coordinates are predicted based on these 2D joints. This approach improves the network's usability and flexibility through staged optimization. With the continuous advancement of the field of 2D human pose estimation, 2D-to-3D approaches have gradually become the mainstream of research. Among these, Transformer-based approaches have demonstrated outstanding performance in temporal pose estimation tasks. For example, PoseFormer is a spatiotemporal Transformer structure used to model the human body joint relationships in frames and the temporal correlation between frames; the MixSTE method alternately uses temporal and spatial Transformer modules to more finely encode the spatiotemporal dependencies of joint points; and MHFormer, MotionAGFormer, PoseFormerV2 and other methods are all variants or improvements of the transformer method, but these methods generally have large computational costs. Although subsequent methods have optimized reasoning by pruning tokens in the transformer, and adopted network architectures such as Uplift and Upsample to reduce computational costs, they are difficult to meet the needs of real-time tasks such as sailboat manipulation based on human posture estimation.
[0004] To further improve computational efficiency and modeling capabilities, research work, represented by Mamba, integrates time-varying parameters into a selective state-space model (SSM) framework based on input, allowing the model to selectively process information, thereby filtering out irrelevant interference signals while retaining important information for a long time, enhancing the model's dynamic adaptability and reasoning capabilities. Furthermore, the visual selective state-space model (VMamba) method can capture multi-dimensional and multi-directional global spatiotemporal dependencies, effectively improving Mamba's modeling capabilities for complex visual tasks. Hamba successfully achieved efficient three-dimensional hand pose reconstruction based on VMamba. The subsequent method, Posemamba, utilizes a joint-sequence-dependent reordering scanning strategy to improve the performance of three-dimensional human pose estimation. However, this method is a sequence-to-sequence estimation strategy that uses both past and future frame information, making it less advantageous for real-time data processing.
[0005] In addition, the existing technology proposes a posture estimation method for human-computer interaction, which uses a two-branch twin supervision network with an efficient channel attention mechanism to perform three-dimensional human posture estimation and verify its effectiveness on a self-made dataset; the existing technology proposes an intelligent interaction system for a smart exhibition hall, which includes a regression-based three-dimensional human posture estimation method; the existing technology proposes a method that integrates two Transformer attention mechanisms to extract multi-scale temporal features, thereby improving model accuracy, but the model architecture is designed as a sequence-to-sequence estimation method, which cannot meet the needs of practical applications such as real-time action recognition. Summary of the Invention
[0006] In order to overcome the problems existing in the related art, the embodiments disclosed in the present invention provide a somatosensory interactive sailboat steering simulation method and system based on three-dimensional human posture estimation, specifically relating to a somatosensory interactive sailboat steering simulator based on a three-dimensional human posture estimation method.
[0007] The technical solution is as follows: A somatosensory interactive sailboat maneuvering simulation method based on three-dimensional human posture estimation includes the following steps:
[0008] S1, uses a camera to capture real-time images and uses a 2D posture detector to obtain 2D human posture;
[0009] S2, inputs the 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate the 3D human pose;
[0010] S3 maps the obtained 3D human posture to the somatosensory control module in the sailing operation virtual scene to realize interactive operations based on somatosensory movements.
[0011] Further, in step S1, using a 2D posture detector to obtain a 2D human posture is to use RTMPose2D to extract the 2D human posture in the real shot picture.
[0012] In step S2, the graph-guided state-space enhanced 3D human pose estimation method includes:
[0013] S21, the PoseFormerV2 model is modified into a baseline model suitable for real-time estimation. The baseline model only takes as input a 2D pose sequence of t frames in the past of the target frame.
[0014] S22, using the SGJMamer layer to extract local spatial features from the target frame’s past f-frame poses, and using the TGJMamer layer to extract the global spatiotemporal features Y from the complete input low-frequency information and the spatial features extracted by SGJMamer;
[0015] Among them, the SGJMamer / TGJMamer layers are composed of the STE-GCN module, the JMamba module, and the Transformer module;
[0016] The Transformer modules in the SGJMamer / TGJMamer layers are spatial Transformer modules and temporal Transformer modules respectively; the JMamba module introduces a hierarchical joint-enhanced selective state-space model to improve the model's ability to extract global features; the STE-GCN module learns the graph structure relationship of joints through a spatiotemporal extended graph convolutional network to enhance the focus on local features.
[0017] S23, constructing a regression head to infer the target frame joint point information, constructing a loss function to train the target frame joint points, and then outputting the target frame 3D pose;
[0018] Furthermore, in step S22, STGJMamer is used to extract the spatiotemporal features of the human posture model to obtain the output Y of the fusion of spatiotemporal features, including:
[0019] S221, construct the graph state space Transformer layer SGJMamer to extract local spatial features;
[0020] S222, constructing a graph state temporal Transformer layer TGJMamer, fusing the low-frequency information of the complete input and the spatial features extracted in step S221, and performing spatiotemporal feature extraction to obtain the output Y of the fused spatiotemporal features;
[0021] Furthermore, in step S221, a graph state space Transformer layer SGJMamer is constructed to extract local spatial features, including:
[0022] The target frame's past f frame information [x t-f ,x t ] is input, and passes through the STE-GCN module and the JMamba module in sequence to obtain the feature X containing rich information SpaGJMamba ;The data will be fed into the spatial Transformer module;
[0023] The STE-GCN module predefines the spatial connections between joints within the same frame and the temporal dependencies of the same joint in different frames by extending the adjacency matrix in time and space. It then extracts the spatiotemporal dependencies between joints through graph convolution. Specifically, the STE-GCN module extracts the spatiotemporal dependencies between joints through the following process:
[0024]
[0025] Where, is the normalized spatiotemporal adjacency matrix, D graph is the degree matrix, and the diagonal elements represent the connection degree of each node; graph is the adjacency matrix, including spatial and temporal relationships; H' graph is the output feature matrix, the dimension is E×T×N×F', σ is the activation function, H original is the input feature matrix with dimensions of E×T×N×F; W is the learnable weight matrix of graph convolution, and Drop() is the random drop mechanism;
[0026] After graph convolution, H' contains local information. A hybrid residual design is introduced to fuse the original features with the graph convolution features using dynamic weighting. The formula is as follows:
[0027] X graph =α·H' graph +(1-α)·H original
[0028] Where, X graph is the feature result after the residual mechanism processing; α is a learnable weight parameter used to dynamically adjust the ratio of the two features; H' graph is the feature after graph convolution processing, H original are the original features without graph convolution processing.
[0029] The JMamba module uses a hierarchical joint enhancement method based on anatomical and kinematic chain reactions based on VMamba to extract features in the information transfer between related joints and complete pose estimation;
[0030] The hierarchical joint strengthening method based on anatomical and kinematic chain reactions includes:
[0031] In the joint enhancement sequence, the enhancement sequence is designed based on anatomical and kinematic chain reactions. It takes into account the characteristics of the linkage between the parent and child joints of the human body, as well as the chain reaction of the joints exhibited by the human body to maintain balance during movement. Through mutual enhancement between joints, information transmission is achieved, thereby improving the model's ability to model complex movements. A mirror flip operation is applied to the input data to generate symmetrical movements. During the pose estimation process, the posture of the original movement and the flipped movement are estimated simultaneously. The final output is obtained by taking the average of the results of the original movement and the flipped movement to achieve enhanced data symmetry.
[0032] Specifically, the specific processing method of the JMamba module to complete the posture estimation includes: combining the hierarchical joint enhancement method based on anatomy and kinematic chain reaction to construct a bidirectional to-be-scanned information X consisting of an enhanced joint point sequence and an original joint point sequence. scan, and calculate the learnable matrices required for selective state space scanning: state matrix A, control matrix B, output matrix C, instruction matrix D and state variable H;
[0033] X scan =f Scan (W in ·X in +b in )
[0034] Where, X scan is the bidirectional information to be scanned, f Scan is the initialization function of the bidirectional scan, which is used to project the original input into the state space; W in is the original joint point sequence, X in is the bias term, b in is the weight matrix;
[0035] Perform local bidirectional scanning and global bidirectional scanning, continuously update the state space, and capture long-term and short-term dependencies. The expression is:
[0036] H'=A·H+B·X scan
[0037] Where H' is the output feature matrix;
[0038] The dimensions are:
[0039] E×T×N×F'
[0040] Where E is the batch size, T is the number of input frames, N is the number of joints, and F' is the feature corresponding to each joint;
[0041] X SSM =C·H'+D·X scan
[0042] Where, X SSM is the intermediate feature generated by selective state space scanning;
[0043] After completing the bidirectional scanning, feature fusion is performed to obtain a complete feature expression rich in global and local information:
[0044] X out =f Merge (X SSM )
[0045] Where, X out is the final output result of this module, f Merge To integrate the bidirectional scanning results into the original dimension;
[0046] The input features of the spatial Transformer module are first normalized and processed with multi-head self-attention to capture the intrinsic connections between joints. Subsequently, the features are normalized and processed with the MLP layer, and nonlinear transformations are introduced through the activation function GELU to learn the complex nonlinear relationships between the input features.
[0047] Spatial feature extraction includes:
[0048] Q,K,V=linear(LayerNorm(X SpaGJMamba ))
[0049]
[0050] x spatial =x res +Drop(x mlp )
[0051] Where Q is the query vector, K is the key vector, V is the value vector, linear() is the fully connected layer, LayerNorm() is the normalization process, and X SpaGJMamba is the spatial feature after processing by the STE-GCN module and the JMamba module, x res is the intermediate processing result, x attn is the feature processed by the attention mechanism, softmax() is to normalize the vector into a probability distribution vector, K T is the transpose of K, d k is the scaling factor, is the intermediate processing, Drop() is the random discarding processing, x mlp is the feature processed by the mlp mechanism, W1 and W2 are both weight matrices of linear transformation, GELU() is the activation function, b1 and b2 are both bias terms, and x spatial It is the spatial dimension feature in the target frame estimation process.
[0052] Furthermore, in step S222, a graph state temporal Transformer layer TGJMamer is constructed to fuse the low-frequency information of the complete input and the spatial features extracted in step S221, and perform spatiotemporal feature extraction to obtain the output Y of the fused spatiotemporal features, including:
[0053] The graph state temporal transformer layer consists of the STE-GCN module, the JMamba module, and the temporal transformer module. The data input is the low-frequency part of the complete action sequence, which captures the changing characteristics of the action from the global information and obtains the overall evolution trend of the action. After being processed by the STE-GCN module and the JMamba module, the frequency domain feature X is obtained. TGJMamba , the frequency domain information will be combined with the spatial feature xspatial The combined spatiotemporal information is sent to the self-attention mechanism for processing; when the processed data passes through the MLP layer, the spatial feature data will be transformed into the frequency domain to convert local detail changes into frequency domain information, thereby effectively supplementing the low-frequency time information.
[0054] Furthermore, in step S23, the regression head is constructed to infer the target frame joint point information, and a loss function is constructed to train the target frame joint points. The joint point error is calculated by comparing the output of the regression head with the real 3D pose, including:
[0055] The spatiotemporal feature Y extracted from the input two-dimensional sequence posture information has the following feature dimensions: represents the set to which Y belongs, f is the number of frames of the input feature; the target frame is the last frame of the sequence, and the dimension is Construct a regression head to infer the target frame joint information; the regression head consists of a weighted average operation and a fully connected layer. The estimated posture information will be calculated under the weight of the learnable weight parameters to calculate the final estimate Y of the target frame. predicate ;
[0056] Y predicate =W head LayerNorm(mean(Y))+b head
[0057] Where W head is the trainable weight, mean(Y) is the mean, b head is the bias term, LayerNorm() is the layer normalization;
[0058] The output result will be the same as the true 3D pose Y target Calculate the joint point error MPJPE, and use MPJPE as the loss function to train the model;
[0059]
[0060] Where y k , are the real information and predicted information of the kth joint point respectively.
[0061] Another object of the present invention is to provide a somatosensory interactive sailboat maneuvering simulation system based on three-dimensional human posture estimation, comprising:
[0062] 2D human posture capture module, used to use the camera to capture real-scene pictures in real time and use the 2D posture detector to obtain the 2D human posture;
[0063] 3D human pose calculation module, used to input 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate 3D human pose;
[0064] The somatosensory interaction module is used to map the obtained 3D human body posture to the somatosensory control module in the virtual scene of sailing operation to realize interactive operation based on somatosensory action.
[0065] Combining all the above technical solutions, the beneficial effects of the present invention are as follows: the present invention is based on the lightweight Transformer model PoseformerV2, integrating the advantages of the spatiotemporal expansion graph convolution model in local feature extraction and the advantages of the hierarchical joint enhanced state space model in global feature extraction. While ensuring the real-time performance of the model, it demonstrates its advantage in accuracy in the test of public datasets, and is more suitable for the somatosensory interactive sailboat manipulation simulation task.
[0066] This paper proposes a 3D human pose estimation network (STGJMamer) that integrates a hierarchical joint-enhanced selective state-space module (JMamba) and a spatiotemporal graph convolutional network (STE-GCN). This paper also proposes a somatosensory interactive sailboat maneuvering simulation based on a 3D human pose estimation method. This paper proposes a hierarchical joint-enhanced selective state-space module (JMamba), which efficiently extracts global features of human pose through hierarchical joint enhancement, thereby improving the accuracy of the model. This paper proposes a spatiotemporal graph convolutional network (STE-GCN), which enhances the model's focus on local information through graph convolution, thereby improving the accuracy of model estimation. This paper deploys STGJMamer in a somatosensory recognition sailboat maneuvering simulation training system and verifies its effectiveness in practical applications. Furthermore, this method only takes past frames as input and can run in real time, thus possessing extremely high practical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure;
[0068] Figure 1 This is a flow chart of a somatosensory interactive sailboat maneuvering simulation method based on three-dimensional human posture estimation provided by an embodiment of the present invention;
[0069] Figure 2 This is a schematic diagram of a somatosensory interactive sailboat maneuvering simulation method based on three-dimensional human posture estimation provided by an embodiment of the present invention;
[0070] Figure 31 is a schematic diagram of the graph-guided state-space enhanced 3D human pose estimation method provided by an embodiment of the present invention, wherein (a) is a diagram of the STGJMamer network structure, (b) is a diagram of the SGJMamer / TGJMamer network architecture, (c) is a diagram of the spatial Transformer network architecture, and (d) is a diagram of the temporal Transformer network architecture;
[0071] Figure 4 This is a schematic diagram of a hierarchical joint enhancement method based on anatomy and kinematic chain reactions provided by an embodiment of the present invention;
[0072] Figure 5 is the spatiotemporal extended adjacency matrix in the STE-GCN provided by an embodiment of the present invention;
[0073] Figure 6 This is the qualitative effect of the posture estimation of STGJMamer provided by the embodiment of the present invention in different input video cases;
[0074] Figure 7 This is the somatosensory interactive action control design of the present invention. DETAILED DESCRIPTION
[0075] To make the above-mentioned objects, features, and advantages of the present invention more readily apparent, specific embodiments of the present invention are described in detail below with reference to the accompanying drawings. The following description sets forth numerous specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways than those described herein, and those skilled in the art may make similar modifications without departing from the scope of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0076] The innovation of this invention lies in: it proposes a graph-guided state-space enhanced 3D pose estimation network (STGJMamer). It uses the lightweight Transformer method PoseformerV2 as the baseline model. Under the premise of ensuring the real-time performance of the model, it combines the spatiotemporal expansion of graph convolution and the selective state space model with hierarchical joint enhancement to enhance the model's ability to model global and local features, thereby improving the accuracy of pose estimation for mobile users.
[0077] Example 1, as Figure 1 As shown, the method for simulating sailboat maneuvering based on somatosensory interaction and 3D human posture estimation provided by an embodiment of the present invention includes:
[0078] S1, uses a camera to capture real-time images and uses a 2D posture detector to obtain 2D human posture;
[0079] S2, inputs the 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate the 3D human pose;
[0080] S3 maps the obtained 3D human posture to the somatosensory control module in the sailing operation virtual scene to realize interactive operations based on somatosensory movements.
[0081] Exemplarily, in step S1, obtaining a 2D human body posture using a 2D posture detector is to extract a 2D human body posture from a real-shot image using RTMPose2D.
[0082] Exemplarily, in step S2, calculating the 3D human body posture includes:
[0083] The accuracy of 3D human pose estimation is reflected by using different core evaluation indicators on the MPI-INF-3DHP dataset;
[0084] Different core evaluation indicators include: the percentage of correct key points PCK under a threshold of 150mm, the area under the curve AUC, and the mean joint position error MPJPE.
[0085] For example, the present invention designs specific somatosensory interactive actions for various sailing operations. The system will continuously monitor the 3D human body posture mapped to the virtual scene and match it with the preset action template. If the match is successful, the sailing operation function in the scene will be triggered to complete the somatosensory interaction. Figure 2 Schematic diagram of the somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation.
[0086] Example 2, in step S2, a graph-guided state space enhanced 3D human posture estimation method, namely, STGJMamer network; Figure 3 As shown; the network includes:
[0087] Figure 3 Figure (a) shows the STGJMamer network structure: This model takes a sequence of t 2D poses as input. SGJMamer extracts local spatial features from the target frame's past f poses, while TGJMamer extracts global spatiotemporal features from the full input's low-frequency information and the SGJMamer's extracted spatial features. Finally, this feature is used to compute the target frame's 3D pose through a regression head. Figure 3 (b) SGJMamer / TGJMamer network architecture: consists of a STE-GCN module, a JMamba module and a Transformer module. Figure 3 (c) Spatial Transformer module in , and Figure 3 (d) Temporal Transformer module in Figure 5.
[0088] This paper takes the PoseFormerV2 model as its basis and transforms it into a model suitable for real-time estimation tasks as the baseline model of this invention. On this basis, the present invention introduces the hierarchical joint-enhanced selective state-space module JMamba, which improves the model's global feature extraction capabilities through a joint-level enhanced bidirectional scanning method. In addition, the present invention also introduces a spatiotemporal extended graph convolution module STE-GCN, which learns the graph structure relationship of joints through spatiotemporal extended graph convolution, enhancing the model's attention to local features, thereby achieving improved model accuracy.
[0089] Specifically, the 3D human body posture estimation method includes the following steps:
[0090] S21, the graph-guided state-space enhanced 3D human pose estimation method is based on the PoseFormerV2 model. However, the PoseFormerV2 model utilizes both past and future information of the target frame and is not suitable for real-time tasks. Therefore, the present invention first modifies the PoseFormerV2 model into a baseline model suitable for real-time estimation. The baseline model only takes the t-frame 2D pose sequence of the target frame as input;
[0091] S22, use the SGJMamer layer to extract local spatial features from the past f frame poses of the target frame, and use the TGJMamer layer to extract the global spatiotemporal features Y from the low-frequency information of the complete input and the spatial features extracted by SGJMamer.
[0092] Among them, the SGJMamer / TGJMamer layers are composed of the STE-GCN module, the JMamba module, and the Transformer module;
[0093] The Transformer modules in the SGJMamer / TGJMamer layers are spatial Transformer modules and temporal Transformer modules respectively; the JMamba module introduces a hierarchical joint-enhanced selective state-space model to improve the model's ability to extract global features; the STE-GCN module learns the graph structure relationship of joints through a spatiotemporal extended graph convolutional network to enhance the focus on local features.
[0094] S23, constructing a regression head to infer the target frame joint point information, constructing a loss function to train the target frame joint points, and then outputting the target frame 3D pose;
[0095] Exemplarily, in step S22, STGJMamer is used to extract the spatiotemporal features of the human posture model to obtain the output Y of the fusion of spatiotemporal features, including:
[0096] S221, construct the graph state space Transformer layer SGJMamer to extract local spatial features;
[0097] S222, constructing a graph state temporal Transformer layer TGJMamer, fusing the low-frequency information of the complete input and the spatial features extracted in step S221, and performing spatiotemporal feature extraction to obtain the output Y of the fused spatiotemporal features;
[0098] Exemplarily, step S221, constructing a graph state space Transformer layer SGJMamer, and performing local space feature extraction includes:
[0099] like Figure 3 As shown in Figure (b), the state space Transformer layer consists of a STE-GCN module, a JMamba module and a spatial Transformer module. The target frame’s past f frame information [x t-f ,x t ] is input, and passes through the STE-GCN module and the JMamba module in sequence to obtain the feature X containing rich information SpaGJMamba ;The data will be fed into the spatial Transformer module;
[0100] The STE-GCN module predefines the spatial connections between joints within the same frame and the temporal dependencies of the same joint in different frames through the spatiotemporal extended adjacency matrix, and then extracts the spatiotemporal dependencies between joints through graph convolution. Specifically, the STE-GCN module extracts the spatiotemporal dependencies between joints through the following process:
[0101]
[0102] Where, is the normalized spatiotemporal adjacency matrix, D graph is the degree matrix, and the diagonal elements represent the connection degree of each node; graph is the adjacency matrix, including spatial and temporal relationships; H' graph is the output feature matrix, the dimension is E×T×N×F', σ is the activation function, H original is the input feature matrix with dimensions E×T×N×F; W is the learnable weight matrix of graph convolution, and Drop() is the random drop mechanism;
[0103] In the spatial Transformer module, the input features are first normalized and processed with multi-head self-attention to capture the intrinsic connections between joints. Subsequently, the features are normalized and processed with the MLP layer, and nonlinear transformations are introduced through the activation function GELU to learn the complex nonlinear relationships between the input features.
[0104] Spatial feature extraction includes:
[0105] Q,K,V=linear(LayerNorm(X SpaGJMamba ))
[0106]
[0107] x spatial =x res +Drop(x mlp )
[0108] Where Q is the query vector, K is the key vector, V is the value vector, linear() is the fully connected layer, LayerNorm() is the normalization process, and X SpaGJMamba is the spatial feature after processing by the STE-GCN module and the JMamba module, x res is the intermediate processing result, x attn is the feature processed by the attention mechanism, softmax() is to normalize the vector into a probability distribution vector, K T is the transpose of K, d k is the scaling factor, is the intermediate processing, Drop() is the random discarding processing, x mlp is the feature processed by the MLP mechanism, W1 and W2 are both weight matrices of linear transformation, GELU () is the activation function processing, b1 and b2 are both bias terms, and x spatial It is the spatial dimension feature in the target frame estimation process.
[0109] In the experiment, since video stream data is input continuously and there is significant temporal correlation between frames, it is necessary to extend graph convolution to the temporal dimension to extract more local feature information. Existing spatiotemporal graph convolution models achieve hierarchical modeling of spatiotemporal dynamic relationships by stacking spatiotemporal convolution and graph convolution operations, but this incurs huge additional computational overhead.
[0110] In order to achieve efficient graph convolution operation, this paper proposes a spatiotemporal extended adjacency matrix, such as Figure 5This is the spatiotemporal extended adjacency matrix in STE-GCN. In the figure, Ji (i = 0, 1, ..., 16) represents the joint point number; tj (j = 0, 1, ...f) represents the frame number; in summary, Ji(tj) represents the i-th joint point in the j-th frame. Spatiotemporal extended graph convolution (STE-GCN) uses this adjacency matrix to predefine the spatial connections between joints within the same frame and the temporal dependencies of the same joint across different frames. Based on this structure, graph convolution can effectively extract the spatiotemporal dependencies between joints, allowing the model to focus on the intrinsic relationships between joints during feature extraction and capture richer local information. This approach essentially extracts temporal features from multiple frames, effectively avoiding redundant high-dimensional temporal modeling.
[0111] After graph convolution processing, H' will contain more local information. In order to keep the original information from being lost, a hybrid residual design is introduced to fuse the original features with the graph convolution features using dynamic weighting. The formula is as follows:
[0112] X graph =α·H′ graph +(1-α)·H original
[0113] Where, X graph is the feature result after the residual mechanism processing; α is a learnable weight parameter used to dynamically adjust the ratio of the two features; H' graph is the feature after graph convolution processing, H original are the original features without graph convolution processing.
[0114] Exemplarily, the JMamba module utilizes a hierarchical joint enhancement method based on anatomical and kinematic chain reactions based on VMamba to perform feature extraction in the information transfer between related joints and complete pose estimation;
[0115] The hierarchical joint strengthening method based on anatomical and kinematic chain reactions includes:
[0116] In the joint enhancement sequence, the enhancement sequence is designed based on anatomical and kinematic chain reactions. It takes into account the characteristics of the linkage between the parent and child joints of the human body, as well as the chain reaction of the joints exhibited by the human body to maintain balance during movement. Through mutual enhancement between joints, information transmission is achieved, thereby improving the model's ability to model complex movements. A mirror flip operation is applied to the input data to generate symmetrical movements. During the pose estimation process, the posture of the original movement and the flipped movement are estimated simultaneously. The final output is obtained by taking the average of the results of the original movement and the flipped movement to achieve enhanced data symmetry.
[0117] For example, VMamba effectively stores and propagates long-term historical information through implicit recursive calculation and bidirectional scanning mechanism, thereby achieving stable capture of long-term dependencies. The present invention integrates VMamba to quickly establish a selective state space layer and adopts bidirectional scanning to model joint sequences. Figure 4 As shown in the average error analysis of different joints in Figure (a) of the hierarchical joint augmentation method based on anatomical and kinematic chain reactions, this method is prone to large deviations in the terminal joints during pose estimation. Therefore, the present invention proposes a hierarchical joint augmentation method based on anatomical and kinematic chain reactions, which strengthens VMamba's information transfer between related joints, further enhances the model's feature extraction capabilities, and improves the accuracy of pose estimation.
[0118] In the design process of the joint enhancement sequence, the present invention takes two aspects into consideration. On the one hand, the present invention takes into account the anatomical sequential relationship, that is, physically connected joints generally exhibit interdependent movements during movement. Therefore, enhancing the parent-child node relationship can effectively amplify the motion characteristics and thus extract richer motion information. On the other hand, the present invention also takes into account the motion chain reaction in complex movements. When one part of the body moves, other parts will also move accordingly in order to maintain the overall coordination of the body. For example, lifting one leg will cause the opposite hip joint to move, and bending the waist will cause the knee joint to move.
[0119] Taking these two aspects into consideration, the final joint enhancement sequence is shown in Table 1. Furthermore, due to the symmetry of human joints, during training and testing, the present invention applies a mirror flip operation to the input data to generate symmetrical movements. During pose estimation, the present invention simultaneously estimates the pose of the original and flipped movements and averages the results to obtain the final output, thereby enhancing data symmetry.
[0120] Table 1 Joint enhancement sequence
[0121] 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 First 0 0 1 2 3 0 4 5 6 8 11 8 11 12 14 15 16 Second 0 0 2 2 3 5 4 7 6 8 11 8 11 12 14 15 16
[0122] like Figure 4 As shown in Figure (b) of the hierarchical joint enhancement method based on the chain reaction of anatomy and kinematics, the hierarchical joint enhancement method is demonstrated, taking the joints of the right leg as an example: During the scanning process, the model first performs a structural enhancement based on the parent-child relationship of the joints and the correlation characteristics of human motion. Joint point 1 is enhanced with joint point 0, joint point 2 is enhanced with node 1, and joint point 3 is enhanced with node 2. Subsequently, a second enhancement is performed on the terminal joint 3, using parent node 2 as a bridge for information transmission to link the information of grandparent node 1. Ultimately, the characteristics of the terminal joint 3 will contain all the information of nodes 1, 2, and 3, so as to fully utilize the information between each node.
[0123] The joint information contains rich local information after the layer-by-layer enhancement process. In order to maintain its complete expression of global information, the present invention performs bidirectional scanning according to the local enhanced information on the one hand, and performs global bidirectional scanning according to the incompletely enhanced information on the other hand. Finally, the two are fused to output richer feature information. The network structure of JMamba is as follows Figure 3 As shown in Figure (b).
[0124] Specifically, the specific processing method of the JMamba module to complete the posture estimation includes: combining the hierarchical joint enhancement method based on anatomy and kinematic chain reaction to construct a bidirectional to-be-scanned information X consisting of an enhanced joint point sequence and an original joint point sequence. scan , and calculate the learnable matrices required for selective state space scanning: state matrix A, control matrix B, output matrix C, instruction matrix D and state variable H;
[0125] X scan =f Scan (W in ·X in +b in )
[0126] Where, X scan is the bidirectional information to be scanned, f Scan is the initialization function of the bidirectional scan, which is used to project the original input into the state space; W in is the original joint point sequence, X in is the bias term, b in is the weight matrix;
[0127] Perform local bidirectional scanning and global bidirectional scanning, continuously update the state space, and capture long-term and short-term dependencies. The expression is:
[0128] H'=A·H+B·X scan
[0129] Where H' is the output feature matrix;
[0130] The dimensions are:
[0131] E×T×N×F'
[0132] Where E is the batch size, T is the number of input frames, N is the number of joints, and F' is the feature corresponding to each joint;
[0133] X SSM =C·H'+D·X scan
[0134] Where, X SSMis the intermediate feature generated by selective state space scanning;
[0135] After completing the bidirectional scanning, feature fusion is performed to obtain a complete feature expression rich in global and local information:
[0136] X out =f Merge (X SSM )
[0137] Where, X out is the final output result of this module, f Merge To integrate the bidirectional scanning results into the original dimension;
[0138] Exemplarily, in step S222, a graph state temporal Transformer layer TGJMamer is constructed to fuse the low-frequency information of the complete input and the spatial features extracted in step S221, and perform spatiotemporal feature extraction to obtain an output Y of the fused spatiotemporal features, including:
[0139] The graph state temporal transformer layer consists of the STE-GCN module, the JMamba module, and the temporal transformer module. The data input is the low-frequency part of the complete action sequence, which captures the changing characteristics of the action from the global information and obtains the overall evolution trend of the action. After being processed by the STE-GCN module and the JMamba module, the frequency domain feature X is obtained. TGJMamba , the frequency domain information will be combined with the spatial feature x spatial The combined spatiotemporal information is sent to the self-attention mechanism for processing; when the processed data passes through the MLP layer, the spatial feature data will be transformed into the frequency domain to convert local detail changes into frequency domain information, thereby effectively supplementing the low-frequency time information.
[0140] For example, after being processed by the STE-GCN module and the JMamba module similar to SGJMamer, the frequency domain feature X is obtained. TGJMamba , the frequency domain information will be combined with the spatial feature x spatial The combined spatiotemporal information is fed into the self-attention mechanism for processing. The difference is that when the processed data passes through the multi-layer perceptron, the spatial feature data will be transformed into the frequency domain through discrete cosine transform to convert local detail changes into frequency domain information, thereby effectively supplementing the low-frequency time information; the discrete cosine transform formula is:
[0141]
[0142] Where k is the frequency index, k = 0, 1, 2…N-1; i is the time index, i = 0, 1, 2…N-1; N is the length of the input signal;
[0143] The inverse discrete cosine transform formula is:
[0144]
[0145]
[0146] Exemplarily, in step S23, a regression head is constructed to infer the target frame joint point information, a loss function is constructed to train the target frame joint points, and the joint point error is calculated between the output of the regression head and the true 3D pose, including:
[0147] The spatiotemporal feature Y extracted from the input two-dimensional sequence posture information has the following feature dimensions: represents the set to which Y belongs, f is the number of frames of the input feature; the target frame is the last frame of the sequence, and the dimension is Construct a regression head to infer the target frame joint information; the regression head consists of a weighted average operation and a fully connected layer. The estimated posture information will be calculated under the weight of the learnable weight parameters to calculate the final estimate Y of the target frame. predicate ;
[0148] Y predicate =W head LayerNorm(mean(Y))+b head
[0149] Where W head is the trainable weight, mean(Y) is the mean, b head is the bias term, LayerNorm() is the layer normalization;
[0150] The output result will be the same as the true 3D pose Y target Calculate the joint point error MPJPE. The present invention follows the setting of the benchmark model and uses MPJPE as the loss function of the model to train the model.
[0151]
[0152] Where y k , are the real information and predicted information of the kth joint point respectively.
[0153] In embodiment 2, the present invention provides a somatosensory interactive sailboat maneuvering simulation system based on three-dimensional human posture estimation, the system comprising:
[0154] 2D human posture capture module, used to use the camera to capture real-scene pictures in real time and use the 2D posture detector to obtain the 2D human posture;
[0155] 3D human pose calculation module, used to input 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate 3D human pose;
[0156] The somatosensory interaction module is used to map the obtained 3D human body posture to the somatosensory control module in the virtual scene of sailing operation to realize interactive operation based on somatosensory action.
[0157] Application Example 1, implementation case on public dataset.
[0158] Step 1, data set. The MPI-INF-3DHP dataset is significantly different from other similar datasets with its diversified indoor and outdoor scenes, wide and dynamic action categories, strong freedom of movement and fast action rhythm, showing a high degree of similarity with the subject's action situation in the real-time recognition environment. Especially in the specific application scenario of sailboat maneuvering simulation, the random and non-periodic movement pattern of the subject puts higher requirements on the adaptability of the model. Therefore, the present invention believes that the performance of the model on the MPI-INF-3DHP dataset is closer to the actual needs of the sailboat maneuvering simulation task. In order to make a fair comparison with other methods, the data input during the experiment follows the unified standard of experimental methods in recent years and adopts real 2D posture input. The data input length is set to 81 frames, and the short sequence interception uses 9 frames.
[0159] Step 2: Evaluation Metrics. This paper focuses on three core evaluation metrics: PCK (Percentage of Correct Keypoints), AUC (Area Under the Curve), and MPJPE (Mean Per Joint Position Error) at a threshold of 150 mm. These metrics reflect the accuracy and stability of the model in human pose estimation tasks from different dimensions.
[0160] Step 3, experimental results.
[0161] Table 2 Experimental results compared with other methods on the MPI-INF-3DHP dataset
[0162]
[0163]
[0164] As shown in Table 2, the model proposed in this paper demonstrates excellent performance in all evaluation indicators. Specifically, its PCK value is as high as 97.2%, which means that within the specified threshold, the model can accurately predict the positions of key points of the human body in most cases, reflecting its high precision in key point detection. At the same time, the AUC value reaches 77.9%, indicating that the model can maintain high detection accuracy at different thresholds and has good robustness. In addition, the MPJPE value is only 29.5mm, further demonstrating the accuracy of the model in joint position estimation.
[0165] In contrast, some existing methods (such as VideoPose3D and MHFormer) perform well in PCK but poorly in AUC and MPJPE, indicating a lack of stability. While these methods can accurately localize keypoints under ideal conditions (such as unobstructed scenes), they perform poorly in complex situations such as dynamic occlusion and rapid movement, resulting in low overall stability and large error fluctuations. Other methods (such as GTA-Net and HDFormer) perform well in AUC but poorly in PCK and MPJPE. This suggests that while these methods can maintain stable output under complex poses and occlusions, they over-rely on global features, resulting in decreased local keypoint localization accuracy and increased error. Furthermore, some methods (such as GLA-GCN and STCFormer) perform well in MPJPE and have low model error, but often require high computational costs in practical applications.
[0166] The model of this invention outperforms most existing methods across all three evaluation metrics. Furthermore, the model's design relies entirely on past information, which enhances its practical application value in real-time pose estimation tasks. In summary, the model of this invention performs well on the MPI-INF-3DHP dataset, validating its ability to estimate human pose in complex dynamic scenes and providing strong technical support for its subsequent application in sailboat maneuvering simulation systems.
[0167] Step 3: Inference speed.
[0168] The model parameters of the model of the present invention are about 14.4M, the amount of calculation is about 0.177GMac, and the frame-by-frame inference speed is about 52FPS. In practical applications, the present invention adopts RTMPose as a 2D detector and is deployed on an NVIDIA RTX4090GPU. Experimental results show that the method of the present invention still maintains an inference speed of 38FPS during the complete estimation process, which can fully meet the needs of real-time tasks. In the above experiments, the present invention enhances the data by following the commonly used flipping method in experiments in recent years, that is, estimating the same data twice. In practical applications, according to the performance requirements of the device, the data flipping function can be turned off. At this time, the average estimation speed from the picture to the 3D human posture can reach 60FPS.
[0169] Step 5: Ablation experiment.
[0170] We conduct systematic ablation experiments on the proposed modules and demonstrate their effectiveness by adding components one by one to the baseline model (see Table 3).
[0171] Table 3 Experimental results comparing existing methods
[0172]
[0173] Step 6: Qualitative presentation.
[0174] like Figure 6 The qualitative display effect diagram shows the processing results of STGJMamer in random frames of different videos. Figure 6 It can be seen that the method of the present invention has a good estimation of human posture in various scenes.
[0175] Application Example 2: Implementation in a somatosensory interactive sailboat maneuvering simulation.
[0176] The first step is the somatosensory interactive sailboat manipulation simulation.
[0177] The present invention builds a Figure 2The somatosensory interaction-based sailboat manipulation simulation system shown in the figure applies the proposed STGJMamer in somatosensory interaction design, providing users with a natural interactive operating environment. The system consists of five modules. Specifically, the user interaction interface provides a simple and intuitive operating interface to help users achieve friendly interaction with the system; the voice control module controls the system UI button functions through voice commands, providing users with a convenient control experience; the sailing environment simulation module is responsible for building a virtual environment, simulating the real ocean environment and weather conditions, ensuring the authenticity and immersion of training; the sailing kinematics module simulates the motion feedback of the sailboat under different environmental conditions and provides real-time operation response; the teaching module and training module provide users with customized training content, helping users gradually improve their operating skills through different sailing tasks. Finally, the somatosensory control module implements interactive operations based on somatosensory movements, accurately capturing the user's body movements and converting them into corresponding sailboat operation instructions.
[0178] The second step is the application of STGJMamer in somatosensory interaction.
[0179] After obtaining the three-dimensional human body posture through the STGJMamer algorithm, the present invention uses TCP Socket communication to realize data transmission between Python and Unity3D, and uses Unity's Animator component to realize the binding of posture data and virtual characters.
[0180] like Figure 7 As shown in FIG4 , according to the sailboat operation task, the present invention designs six types of actions, namely, raising the canvas, lowering the canvas, turning the canvas left, turning the canvas right, steering, and resetting, and produces corresponding action feature descriptions (as shown in Table 4). Action recognition is performed through feature matching, thereby driving the sailboat in the virtual scene and realizing somatosensory interaction.
[0181] Table 4 Description of somatosensory interaction action characteristics
[0182]
[0183] The present invention uses a simple and intuitive performance evaluation metric in its experiments: the gesture recognition success rate, which is the number of successfully recognized actions divided by the total number of actions performed. This metric is calculated by repeatedly executing each navigation simulation control action.
[0184] The present invention uses a camera with a resolution of 1920×1080 and a frame capture rate of 30FPS, and tests each of the six control actions 200 times. During the experiment, the present invention uses the RTMPose 2D detector to extract 2D human postures from real-world scenes as data preprocessing. The extracted posture information is input into STGJMamer for three-dimensional human posture estimation, and then mapped to the sailboat manipulation simulation system. As shown in Table 4, the present invention defines specific evaluation criteria for each action at different execution stages. During the operation of the system, the system will continuously evaluate whether the current posture meets the predefined criteria. If it meets the criteria, the recognition action is successful; otherwise, it is considered unsuccessful. The experimental results are shown in Table 5. It can be seen that the minimum recognition accuracy of the posture recognition algorithm designed by the present invention reached 86.5%, and the recognition accuracy of the remaining actions was above 90%, which fully verified the good performance of the algorithm.
[0185] Table 5 Action recognition experimental results
[0186] action Execution times / times Number of successes / times Accuracy % Raised canvas 200 193 96.5 Lowering canvas 200 191 95.5 Turn right canvas 200 186 93.0 Turn left canvas 200 189 94.5 at the helm 200 173 86.5 Reset 200 188 94.0
[0187] The above description is only a preferred specific implementation method of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A somatosensory interactive sailboat maneuvering simulation method based on three-dimensional human posture estimation, characterized in that: The method comprises the following steps: S1, uses a camera to capture real-time images and uses a 2D posture detector to obtain 2D human posture; S2, inputs the 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate the 3D human pose; S3 maps the obtained 3D human body posture to the somatosensory control module in the sailing operation virtual scene to realize interactive operations based on somatosensory movements.
2. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 1, characterized in that: In step S1, the 2D human body posture is obtained by using a 2D posture detector to extract the 2D human body posture from the real-shot picture using RTMPose2D.
3. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 1, characterized in that: In step S2, the graph-guided state-space enhanced 3D human pose estimation method includes: S21, modify the PoseFormerV2 model into a baseline model suitable for real-time estimation, which takes as input a 2D pose sequence of t frames in the past of the target frame; S22, uses the SGJMamer layer to extract local spatial features from the past f-frame poses of the target frame, and uses the TGJMamer layer to extract the global spatiotemporal feature Y from the low-frequency information of the complete input and the spatial features extracted by SGJMamer; among them, the SGJMamer / TGJMamer layers are composed of the STE-GCN module, the JMamba module and the Transformer module; the Transformer modules in the SGJMamer / TGJMamer layers are the spatial Transformer module and the temporal Transformer module respectively; the JMamba module introduces a hierarchical joint-enhanced selective state space model to improve the model's ability to extract global features; the STE-GCN module learns the graph structure relationship of the joints through a spatiotemporal extended graph convolutional network; S23, constructing a regression head to infer the target frame joint point information, constructing a loss function to train the target frame joint points, and then outputting the target frame 3D posture.
4. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 3, characterized in that: In step S22, STGJMamer is used to extract the spatiotemporal features of the human posture model to obtain the output Y of the fusion of spatiotemporal features, including: S221, construct the graph state space Transformer layer SGJMamer to extract local spatial features; S222, construct the graph state time Transformer layer TGJMamer, fuse the low-frequency information of the complete input and the spatial features extracted in step S221, and perform spatiotemporal feature extraction to obtain the output Y of the fused spatiotemporal features.
5. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 3, characterized in that: Step S221, constructing a graph state space Transformer layer SGJMamer, and performing local space feature extraction includes: The target frame's past f frame information [x t-f ,x t ] is input, and passes through the STE-GCN module and the JMamba module in sequence to obtain the feature X SpaGJMamba ;The data will be fed into the spatial Transformer module; The STE-GCN module predefines the spatial connections between joints in the same frame and the temporal dependencies of the same joint in different frames by extending the adjacency matrix in time and space, and then extracts the spatiotemporal dependencies between joints through graph convolution. The process of the STE-GCN module to extract the spatiotemporal dependencies between joints is as follows: Where, is the normalized spatiotemporal adjacency matrix, D graph is the degree matrix, and the diagonal elements represent the connection degree of each node; graph is the adjacency matrix, including spatial and temporal relationships; H' graph is the output feature matrix, the dimension is E×T×N×F', σ is the activation function, H original is the input feature matrix with dimensions of E×T×N×F; W is the learnable weight matrix of graph convolution, and Drop() is the random drop mechanism; After graph convolution processing, H' will contain local information. We introduce hybrid residual design and use dynamic weighting to fuse the original features with graph convolution features. The formula is: X graph =α·H′ graph +(1-α)·H original Where, X graph is the feature result after the residual mechanism processing; α is a learnable weight parameter used to dynamically adjust the ratio of the two features; H' graph is the feature after graph convolution processing, H original are the original features without graph convolution processing.
6. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 5, characterized in that: The JMamba module uses a hierarchical joint enhancement method based on anatomical and kinematic chain reactions based on VMamba to extract features in the information transfer between related joints and complete pose estimation; The specific processing method of the JMamba module to complete the posture estimation includes: combining the hierarchical joint enhancement method based on anatomy and kinematic chain reaction, constructing a bidirectional to-be-scanned information X consisting of the enhanced joint point sequence and the original joint point sequence scan , and calculate the learnable matrices required for selective state space scanning: state matrix A, control matrix B, output matrix C, instruction matrix D and state variable H; X scan =f Scan (W in ·X in +b in ) Where, X scan is the bidirectional information to be scanned, f Scan is the initialization function of the bidirectional scan, which is used to project the original input into the state space; W in is the original joint point sequence, X in is the bias term, b in is the weight matrix; Perform local bidirectional scanning and global bidirectional scanning, continuously update the state space, and capture long-term and short-term dependencies. The expression is: H′=A·H+B·X scan Where H' is the output feature matrix; The dimensions are: E×T×N×F′ Where E is the batch size, T is the number of input frames, N is the number of joints, and F' is the feature corresponding to each joint; X SSM =C·H'+D·X scan Where, X SSM is the intermediate feature generated by selective state space scanning; After completing the bidirectional scanning, feature fusion is performed to obtain a complete feature expression rich in global and local information: X out =f Merge (X SSM ) Where, X out is the final output result of this module, f Merge To integrate the bidirectional scanning results into the original dimension; In the spatial Transformer module, the input features are first normalized and processed with multi-head self-attention to capture the intrinsic connections between joints. Subsequently, the features are normalized and processed with the MLP layer, and nonlinear transformations are introduced through the activation function GELU to learn the complex nonlinear relationships between the input features. Spatial feature extraction includes: Q,K,V=linear(LayerNorm(X SpaGJMamba )) x spatial =x res +Drop(x mlp ) Where Q is the query vector, K is the key vector, V is the value vector, linear() is the fully connected layer, LayerNorm() is the normalization process, and X SpaGJmamba is the spatial feature after processing by the STE-GCN module and the JMamba module, x res is the intermediate processing result, x attn is the feature processed by the attention mechanism, softmax() is to normalize the vector into a probability distribution vector, K T is the transpose of K, d k is the scaling factor, is the intermediate processing, Drop() is the random discarding processing, x mlp is the feature processed by the mlp mechanism, W1 and W2 are both weight matrices of linear transformation, GELU() is the activation function processing, b1 and b2 are both bias terms, and x spatial It is the spatial dimension feature in the target frame estimation process.
7. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 6, characterized in that: A hierarchical approach to joint strengthening based on anatomy and kinematics chain reactions includes: In the joint enhancement sequence, the enhancement sequence is designed based on the anatomical and kinematic chain reaction. The characteristics of the linkage between the father and son joints of the human body are utilized, and the joint chain reaction exhibited by the human body to maintain balance during movement is combined to improve the model feature extraction capability through mutual enhancement between joints. At the same time, symmetry enhancement is adopted in the data processing process, that is, a mirror flip operation is applied to the input data to generate symmetrical movements. In the posture estimation process, the postures of the original movement and the flipped movement are estimated at the same time. The final output is obtained by taking the average of the results of the original movement and the flipped movement to complete the enhancement of data symmetry.
8. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 4, characterized in that: In step S222, a graph state temporal Transformer layer TGJMamer is constructed to fuse the low-frequency information of the complete input and the spatial features extracted in step S221, and perform spatiotemporal feature extraction to obtain the output Y of the fused spatiotemporal features, including: The graph state temporal transformer layer consists of the STE-GCN module, the JMamba module, and the temporal transformer module. The data input is the low-frequency part of the complete action sequence, which captures the changing characteristics of the action from the global information and obtains the overall evolution trend of the action. After being processed by the STE-GCN module and the JMamba module, the frequency domain feature X is obtained. TGJMamba , the frequency domain information will be combined with the spatial feature x spatial The combined spatiotemporal information is sent to the self-attention mechanism for processing; when the processed data passes through the MLP layer, the spatial feature data will be transformed into the frequency domain to complete the conversion of local detail changes into frequency domain information, and to effectively supplement the low-frequency time information.
9. The somatosensory interactive sailboat maneuvering simulation method based on 3D human posture estimation according to claim 3, characterized in that: In step S23, the regression head is constructed to infer the target frame joint information, and the loss function is constructed to train the target frame joints. The joint point error is calculated by comparing the output of the regression head with the real 3D pose, including: The spatiotemporal feature Y extracted from the input two-dimensional sequence posture information has the following feature dimensions: represents the set to which Y belongs, f is the number of frames of the input feature; the target frame is the last frame of the sequence, and the dimension is Construct a regression head to infer the target frame joint information; the regression head consists of a weighted average operation and a fully connected layer. The estimated posture information will be calculated under the weight of the learnable weight parameters to calculate the final estimate Y of the target frame. predicate ; Y predicate =W head ·LayerNorm(neab(Y))+b head Where W head is the trainable weight, mean(Y) is the mean, b head is the bias term, LayerNorm() is the layer normalization; The output result will be the same as the true 3D pose Y tar get Calculate the joint point error MPJPE, and use MPJPE as the loss function to train the model; Where, are the real information and predicted information of the kth joint point respectively.
10. A somatosensory interactive sailboat maneuvering simulation system based on 3D human posture estimation, characterized in that: The system implements the somatosensory interactive sailboat maneuvering simulation method based on three-dimensional human posture estimation according to any one of claims 1 to 9, and the system comprises: 2D human posture capture module, used to use the camera to capture real-scene pictures in real time and use the 2D posture detector to obtain the 2D human posture; 3D human pose calculation module, used to input 2D human pose into the graph-guided state space enhanced 3D human pose estimation method to calculate 3D human pose; The somatosensory interaction module is used to map the obtained 3D human body posture to the somatosensory control module in the virtual scene of sailing operation to realize interactive operation based on somatosensory action.
Citation Information
Patent Citations
Motion mapping method and device of motion sensing game and computer storage medium
CN117547810A
Layer chain constraint three-dimensional human body posture estimation method based on monocular video stream
CN119131904A
Three-dimensional motion capture and intelligent analysis system and method based on monocular camera
CN119169701A
Cited By
Six-degree-of-freedom pose estimation system and method based on Mamba network
CN121505037A