Posture assessment method based on improved YOLOv8 key point detection and triangulation three-dimensional reconstruction
By improving YOLOv8 key point detection network and three-dimensional reconstruction technology, combined with dynamic multi-scale cross-attention and confidence weighting methods, the problems of high cost, low efficiency and insufficient three-dimensional reconstruction accuracy of traditional health monitoring methods are solved, and a high-precision, low-cost and easy-to-deploy body evaluation method is achieved.
Patent Information
- Application Number
- CN202510570789.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-17
AI Technical Summary
Traditional health monitoring methods have high cost, low efficiency, insufficient three-dimensional reconstruction accuracy and poor adaptability to complex scenarios, making it difficult to meet the needs of large-scale application scenarios, especially in the field of home or telemedicine.
Using the improved YOLOv8 key point detection network, high-precision body posture evaluation is achieved through DCA-ShuffleNetv2 lightweight network design, DEMA dynamic multi-scale cross attention module, SA-DWIOU structure-aware dynamic weighted loss function and confidence-weighted three-dimensional reconstruction technology, combined with the attitude-independent evaluation mechanism.
While ensuring high accuracy, it significantly reduces the amount of model parameters and calculation costs, and is suitable for real-time operation of low-computing equipment, improves detection stability and three-dimensional reconstruction accuracy in complex scenarios, and meets the needs of low cost and easy deployment.
Smart Images

Figure CN120164237A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields such as key point detection in the field of computer vision, and specifically relates to a body posture evaluation method based on improved YOLOv8 key point detection and triangulation 3D reconstruction. Background Art
[0002] Traditional health monitoring methods mostly rely on dedicated devices or physically contact sensors. Although they can provide a certain degree of accuracy, their complex operation processes and high deployment costs limit their popularity. Users need to wear devices or are restricted by specific environments, resulting in insufficient comfort and flexibility in actual use. In addition, such technologies are difficult to meet the needs of large-scale application scenarios, especially in the fields of home or telemedicine, where there are obvious shortcomings.
[0003] In recent years, non-contact monitoring technologies have gradually become a research hotspot, which realizes remote capture of human postures through visual information analysis. Such methods do not require complex hardware support, significantly reduce the usage threshold, and at the same time take into account user privacy and convenience. However, existing solutions still face challenges in terms of stability in complex scenarios, accuracy of 3D posture reconstruction, and the balance between real-time performance and computational efficiency. For example, factors such as dynamic occlusion and light changes are likely to cause misdetection of key points, and the sensitivity of 3D reconstruction algorithms to low-quality observation data also affects the reliability of the final evaluation.
[0004] In view of the above problems, there is an urgent need for a technical solution that takes into account lightweight, robustness and practicality, can achieve high-precision body posture evaluation under non-ideal conditions, and at the same time meet the requirements of low cost and easy deployment, providing a better solution for fields such as health management and rehabilitation training. Summary of the Invention
[0005] Aiming at the defects and deficiencies of the existing technology, the present invention provides a body posture evaluation method based on improved YOLOv8 key point detection and triangulation 3D reconstruction, which solves the problems of high cost, low efficiency, insufficient 3D reconstruction accuracy and poor adaptability to complex scenarios of traditional methods through lightweight network design, dynamic multi-scale feature fusion, confidence-weighted reconstruction and posture-independent evaluation mechanism. Specifically, it includes:
[0006] Lightweight key point detection network: Replace the YOLOv8 backbone network with DCA-ShuffleNetv2, and optimize feature extraction through dynamic channel attention mechanism and adaptive channel shuffle, significantly reducing the number of parameters and improving computational efficiency;
[0007] Multi-scale feature enhancement: Introduce the DEMA dynamic multi-scale cross-attention module in the neck network, and fuse multi-scale features through dynamic grouping strategy and parallel local / global / dilated pooling to enhance the robustness of key point detection under complex postures;
[0008] Bounding Box Regression Optimization: Design the SA-DWIOU structure-aware dynamic weighted loss function, dynamically adjust gradient allocation by combining the variance of the feature map and the outlier degree, and alleviate the interference of low-quality samples;
[0009] Confidence-Weighted 3D Reconstruction: Construct a weighted projection matrix based on the confidence of key points, and preferentially fuse high-confidence observation data through SVD decomposition to improve the accuracy of 3D coordinate reconstruction;
[0010] Posture-Independent Evaluation System: Calculate the human body orientation vector using the pelvic stability feature, combine the piecewise scoring mechanism of joint angles and slope parameters to achieve standardized body posture evaluation, and dynamically recommend personalized correction plans through web crawler technology in the optimized solution.
[0011] While ensuring high accuracy, this solution supports real-time operation on low-computing-power devices, is applicable to home, medical, and sports scenarios, and provides low-cost and highly robust technical support for body posture health management.
[0012] The technical solution specifically adopted by the present invention to solve its technical problems is:
[0013] A body posture evaluation method based on improved YOLOv8 key point detection and triangulation 3D reconstruction, including:
[0014] Dynamic Lightweight Key Point Detection:
[0015] Use the improved DCA-ShuffleNetv2 lightweight network to replace the backbone network of YOLOv8-Pose, and optimize feature extraction through dynamic channel attention and adaptive channel interaction;
[0016] Integrate the DEMA dynamic multi-scale cross-attention module in the neck network, and enhance multi-scale feature fusion through dynamic feature grouping and cross-group interaction strategies;
[0017] Adopt the SA-DWIOU structure-aware dynamic weighted loss function to jointly optimize bounding box regression through the variance of the feature map and dynamic attention weights;
[0018] Confidence-Weighted 3D Reconstruction:
[0019] Based on the 2D coordinates and confidence of key points from the front and side views, construct a weighted projection matrix, solve the 3D coordinates through matrix decomposition, where the weight factor of the observation data is positively correlated with its confidence;
[0020] Pelvic Stability Orientation Correction and Body Posture Evaluation:
[0021] Define the human body orientation using the outer product of the lumbar and hip joints in the pelvic region, and achieve posture-independent standardized evaluation through coordinate system rotation; generate body posture scores by combining the differences between joint angles, slopes and standard values.
[0022] Furthermore, the confidence-weighted three-dimensional reconstruction specifically includes:
[0023] Using the confidence of the frontal and lateral key points as a weight factor to construct a weighted projection equation, where the weight of each observation point is positively correlated with its confidence;
[0024] Solving the least squares solution of the weighted projection equation through singular value decomposition (SVD), so that the observation data with higher confidence has a greater contribution to the calculation result of the three-dimensional coordinates.
[0025] Furthermore, using the confidence of the two-dimensional key points as a weight factor;
[0026] The three-dimensional coordinates are solved from the weighted projection equation through singular value decomposition (SVD).
[0027] Furthermore, the method of defining the human body orientation by the outer product of the lumbar joint and the hip joint in the pelvic region and realizing pose-independent standardized evaluation through coordinate system rotation is specifically as follows:
[0028] Based on the spatial coordinates of the lumbar joint, right hip joint, and left hip joint, calculate the human body orientation vector through the outer product;
[0029] Rotate the three-dimensional model around the world coordinate system to make the front of the human body face the standardized direction;
[0030] The coordinate system rotation calculates the human body orientation vector through the outer product and realizes standardized alignment through mathematical transformations of sine and cosine functions to ensure that the parameters after pose standardization are used for body posture score calculation.
[0031] Furthermore, the improved lightweight DCA-ShuffleNetv2 network stacks 6 DCA-ShuffleNetv2 modules on the backbone network, including:
[0032] Dynamic Channel Attention (DCA):
[0033] Extract the global information of the feature map through global average pooling, and dynamically adjust the activation function threshold in combination with learnable parameters to avoid feature loss caused by fixed thresholds;
[0034] The output of the dynamic gating activation function is the pointwise product of the ReLU activation result and the Sigmoid gating weight;
[0035] Adaptive Channel Shuffle:
[0036] Dynamically rearrange the channels according to the channel importance score based on learnable weights and global average pooling. The channel importance score is calculated by the product of a learnable weight matrix of dimension C×C and the output of global average pooling;
[0037] When performing channel shuffle, channels with higher scores are preferentially retained to promote cross-group information interaction.
[0038] Furthermore, the DEMA dynamic multi-scale cross-attention module includes:
[0039] Dynamic feature grouping:
[0040] The number of groups for the input feature map is dynamically calculated through learnable weights, specifically as follows:
[0041] Perform global average pooling on the input feature map to generate a one-dimensional vector in the channel dimension;
[0042] Multiply the one-dimensional vector by a learnable weight matrix of dimension C×1, and scale it to a preset maximum number of groups after normalization through the Softmax function to obtain the dynamic number of groups;
[0043] Divide the feature map into multiple groups according to the dynamic number of groups, and the feature dimension of each group is adaptively adjusted.
[0044] Cross-group interaction strategy:
[0045] Perform the following operations in parallel on each group of feature maps:
[0046] Local pooling: Extract local context information through 3×3 average pooling;
[0047] Global pooling: Extract global context information through adaptive size pooling;
[0048] Dilated pooling: Extract large-range context information through pooling with a dilation rate of 2;
[0049] Concatenate the results of local, global, and dilated pooling, and generate cross-attention weights through 1×1 convolution;
[0050] Perform cross-group information interaction on the multi-scale features according to the attention weights to achieve feature fusion.
[0051] Furthermore, the structure-aware dynamic weighted loss function includes:
[0052] Quantification of regional semantic complexity: Calculate the pixel value dispersion of each region based on the feature map output by the neck network to characterize semantic complexity;
[0053] Coupling of localization error and complexity: Combine the feature map dispersion with the offset of the center point of the predicted box, and generate a dynamic adjustment coefficient through an exponential function, so that the loss weight of complex regions decays and the loss weight of simple regions increases; The dynamic adjustment coefficient is calculated through the outlier degree, and the outlier degree is the ratio of the center offset between the predicted box and the ground truth box to the feature variance;
[0054] Loss weighted fusion: Combine the dynamic adjustment coefficient with the attention weight to weight the IoU loss for the joint optimization of feature expression and bounding box regression.
[0055] And, a body posture evaluation system based on improved YOLOv8 keypoint detection and triangulation 3D reconstruction, comprising:
[0056] Dynamic lightweight keypoint detection module:
[0057] Deploy an improved DCA-ShuffleNetv2 lightweight network to replace the backbone network of YOLOv8-Pose, optimize feature extraction through the dynamic channel attention mechanism and adaptive channel interaction. Among them, the dynamic channel attention dynamically adjusts the activation function threshold through global average pooling and learnable parameters, and the adaptive channel interaction dynamically rearranges the channels based on the channel importance score;
[0058] Integrate the DEMA dynamic multi-scale cross-attention module in the neck network, realize the adaptive division of multi-scale features through dynamic feature grouping, and fuse multi-scale features of local pooling, global pooling and dilated pooling through the cross-group interaction strategy;
[0059] Adopt the SA-DWIOU structure-aware dynamic weighted loss function, generate a dynamic adjustment coefficient through the coupling of the feature map variance and the prediction box center offset, and optimize the bounding box regression training;
[0060] Confidence weighted 3D reconstruction module:
[0061] The input unit receives the two-dimensional coordinates and confidence of the keypoints from the front and side views. The weight calculation unit generates a weighted projection matrix based on the confidence, where the confidence of the two-dimensional keypoints is used as the weight factor for each observation point;
[0062] The solving unit solves the three-dimensional coordinates through matrix decomposition methods, including performing singular value decomposition on the weighted projection matrix, so that high-confidence observation data is given higher weight during the solving process.
[0063] Pelvic stability orientation correction and body posture evaluation module:
[0064] The orientation correction unit calculates the human body orientation vector based on the three-dimensional coordinates of the lumbar and hip joints, and aligns the three-dimensional model to the preset standard direction through coordinate system rotation;
[0065] The parameter calculation unit extracts joint angle and slope parameters and compares the differences with the anatomical standard values;
[0066] The scoring generation unit generates a body posture score through piecewise function mapping according to the difference value, and the scoring level gradually decreases when the error is within the preset threshold range;
[0067] Data interface module:
[0068] The input interface is connected to the image acquisition device, receives the front and side human body images and transmits them to the detection module;
[0069] The output interface transmits the body posture score and the standardized evaluation result to the display terminal or the user device.
[0070] Moreover, an electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned method are implemented.
[0071] A non-transitory computer-readable storage medium stores a computer program thereon. When the computer program is executed by the processor, the steps of the above-mentioned method are implemented.
[0072] Compared with the prior art, the present invention and its preferred solutions at least include the following beneficial effects:
[0073] Balance between lightweight and high precision: By optimizing the backbone network structure through lightweight network design, while significantly reducing the number of model parameters, the detection accuracy is maintained, taking into account both computational efficiency and performance, and is suitable for deployment on low-computing-power devices;
[0074] Enhanced robustness of multi-scale features: Introduce a dynamic multi-scale attention mechanism to strengthen the feature focusing ability on the human body key point areas, reduce complex background interference, and improve the detection stability in scenarios with occlusion or variable postures;
[0075] Optimization of the training process: Improve the loss function design, and alleviate the negative impact of sample quality differences on model training by dynamically adjusting the gradient allocation strategy, enhancing the effectiveness and generalization of parameter optimization;
[0076] Improved 3D reconstruction accuracy: Dynamically weighted fusion of multi-view observation data based on key point confidence, preferentially using high-quality matching points to optimize 3D coordinate calculation, reducing noise interference, and improving the reliability and consistency of the reconstruction results.
[0077] Through the collaborative design of lightweight network, dynamic feature fusion, training strategy optimization, and confidence-driven reconstruction, this solution provides an efficient, robust, and easy-to-implement technical path for body posture assessment. Description of the Drawings
[0078] The following further details the present invention in conjunction with the drawings and specific embodiments:
[0079] Figure 1 It is the schematic diagram for implementing the solution of the embodiment of the present invention. Specific Embodiments
[0080] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are given below for detailed description as follows:
[0081] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0082] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0083] As Figure 1 shown, this embodiment provides a body posture evaluation scheme based on improved YOLOv8 key point detection and triangulation 3D reconstruction. The implementation process of this scheme specifically includes the following steps:
[0084] Step S1: Obtain the Human36m public dataset. Using YOLOv8-Pose as the basic framework, transform the backbone network with the DCA-ShuffleNetv2 lightweight network, transform the neck network with the DEMA attention mechanism, change the loss function with SA-DWIOU, and perform multiple rounds of iterative training to obtain the optimal model. Input the front and side views of the human body, and through the inference of the transformed YOLOv8-Pose network, obtain the two-dimensional coordinates and their confidence levels of 16 joints on the front and side of the human body.
[0085] Step S2: According to the principle of pose transformation, define two rotation matrices for the front and side, construct the camera projection matrix, and use the key point confidence levels calculated in Step S1 as weight factors to restore the positions of the three-dimensional coordinates of the human joints in the world coordinate system based on the Triangulation method in the traditional ORB-SLAM system.
[0086] Step S3: Utilize the characteristics of the relative stability of the pelvis to correct the human body orientation. Then, conduct body posture evaluation through parameters such as joint angles, slopes, and outer products, and statistically analyze the evaluation results. In a preferred implementation expansion scheme, personalized suggestions and video resources are also provided to users in combination with crawler technology.
[0087] In this embodiment, Step S1 specifically includes the following steps:
[0088] Step S11: Obtain the Human36m public dataset, perform data preprocessing, and complete label extraction.
[0089] Step S12: In the backbone of YOLOv8-Pose, the original CSS (Cross Stage Partial Connections) structure of YOLOv8-Pose is replaced by DCA-ShuffleNetv2 modified based on ShuffleNetv2. Through the dynamic channel attention (DCA) mechanism and lightweight feature extraction design, the feature expression ability and computational efficiency of the network are significantly improved. DCA-ShuffleNetv2 adopts a "split-recombine" structure, including two branches, branch1 and branch2.
[0090] For branch1, if the stride is 1, no operation is performed; if the stride is 2, the input passes through a 3×3 depthwise separable convolution (DWConv), a batch normalization layer (BatchNorm2d), a 1×1 grouped convolution (GConv), and a batch normalization layer and a dynamic gating activation function DGR. The dynamic gating activation function DGR is:
[0091] DGR = ReLU(x DGR-input )·Sigmoid(W DGR ·GAP(x DGR-input ) + b DGR )
[0092] where W DGR and b DGR are learnable parameters, and GAP(·) is global average pooling. The DGA activation function is used to replace the ReLU activation function in the original ShuffleNetv2, and the activation threshold is dynamically adjusted through global features to avoid feature loss caused by the fixed threshold of ReLU.
[0093] For branch2, the cross-dimensional hybrid attention CDMA is used to replace the static convolutional chain structure adopted by the original branch2. The input x branch2 is compressed in channels through a 1x1 ordinary convolution Conv, and then the spatial attention A space and the channel attention A channel are executed in parallel, and finally cross-dimensional feature fusion is performed to obtain Fusion sc . The specific calculation method is as follows:
[0094] A space = Sigmoid(Conv 3×3 (ReLU(Conv 1×1 (x branch2 ))))
[0095] A channel= Sigmoid(MLP(GAP(x branch2 )))
[0096] Fusion sc = x branch2 *(A space + A channel )
[0097] Among them, MLP(·) is a multi-layer perceptron used to learn the dependencies between channels. * represents element-wise multiplication, used to dynamically adjust the feature response. Replace the static convolutional chain structure of the original branch2 with CDMA, while capturing the dependencies in both the spatial and channel dimensions, making up for the deficiency of the single attention mechanism in ShuffleNetv2.
[0098] Finally, perform Concat feature fusion on the outputs of branch1 and branch2 to obtain F concat , use Adaptive Shuffle to replace the original static channel shuffle for the case of stride 2, and then reshape the output tensor F shuffled back to the original shape (n, c, h, w). The specific process of Adaptive Shuffle is as follows:
[0099] F shuffled = persume(F concat , argsort(W rank ·GAP(F concat )))
[0100] Among them, W rank ∈R c×c is a learnable weight matrix, W rank ·GAP(F concat ) is the channel importance score, argsort(·) dynamically rearranges the channels according to the score, preferentially retaining the high-response channels, and persume(·) is the channel shuffle, which rearranges the channel dimension to promote information exchange between different groups. Adaptive Shuffle dynamically adjusts the shuffle strategy based on channel importance, abandons the original fixed uniform shuffle, and promotes cross-group interaction of key information.
[0101] In this embodiment, 6 DCA-ShuffleNetv2 modules are stacked in the backbone network part of YOLOv8-Pose, gradually extracting features and reducing the spatial resolution. The number of parameters of the CCS structure of the YOLOv8-Pose network is significantly reduced, and the overall performance of the network model is comprehensively considered, featuring high precision and small computational load.
[0102] Step S13: For the Neck part of YOLOv8-Pose, add the Dynamic Multi-Scale Cross Attention DEMA module based on the improved Efficient Multi-Scale Attention EMA. Utilize parallel sub-structures and cross-space feature fusion to make the network more sensitive to the human body and its key points. The core design of the DEMA attention mechanism is to capture dependencies in the spatial and channel dimensions through grouping operations and adaptive pooling. The specific steps are as follows.
[0103] The first step is dynamic feature grouping. For the input feature map X Neck ∈R b×c×h×w , dynamically calculate the number of groups g, breaking through the limitation of the fixed grouping in traditional EMA. Dynamically divide the feature groups through learnable parameters to achieve multi-scale adaptability. After grouping, the feature map X groups ∈R b×g×c / g×h×w .
[0104] g = Softmax(W g ·GAP(X Neck ))·G max
[0105] where W g ∈R c is the learnable weight, G max is the maximum number of groups, and finally g is rounded and restricted to 2 ≤ g ≤ G max .
[0106] The second step is multi-scale pyramid pooling. Perform multi-scale pooling on X groups in parallel, including local pooling X local , global pooling X global , and dilated pooling X dilated .
[0107] X local = AvgPool 3×3 (X groups )
[0108] X global = AvgPool adaptive (X groups )
[0109] X dilated = AvgPool dilated=2 (X groups )
[0110] The third step is cross-attention fusion. Concatenate and fuse the multi-scale pooling results to obtain A cross ∈R b×g×3×1×1 . This cross-group interaction strategy solves the problem of information isolation between groups in traditional EMA grouped attention.
[0111] X multi = Concat(X local , X global , X dilated )
[0112] A cross = Sigmoid(Conv 1×1 (X multi ))
[0113] Step 4, cross-group feature interaction, perform cross-group information exchange on the weighted features to obtain X interact .
[0114]
[0115] Step 5, dynamic weight generation, the final attention weight is obtained through channel-spatial joint calibration to get the DEMA learning weight W DEMA .
[0116] W DEMA = Sigmoid(MLP(GAP(X interact )) + Conv 3×3 (X interact ))
[0117] Apply the weight to the original output
[0118]
[0119] In this embodiment, the integration position of the DEMA module is between the three connection lines from the C2f module in the Neck network layer to the Head module in the Head network layer. By using the parallel sub-structure and cross-space feature fusion of the DEMA attention mechanism, the computational weight of the human body and its key point regions during the network prediction process is increased, enabling the network to better utilize the feature information related to the human body and its key points, and reducing the impact of uneven sample distribution and background on recognition.
[0120] Step S14: For the Head network part of YOLOv8-Pose, introduce the SA-DWIOU loss function modified based on the Wise-IoU (WIOU) structure-aware dynamic weighting to enhance the bounding box regression ability of the network model. SA-DWIOU uses a dynamic non-monotonic focusing mechanism to use "outlier degree" to replace IoU for quality assessment of anchor boxes and provides a wise gradient gain allocation strategy. The SA-DWIOU loss function L SA-DWIOU The specific formula is
[0121]
[0122] Among them, Atten iRepresents the dynamic attention weight, Represents the feature map output by the i-th key point of the neck network. R DWIOU Is the dispersion adjustment coefficient, Is the feature map The variance of, representing the regional semantic complexity, γ is a learnable parameter, initialized to 1.0, ∈ is a minimum value to prevent division-by-zero errors, (point x , point y ) and Respectively represent the center point of the predicted bounding box and the center point of the ground truth bounding box.
[0123] SA-DWIOU uses the feature map generated by the neck network for the adjustment of the dynamic attention weight Atten i And the dispersion adjustment coefficient R DWIOU , not only quantifies the regional semantic complexity, but also realizes the deep coupling of bounding box regression and feature expression, breaking through the problems of traditional WIOU's static weight allocation and insufficient regional sensitivity.
[0124] Step S15: Divide the dataset into a training set and a test set according to a ratio of 7:3. During the training process, use ablation experiments to continuously save the optimal model. Finally, input the front and side views of the human body, and through the inference of the modified YOLOv8-Pose network, obtain the two-dimensional coordinates and confidence levels of 16 joints of the front and side views of the human body.
[0125] In this embodiment, step S2 specifically includes the following steps:
[0126] Step S21: The world coordinate system takes the center of the ankle when the human body stands as the origin, where the positive direction of the x-axis points to the front of the human body facing, the positive direction of the y-axis points to the left side of the human body, and the positive direction of the z-axis points to the top of the human body. According to the conversion principle of the world coordinate system and the camera coordinate system, two rotation matrices are customized.
[0127]
[0128] Rotation1 is the rotation matrix of the camera when shooting from the front.
[0129]
[0130] Rotation2 is the rotation matrix of the camera when shooting from the side (left side of the human body).
[0131] Step S22: Based on the Triangulation method in the traditional ORB-SLAM system, combined with the two-dimensional coordinates and confidence information of 16 key points of the front and side views of the human body obtained in S1, calculate the three-dimensional coordinate information of the human body key points.
[0132]
[0133] where f x and f y are the focal lengths, c x and c y are the image center coordinates, Rotation is the rotation matrix, Translation is the translation vector, and the subscripts 1 and 2 refer to the camera parameters of the front and side respectively.
[0134] CameraProjection1 = intrinsics1 × extrinsics1
[0135] CameraProjection2 = intrinsics2 × extrinsics2
[0136] Multiply the camera intrinsic (intrinsics) and extrinsic (extrinsics) matrices to obtain the camera projection matrix (CameraProjection), where the subscripts 1 and 2 refer to the camera parameters of the front and side respectively. Then perform Triangulation 3D reconstruction, and the specific calculation method is as follows
[0137] A weight × Point 3D = 0
[0138]
[0139] where (v, u) and (v', u') are the 2D coordinates of the same key point of the human body at two angles of the front and side respectively, and conf1 and conf2 are the key point confidences. Point 3D is the 3D coordinate of the key point to be solved.
[0140] Step S23: Use SVD decomposition to find the least squares solution of A and solve AP = 0 to obtain the solution.
[0141]
[0142] where U Orthogonal ∈ R 3×3 is an orthogonal matrix, Σ ∈ R 3×4 is a diagonal matrix, and V Orthogonal ∈ R 4×4 is an orthogonal matrix. Scale the last column of V Orthogonal to homogeneous coordinate form to obtain the 3D coordinate of the key point Point 3D .
[0143] In this embodiment, step S3 specifically includes the following steps:
[0144] Step S31: According to the relative stability of the pelvis, a vector composed of the lumbar joint and the left and right hip joints near the pelvis is used to define the human body orientation. The specific calculation formula is as follows:
[0145]
[0146] Where O, R, and L are the lumbar joint, right hip joint, and left hip joint on the human bone model, respectively, then The corresponding direction is called the human body orientation. Essentially, the human body orientation is obtained by using the cross product, and through the mathematical transformation of the sine and cosine functions, the coordinate system of the three-dimensional model is rotated to ensure that all human body samples face the positive x-axis direction.
[0147] Step S32: By calculating the differences between parameters such as joint angles and slopes and the standards, within a certain error range, the human joints are scored. The scoring mechanism uses a piecewise function, with a score of A within one error range, a score of B between one and two error ranges, and so on. The detected joint disorders include "head tilt", "forward neck", "unequal shoulders", "hunchback", "scoliosis", "pelvic anterior / posterior tilt", "X / O-shaped legs", "knee hyperextension", etc.
[0148] As an extended preferred solution, this embodiment also provides Step S33: For the detected joint disorders, treatment suggestions are provided. At the same time, using the crawler technology, according to the keywords of the disorders, the video with the highest score among the top x num videos in the comprehensive ranking on the Bilibili platform is crawled, providing more diversified, easy-to-understand, and more efficient auxiliary treatment information and rehabilitation guidance for patients. The score recommend calculation formula is as follows: recommend The calculation formula is as follows:
[0149]
[0150] Where views represents the number of views, likes represents the number of likes, coins represents the number of coins, favorites represents the number of favorites, shares represents the number of shares, danmu represents the number of bullet screens, and w1, w2, w3, w4, w5, w6 are the weight coefficients of each interaction index respectively.
[0151] x num The calculation of x is dynamically adjusted according to the http request, content acquisition, and parsing time time response That is, the smaller the time response is, the larger the x num is.
[0152] Based on the same inventive concept, the present invention further provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically for loading and executing one or more instructions in the computer storage medium to implement the above method.
[0153] It should be further noted that, based on the same inventive concept, the present invention further provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.
[0154] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "linked" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to indicate relative position relationships, and when the absolute position of the object being described changes, the relative position relationship may also change accordingly.
[0155] As described above, it is only the preferred embodiment of the present invention, and it is not a limitation to the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
[0156] The present invention is not limited to the above best implementation mode. Anyone inspired by the present invention can obtain various other forms of body posture assessment methods based on improving YOLOv8 key point detection and triangulation 3D reconstruction. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. A posture assessment method based on improved YOLOv8 key point detection and triangulated 3D reconstruction, characterized in that: include: Dynamic lightweight key point detection: Use the improved DCA-ShuffleNetv2 lightweight network to replace the YOLOv8-Pose backbone network, and optimize feature extraction through dynamic channel attention and adaptive channel interaction; Integrate the DEMA dynamic multi-scale cross-attention module in the neck network to enhance multi-scale feature fusion through dynamic feature grouping and cross-group interaction strategy; The SA-DWIOU structure-aware dynamic weighted loss function is used to jointly optimize bounding box regression by using feature map variance and dynamic attention weights. Confidence-weighted 3D reconstruction: Based on the two-dimensional coordinates of the key points of the front and side view and their confidence, a weighted projection matrix is constructed, and the three-dimensional coordinates are solved by matrix decomposition, where the weight factor of the observation data is positively correlated with its confidence; Pelvic stability orientation correction and posture assessment: The body orientation is defined by the external products of the waist and hip joints in the pelvic area, and a posture-independent standardized assessment is achieved through coordinate system rotation. The posture score is generated by combining the differences in joint angles and slopes with standard values.
2. The posture assessment method based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The confidence-weighted three-dimensional reconstruction specifically includes: The confidence of the key points on the front and side surfaces is used as a weight factor to construct a weighted projection equation, in which the weight of each observation point is positively correlated with its confidence; The least square solution of the weighted projection equation is solved by singular value decomposition (SVD), so that observation data with higher confidence can make a greater contribution to the solution result of the three-dimensional coordinates.
3. The posture assessment method based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The confidence of the two-dimensional key points is used as the weight factor; The weighted projection equation is solved for three-dimensional coordinates by singular value decomposition (SVD).
4. The method for posture assessment based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The method of defining the human body orientation by using the outer product of the waist joint and the hip joint in the pelvic region and realizing posture-independent standardized evaluation by rotating the coordinate system is specifically as follows: Based on the spatial coordinates of the waist joint, right hip joint, and left hip joint, the body orientation vector is calculated by the outer product; Rotate the 3D model around the world coordinate system so that the front of the human body faces the standardized direction; The coordinate system rotation calculates the human body orientation vector through the outer product, and combines the mathematical transformation of sine and cosine functions to achieve standardized alignment to ensure that the parameters after posture standardization are used for posture score calculation.
5. The method for posture assessment based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The improved DCA-ShuffleNetv2 lightweight network uses 6 DCA-ShuffleNetv2 modules stacked on the backbone network, including: Dynamic Channel Attention DCA: The global information of the feature map is extracted through global average pooling, and the threshold of the activation function is dynamically adjusted in combination with learnable parameters to avoid feature loss caused by a fixed threshold. The output of the dynamic gate activation function is the point-by-point product of the ReLU activation result and the Sigmoid gate weight; Adaptive channel shuffling: Dynamically reorder channels according to a channel importance score based on learnable weights and global average pooling, where the channel importance score is calculated by multiplying a learnable weight matrix of dimension C×C with the output of global average pooling; When shuffling channels, channels with higher scores are prioritized to promote cross-group information interaction.
6. The posture assessment method based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The DEMA dynamic multi-scale cross-attention module includes: Dynamic feature grouping: The number of groups of the input feature map is dynamically calculated using learnable weights, specifically: Perform global average pooling on the input feature map to generate a one-dimensional vector of channel dimension; The one-dimensional vector is multiplied by a learnable weight matrix of dimension C×1, and the result is normalized by a Softmax function and then scaled to a preset maximum number of groups to obtain a dynamic number of groups; The feature map is divided into multiple groups according to the number of dynamic groups, and the feature dimension of each group is adaptively adjusted; Cross-group interaction strategy: For each set of feature maps, the following operations are performed in parallel: Local pooling: extract local context information through 3×3 average pooling; Global pooling: extracting global context information through adaptive size pooling; Dilated pooling: extracting large-scale context information through pooling with a dilation rate of 2; Concatenate the results of local, global, and dilated pooling, and generate cross attention weights through 1×1 convolution; The multi-scale features are interacted across groups according to the attention weights to achieve feature fusion.
7. The posture assessment method based on improved YOLOv8 key point detection and triangulated 3D reconstruction according to claim 1, characterized in that: The structure-aware dynamic weighted loss function includes: Quantification of regional semantic complexity: The discrete degree of pixel values in each region is calculated based on the feature map output by the neck network to characterize the semantic complexity; Coupling of positioning error and complexity: The feature map discreteness is combined with the offset of the center point of the prediction box, and a dynamic adjustment coefficient is generated through an exponential function, so that the loss weight of the complex area is attenuated and the loss weight of the simple area is enhanced; the dynamic adjustment coefficient is calculated by the outlier degree, which is the ratio of the center offset of the prediction box and the real box to the feature variance; Loss weighted fusion: The dynamic adjustment coefficient is combined with the attention weight to weight the IoU loss to achieve joint optimization of feature expression and bounding box regression.
8. A posture assessment system based on improved YOLOv8 key point detection and triangulated 3D reconstruction, characterized in that: include: Dynamic lightweight key point detection module: Deploy the improved DCA-ShuffleNetv2 lightweight network to replace the YOLOv8-Pose backbone network, and optimize feature extraction through dynamic channel attention mechanism and adaptive channel interaction. Dynamic channel attention dynamically adjusts the activation function threshold through global average pooling and learnable parameters, and adaptive channel interaction dynamically rearranges channels based on channel importance scores. The DEMA dynamic multi-scale cross-attention module is integrated in the neck network to achieve adaptive division of multi-scale features through dynamic feature grouping, and to fuse multi-scale features of local pooling, global pooling and dilated pooling through cross-group interaction strategy; The SA-DWIOU structure-aware dynamic weighted loss function is used to generate a dynamic adjustment coefficient by coupling the feature map variance with the prediction box center offset to optimize the bounding box regression training. Confidence-weighted 3D reconstruction module: The input unit receives the two-dimensional coordinates of the key points of the front and side views and their confidences, and the weight calculation unit generates a weighted projection matrix based on the confidences, wherein the confidence of the two-dimensional key points of each observation point is used as a weight factor; The solving unit solves the three-dimensional coordinates by matrix decomposition method, including singular value decomposition of weighted projection matrix, so that high confidence observation data are given higher weight in the solving process; Pelvic stability orientation correction and posture assessment module: The orientation correction unit calculates the human body orientation vector based on the three-dimensional coordinates of the waist joint and the hip joint, and aligns the three-dimensional model to a preset standardized direction by rotating the coordinate system; The parameter calculation unit extracts the joint angle and slope parameters and compares the differences with the anatomical standard values; The scoring generating unit generates a posture score through piecewise function mapping according to the difference value, and the score level decreases step by step when the error is within the preset threshold range; Data interface module: The input interface is connected to the image acquisition device to receive the front and side human body images and transmit them to the detection module; The output interface transmits the posture score and standardized assessment results to the display terminal or user device.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the methods of claims 1-7 when executing the program.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.