Intelligent medical emergency teaching system based on key point detection

By adopting multimodal key point detection technology in the intelligent medical emergency teaching system, the insufficient detection accuracy of human key point detection in complex backgrounds and occlusion problems in multi-person scenarios are solved, and a high-precision 3D action tracking and quantitative evaluation of monocular objects is achieved, meeting the needs of real-time and low-cost equipment deployment.

CN120148121APending Publication Date: 2025-06-13FUZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510571613.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has insufficient detection accuracy of human key points in complex backgrounds, and the occlusion problem is prominent in multi-person scenarios, and the real-time requirements are difficult to meet.

Method used

An intelligent medical emergency teaching system based on multimodal key point detection is adopted, including dynamic light compensation module, key point detection with anatomical constraint enhancement, multi-assumption cross-modal 3D pose reconstruction and contactless multi-dimensional scoring feedback.

Benefits of technology

It improves the robustness and accuracy of key point detection, reduces the error detection rate and joint positioning deviation in multi-person scenarios, and realizes high-precision 3D action tracking and quantitative evaluation of monocular objects, meeting the needs of real-time and low-cost equipment deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148121A_ABST
    Figure CN120148121A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent medical emergency teaching system based on key point detection, and the system comprises a self-adaptive image enhancement module which is used for carrying out local-global layered optimization on an input video frame, and comprises the steps: simulating RAW-RGB data through an inverse mapping function and Gaussian noise; the cascaded lightweight PEM module extracts local features and enhances illumination; generating a global color matrix and a Gamma value based on a dynamic Query mechanism, and performing end-to-end optimization through orthogonal constraint and joint loss; the key point detection module is used for extracting human body key points of an operator from the enhanced video frame; the three-dimensional attitude reconstruction module is used for mapping the key points to a three-dimensional space and comprises the steps that a multi-hypothesis generator generates initial 3D attitude hypotheses in parallel; a multi-head cross attention mechanism of cross-hypothesis interaction is fused with spatial-temporal features; optimizing an attitude sequence based on geometric constraint loss of a joint length error and an angle error; and the multi-modal scoring module is used for quantitatively evaluating the operation action.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields such as computer vision, and particularly relates to an intelligent medical emergency teaching system based on key point detection. Background Art

[0002] In recent years, the human key point detection technology has attracted extensive research attention in the field of computer vision, and has had a profound impact on fields such as intelligent healthcare, motion analysis, and human-computer interaction. The human key point detection task aims to locate the positions of human joint points (such as the head, shoulders, elbows, knees, etc.) in an image or video, and further infer the pose information of the human body. This is a challenging task, and its difficulties lie in: (1) the diversity and complexity of human poses, such as different poses, perspectives, occlusions, etc.; (2) large intra-class differences and small inter-class differences: the same pose may show significant differences under different lighting, background, or occlusion conditions, while there may be only subtle differences between different poses; (3) key point detection in a multi-person scenario needs to handle the interaction and occlusion problems between people.

[0003] Benefiting from the powerful perception ability of deep neural networks, the human key point detection technology has made significant progress.

[0004] For multi-person key point detection methods, they are mainly divided into two paradigms: top-down and bottom-up. The top-down method first detects each human body in the image, and then crops it proportionally and identifies each key point. In this method, missed detections, false detections, and the intersection over union (IoU) of the human body boxes will all affect the final detection results. The bottom-up method first detects all key points in the image, and then clusters them into different human body instances. The Crowdpose network model designs a corresponding loss function using the candidate point idea, thereby suppressing the key points that do not belong to the current human body and completing the detection of key points for dense crowds. With the continuous development of deep learning technology, the human key point detection technology has made significant progress. However, in practical applications, there are still many challenges, such as the detection accuracy under complex backgrounds, the occlusion problem in multi-person scenarios, and the real-time requirements, etc. Future research will pay more attention to the robustness, efficiency, and generalization ability of algorithms to meet the needs of more practical scenarios. Summary of the Invention

[0005] Aiming at the defects and deficiencies existing in the prior art, the present invention provides an intelligent medical emergency teaching system based on multi-modal key point detection, and its core innovative design points include:

[0006] Dynamic light compensation module: By means of a local-global hierarchical optimization strategy (a lightweight PEM module extracts local features + a dynamic Query mechanism generates global parameters), it solves the problem of image feature loss under complex lighting, and improves the robustness of key point detection;

[0007] Anatomical Constraint-Enhanced Key Point Detection: Based on a spatial attention mechanism and predefined anatomical length / angle constraints, combined with a candidate box aspect ratio screening algorithm, it significantly reduces the false detection rate and joint positioning deviation in multi-person scenarios;

[0008] Multi-Hypothesis Cross-Modal 3D Pose Reconstruction: Using a multi-hypothesis generator to predict 3D poses in parallel, through cross-hypothesis attention interaction and human skeleton geometric constraints (joint length, angle smoothness), it breaks through the depth ambiguity limitation of monocular vision and achieves millimeter-level precision tracking;

[0009] Contactless Multi-Dimensional Scoring Feedback: Fusing dynamic time warping (DTW) spatio-temporal trajectory matching, clinical indicators (compression depth / frequency), and biomechanical parameters (trunk inclination, arm straightness) to generate a quantitative evaluation result, and through the geometric mapping of the trunk plane normal vector and the ground coordinate system, it provides real-time feedback on the perpendicularity deviation of the compression axis.

[0010] Its significant technical advantages are reflected in:

[0011] Monocular High Precision: Only requires an ordinary RGB camera, replacing depth sensors such as Kinect, reducing equipment costs and deployment thresholds;

[0012] Multi-Modal Collaboration: Deep integration of dynamic light optimization, anatomical priors, and cross-hypothesis modeling to solve the pain points of light sensitivity, high false detection rate, and insufficient accuracy in traditional solutions;

[0013] Clinical Guidance: Quantitative feedback based on medical standards (such as CPR compression perpendicularity requirements) to improve the standardization of training actions and the effectiveness of mechanical conduction.

[0014] The technical solution specifically adopted by the present invention to solve its technical problems is:

[0015] An intelligent medical emergency teaching system based on key point detection, comprising:

[0016] An adaptive image enhancement module for performing local-global hierarchical optimization on the input video frame, including:

[0017] Simulating RAW-RGB data through an inverse mapping function and Gaussian noise;

[0018] A cascaded lightweight PEM module for extracting local features and enhancing light, where the PEM module is composed of depthwise separable convolution and a lightweight Transformer;

[0019] Generating a global color matrix and Gamma value based on a dynamic Query mechanism, and performing end-to-end optimization through orthogonal constraints and joint loss;

[0020] The key point detection module is used to extract the operator's body key points from the enhanced video frames, including:

[0021] Enhancing the key point heat map regression with a spatial attention mechanism;

[0022] Optimizing the key point localization with a structural consistency loss based on predefined anatomical length and angle constraints;

[0023] Screening the target operator's key points according to the aspect ratio of the candidate box;

[0024] The 3D pose reconstruction module is used to map the key points to the 3D space, including:

[0025] The multi-hypothesis generator generates initial 3D pose hypotheses in parallel;

[0026] Fusing spatio-temporal features with a multi-head cross-attention mechanism for cross-hypothesis interaction;

[0027] Optimizing the pose sequence with a geometric constraint loss based on joint length error and angle error;

[0028] The multi-modal scoring module is used to quantitatively evaluate the operation actions, including:

[0029] Matching the spatio-temporal similarity between the operator's trajectory and the standard template with the dynamic time warping algorithm;

[0030] Multi-dimensional parameter weighted scoring by fusing the pressing depth, frequency, torso inclination angle, and arm straightness;

[0031] Generating real-time feedback instructions according to the scoring difference.

[0032] Furthermore, the cascaded lightweight PEM module includes:

[0033] The depthwise separable convolutional layer is used to extract local illumination features;

[0034] The lightweight Transformer layer models the global illumination dependence relationship through the multi-head self-attention mechanism;

[0035] The feature fusion unit spatially aligns the local features output by the depthwise separable convolutional layer and the global features output by the Transformer layer to generate a multiplication graph M and an addition graph A, where:

[0036] The multiplication graph M is multiplied element-wise with the original input image X to enhance the local illumination contrast;

[0037] The addition graph A is added element-wise to the multiplication result to compensate for the illumination deviation;

[0038] Finally, the locally enhanced image X is output local ;

[0039] The depthwise separable convolution layer extracts local illumination features through a 3×3 convolution kernel.

[0040] Furthermore, in the dynamic Query mechanism, the generation of the global Gamma value includes the following steps:

[0041] Feature extraction: Extract the illumination-sensitive channels w of the multi-head attention intermediate feature vectors from the intermediate feature vectors output by the multi-head attention mechanism γ ;

[0042] Nonlinear mapping: Input w γ into the Sigmoid function to constrain its value range to the interval (0,1);

[0043] Dynamic adjustment: Perform a linear transformation on the Sigmoid output to map the Gamma value range to [0.5, 2.0], and the formula is expressed as:

[0044] Gamma value = 1.5×Sigmoid(w γ ) + 0.5.

[0045] Furthermore, the anatomical structure length constraint is implemented based on the predefined anatomical length and includes:

[0046] Set the predefined anatomical length between adjacent key points according to the standard human anatomical data;

[0047] Calculate the square of the difference between the predicted key point spacing and the predefined anatomical length to generate a structural consistency loss;

[0048] The optimization method of the anatomical structure length constraint includes the following steps:

[0049] Definition of adjacent key point pairs: According to the prior knowledge of human anatomy, predefined the set of key point pairs to be constrained, including connected joints or bone endpoints;

[0050] Setting of the standard length: Set the corresponding standard anatomical length lc for each key point pair (ti, tj) ti-tj ;

[0051] Calculation of the distance error: Extract the predicted key point coordinates p ti and p tj , and calculate the square of the difference between its Euclidean distance and the standard length;

[0052] Structural consistency loss: Sum the squares of the differences for all key point pairs, and the formula is expressed as:

[0053] Structural consistency loss = Σ(predicted distance - standard anatomical length) 2 ;

[0054] Model optimization: By minimizing the structural consistency loss, the key point localization is constrained to conform to the human anatomical structure.

[0055] Furthermore, the geometric constraint loss optimizes the physical rationality of 3D pose reconstruction through the following steps:

[0056] Joint length constraint:

[0057] Pre-define the medical standard lengths of connected joint pairs in the human skeleton;

[0058] Calculate the Euclidean distance between the predicted joint coordinates and obtain the absolute difference from the medical standard length;

[0059] Sum up the differences of all connected joint pairs to generate the joint length error;

[0060] Joint angle constraint:

[0061] According to the human kinematics standard, pre-define the medical standard angle ranges of human key joints;

[0062] Calculate the included angle between adjacent bone vectors based on the predicted joint coordinates and obtain the cosine similarity difference from the medical standard angle range;

[0063] Sum up the squared differences of all key joints to generate the joint angle error;

[0064] Global position constraint:

[0065] Calculate the squared Euclidean distance between the predicted 3D joint coordinates and the true coordinates;

[0066] Sum up the squared distances of all joints and take the average to generate the global position error;

[0067] Multi-loss collaborative optimization: Weightedly sum up the joint length error, joint angle error and global position error to drive the model to generate a 3D pose sequence that conforms to human biomechanics.

[0068] Furthermore, the weighted score is generated through the following steps:

[0069] Multi-modal parameter extraction: Extract clinical indicators including pressing depth and frequency and biomechanical parameters including trunk inclination and arm straightness from the 3D pose sequence;

[0070] Dynamic weight assignment: Automatically generate the weight w of each parameter through a pre-trained model according to the influence degree of the parameter on the operation standardization; fi The pre-trained model optimizes the weight parameters through the expert score annotation in the CPR operation dataset;

[0071] Weighted score calculation: Multiply the score sc of each parameter by its corresponding weight w; fiMultiply with the corresponding weight w fi Sum them up and output the final score sc final ;

[0072] Feedback instruction generation: Based on the final score sc final ; Generate real-time correction instructions for pressing force, frequency or posture according to the deviation from the standard threshold.

[0073] Furthermore, in the key point detection module, the candidate box screening realizes attention focusing through the following steps:

[0074] Calculation of the aspect ratio of the candidate box: For all detected candidate boxes, calculate the ratio of their width to height;

[0075] Target operator screening: Select the candidate box with the aspect ratio closest to 1:1 as the target operator, and the calculation formula is:

[0076] Screening score = min(candidate box width, candidate box height) / max(candidate box width, candidate box height);

[0077] False detection suppression: Exclude the interference targets with abnormal aspect ratios through the screening score to improve the accuracy of key point detection; The aspect ratio screening threshold is set to 0.8 - 1.2.

[0078] Furthermore, the cross-hypothesis interaction fuses multi-branch features through the following steps:

[0079] Hypothesis feature extraction: Extract the spatio-temporal features z qi , z qj , z qk ; containing joint velocity and temporal displacement from multiple parallel hypothesis branches

[0080] Interactive attention construction:

[0081] Take the features of different hypothesis branches as query, key, and value and input them into the multi-head cross-attention layer respectively;

[0082] Model the complementary dependence relationship between different hypotheses through attention weight calculation;

[0083] Feature fusion output: Weightedly fuse the cross-hypothesis attention output with the original features of each hypothesis branch to generate a physically reasonable 3D pose sequence.

[0084] Furthermore, the system realizes contactless 3D motion tracking through a monocular RGB camera and outputs the included angle between the pressing axis and the ground space with millimeter-level accuracy.

[0085] Furthermore, the calculation method of the included angle between the pressing axis and the ground space includes:

[0086] Trunk plane construction: Based on the key point coordinates of the shoulders and hips in the three-dimensional pose sequence, the trunk plane equation is fitted by the least squares method;

[0087] Normal vector generation: Extract the normal vector from the trunk plane equation to represent the spatial direction of the pressing axis;

[0088] Ground coordinate system mapping: Pre-define the standard normal vector direction of the ground coordinate system;

[0089] Spatial angle calculation: Determine the perpendicularity deviation between the pressing axis and the ground according to the angle between the trunk normal vector and the ground standard normal vector. The calculation formula is:

[0090] Angle = arccos((trunk normal vector × ground normal vector) / (||trunk normal vector|| × ||ground normal vector||));

[0091] Real-time feedback: Compare the angle with a preset perpendicularity threshold to generate a pressing posture correction instruction.

[0092] Compared with the prior art, the present invention and its preferred solutions at least include the following beneficial effects:

[0093] Low-cost non-contact evaluation: The present invention uses monocular vision technology to achieve non-contact action evaluation. The entire process analysis can be completed through an ordinary camera, which breakthroughly solves the high-cost restriction brought by depth sensors such as Kinect in traditional solutions and significantly reduces the equipment deployment threshold.

[0094] High-precision quantitative feedback: Quantitatively analyze biomechanical parameters such as chest compression depth, frequency, and trunk inclination angle through a high-precision motion capture algorithm, realize millimeter-level action error detection, and break through the subjectivity and efficiency limitations of traditional manual evaluation in the judgment of CPR action standardization.

[0095] Heterogeneous illumination robustness: Based on a heterogeneous illumination compensation algorithm with local-global hierarchical optimization, enhance local texture features through adaptive histogram equalization, and combine the Retinex theory to achieve global illumination reconstruction, effectively overcoming the feature loss problem under complex illumination scenarios and improving the stability of key point detection.

[0096] Guarantee of mechanical conduction effectiveness: Innovatively adopt a spatio-temporal constraint 3D reconstruction model to map 2D joint points to three-dimensional space. By establishing the geometric relationship between the trunk plane normal vector and the ground coordinate system, accurately calculate the spatial angle between the arm pressing axis and the ground to ensure that the mechanical conduction of chest compressions conforms to medical operation specifications. Description of the Drawings

[0097] The following further elaborates on the present invention in detail in conjunction with the drawings and specific embodiments:

[0098] Figure 1 This is the schematic diagram for implementing the solution of the embodiment of the present invention. Specific implementation manners

[0099] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are hereinafter given and described in detail as follows:

[0100] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.

[0101] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0102] The system provided by the embodiment of the present invention first obtains the operator's action data through real-time acquisition by a camera or video input, and performs dynamic light compensation through a light optimization model to improve the image quality. Subsequently, the preprocessed frame image is input into a person detection model to generate candidate boxes, and a parameterized non-maximum suppression algorithm is used to screen the optimal detection results. Based on the accurately located person area, the key point detection model extracts the coordinates of 17 core joint points of the human body frame by frame, and synchronously constructs a 2D pose skeleton map for visual feedback. To achieve stereo action tracking, the system maps the planar coordinates to the three-dimensional space through a 2D to 3D key point model, combines the spatio-temporal attention mechanism to deduce the depth information, and significantly improves the pose tracking accuracy. The core evaluation module adopts a multi-modal scoring algorithm to compare the spatio-temporal matching degree between the operator's 3D key point trajectory and the standard CPR action template in real time, calculates the joint motion similarity through the dynamic time warping algorithm, and conducts multi-dimensional evaluation in combination with action indicators such as the straightness of the arm and the verticality of the shoulder to the ground, and clinical indicators such as the pressing depth and frequency. This solution also supports real-time display of the operation score and improvement suggestions, and generates a comprehensive report including action decomposition evaluation after the training ends, specifically points out the optimization directions of each operation link, and forms a closed-loop training system of "acquisition - analysis - feedback - improvement".

[0103] As Figure 1 shown, this embodiment provides a design and working process of an intelligent medical emergency teaching system based on a key point detection algorithm, specifically including the following steps:

[0104] Step S1: Decompose the input training video into frame images, input them into the IAT light detection model, and after local and global adjustments, output the frame images with optimized light.

[0105] Step S2: Input the frame image with optimized light into the OpenPose keypoint detection model. Use the spatial attention mechanism to enhance the localization ability of the keypoint area, use the generated heatmap to perform keypoint localization on the original image, adjust the located image area, and input it into the network again for refined keypoint detection. Combine multi-scale features and structural information to obtain the multi-person keypoint detection result of the image;

[0106] Step S3: Based on the multi-person keypoints generated by OpenPose, use the attention concentration algorithm to exclude the interference of other non-training personnel, track the motion posture of the trainer, output the single-person keypoint frame image of the trainer, and calculate the operation score of the trainer;

[0107] Step S4: Input the single-person keypoint frame image of the trainer into the 2D to 3D keypoint inference algorithm MHFormer. After single-hypothesis temporal modeling, infer multiple possible 3D keypoint coordinates. Combine the cross-hypothesis interaction algorithm for adaptive weighted fusion, and add the human skeleton geometric constraint loss to generate the final optimized 3D human pose sequence;

[0108] Step S5: Based on the 2D and 3D keypoint coordinates, infer the human action posture, give the scores of each action, calculate the total score by weighted calculation, and give the overall evaluation and suggestions in combination with the temporal score.

[0109] In this embodiment, the implementation of step S1 specifically includes the following steps:

[0110] Step S11: Normalize the original input RGB image X to the interval [0,1]. Then inverse map it to the RAW space, and simulate the RAW-RGB domain data through the inverse mapping function F(X). The calculation formula is:

[0111]

[0112] where X is the original input image, γ 0 is the preset Gamma value, and Gauss noise is the simulation module for adding noise;

[0113] Step S12: Perform local illumination adjustment. Adopt 3 cascaded lightweight PEM modules. Each module includes: depthwise separable convolution, which is used to extract local features and reduce the computational amount; lightweight Transformer layer, which captures long-range dependencies through multi-head self-attention. Output the multiplication map M and the addition map A, and calculate the local enhancement result X of the image according to the following formula local :

[0114] X local= M⊙X + A

[0115] Where M is the multiplication graph, A is the addition graph, and X is the original input image. ⊙ is the Hadamard product, which multiplies the corresponding elements of the M and X matrices.

[0116] Step S13: Generate global parameters. Use the dynamic Query mechanism and the multi-head attention mechanism to randomly initialize the learnable Query vector Q, and interact with the Key K and Value V generated by the image features. The formula is as follows:

[0117] MultiHead(Q, K, V) = Concat(head 1 , …, head he )B out

[0118]

[0119] Where Q, K, and V are the Query (query vector), Key (key matrix), and Value (value matrix) respectively, which are derived from the image features and learnable parameters. he is the number of attention heads, indicating that the feature space is divided into he subspaces, and each subspace calculates the attention independently. represents the Query projection matrix of the hi-th attention head, which projects Q into a specific subspace. is the Key projection matrix of the hi-th attention head, which projects K into the corresponding subspace. is the Value projection matrix of the hi-th attention head, which projects V into the corresponding subspace. head hi is the output of the hi-th attention head, which is obtained by calculating the attention of the projected Q, K, and V. Concat is a concatenation function used to concatenate the outputs of he attention heads to form a comprehensive feature. B out is the output projection matrix, which maps the concatenated features to the final global parameter space. D is the feature dimension, which is used to scale the dot product in the attention calculation. The Softmax function is used to calculate the attention weights between the projected Query and Key. K T is the transpose matrix of the Key (key matrix).

[0120] The global parameters are output after the interaction: a 3×3 color matrix c and a Gamma value γ. The Gamma value is constrained to [0.5, 2.0] through the Sigmoid function. The calculation formula is as follows:

[0121] γ = 1.5·σ(w γ ) + 0.5

[0122] Where, w γis the output from the multi-head attention mechanism and serves as the weight coefficient, representing the intermediate features related to the Gamma value. σ is the Sigmoid function used to map w γ to the range (0, 1).

[0123] The color matrix C is initialized orthogonally to ensure numerical stability, and the formula is as follows:

[0124]

[0125] where C is a 3×3 color transformation matrix; C T is the transpose matrix of C; E is a 3×3 identity matrix; is the Frobenius norm, which is used to measure the difference between C T and the identity matrix, representing the orthogonality deviation of C. λ is the regularization coefficient used to control the strength of orthogonal regularization; C ortho is the color matrix after orthogonal initialization.

[0126] Step S14: Finally, perform joint optimization and inference, and optimize the results according to the loss function. The L1 reconstruction loss L rec , supervise the enhanced result and the ground truth, and the calculation formula is as follows:

[0127]

[0128] where is the result output by the image enhancement network. X gt is the real image. |·| 1 represents the L1 norm (Manhattan distance), which is used to calculate the sum of the absolute values of the pixel differences between two images and measure the pixel-level differences.

[0129] The perceptual loss L perc , uses the pre-trained VGG network to extract feature similarity, and the calculation formula is as follows:

[0130]

[0131] where φ(·) is the pre-trained VGG network used to extract the high-level feature representation of the image, such as texture, structure, or semantic information. is the enhanced image generated by the model is the feature vector extracted by the VGG network. φ(X gt ) is the feature vector extracted by the VGG network for the real image X gt . represents the square of the L2 norm (Euclidean distance), which is used to calculate the difference between two feature vectors.

[0132] Finally, high-level task joint training is carried out, sharing gradients with the detection and segmentation network, and optimizing the combined loss. The formula is as follows:

[0133] L total1 = λ 1 L rec + λ 2 L task

[0134] where L task is the detection mAP loss or the segmentation mIoU loss. λ 1 is the weight coefficient of the reconstruction loss L rec , used to adjust its contribution to the total loss. λ 2 is the weight coefficient of the task loss L task , used to adjust its contribution to the total loss. L total1 is the total loss.

[0135] In this embodiment, the implementation of step S2 specifically includes the following steps:

[0136] Step S21: Obtain the image dataset, input it into the OpenPose backbone network, integrate the feature information of different scales through the multi-scale feature extraction strategy to enhance the feature representation, and calculate the key point localization loss L heatmap , and the specific formula is as follows:

[0137]

[0138] where H r is the predicted heat map, TH r is the true heat map, and R is the number of key points.

[0139] Step S22: Use the spatial attention mechanism in OpenPose to enhance the localization ability of the key point area, generate the probability distribution map of the key points through heat map regression, use the generated heat map to perform key point localization on the original image, adjust the located image area to the specified size, and input these areas into the network again for refined key point detection;

[0140] Step S23: Based on the key point mask generated by OpenPose, use the mask to model the geometric relationship between key points, and calculate the structural consistency loss L structure , and the specific formula is as follows:

[0141]

[0142] where p ti and p tj respectively represent the predicted coordinates of the ti-th and tj-th key points. lc ti-tjRepresents a predefined anatomical structure length constraint. Ga(ti) represents the set of key points connected to the key point ti. R is the total number of key points.

[0143] Step S24: Perform iterative training according to the specified training parameters, and update the model parameters by optimizing the combined loss.

[0144] Optimized combined loss L total2 , and its calculation formula is:

[0145] L total2 = L heatmap + λ str L structure + λ reg L reg

[0146] Where L heatmap is the key point heatmap regression loss; λ str is the weight coefficient of the structural consistency loss; L structure is the structural consistency loss; λ reg is the weight coefficient of the regularization loss; l reg is the regularization loss, which prevents the model from overfitting and improves the generalization ability.

[0147] Continuously save the optimal model according to the key point detection accuracy of the validation set, and use the key point heatmap and geometric relationship information output by the final model to make full use of the complementarity of multi-scale features and structural information to obtain the final multi-person key point detection result, and its calculation formula is as follows:

[0148]

[0149] Where Pre final is the final key point prediction result, Pre si is the prediction result of the si-th scale, w si is the weight coefficient of the si-th scale, and AS is the total number of scales.

[0150] In this embodiment, the implementation of step S3 specifically includes the following steps:

[0151] Step S31: Calculate the size of each person's candidate box in the multi-person key point detection result.

[0152] Step S32: Compare and select the key points of the candidate box of the person with the largest aspect ratio closest to 1:1 as the attention. For the width width sci and height height sci of each candidate box sci, its score sc sci is:

[0153]

[0154] The index sci of the finally selected candidate box * is:

[0155]

[0156] where sci * is the index of the finally selected candidate box; sc sci is the score of the sci-th candidate box; is used to return the value of the index i that makes sc sci reach the maximum value.

[0157] Step S33: Calculate the operator score according to the attention key point tracking, and mark the operator key points in the frame image.

[0158] In this embodiment, the implementation of step S4 specifically includes the following steps:

[0159] Step S41: Input the 2D human key point sequence of the frame image, use a fully connected layer to construct a multi-hypothesis generator, and generate multiple initial 3D pose hypotheses in parallel. Perform spatial modeling on each hypothesis through a multi-layer Transformer encoder, and calculate the initial 3D coordinate reconstruction loss, that is, the L1 loss. The calculation formula is as follows:

[0160] Z zi = LayerNorm(M + Transformer zi (M + E spatial ))

[0161] where Z zi represents the feature of the zi-th hypothesis. E spatial represents the spatial position encoding, which is used to enhance the spatial structure information of the 2D key points. Transformer zi represents the zi-th layer Transformer encoder. LayerNorm represents the layer normalization operation. M is the multiplication graph.

[0162] Step S42: Perform single-hypothesis temporal modeling. Specifically, in each hypothesis branch, use the multi-head self-attention mechanism to capture the temporal dependence within the single hypothesis. The calculation formula is as follows:

[0163] Z′ mhi = MH - SA(Z mhi ) + Z mhi , Z mix = Concat(Z′ 1 , Z′ 2 , Z′ 3 )

[0164] {Z″1 , Z″ 2 , Z″ 3}, = split(HM - MLP(Z mix ))

[0165] where Z' mhi is the mhi - th hypothesized feature enhanced by the multi - head self - attention (MH - SA) module. MH - SA, namely the multi - head self - attention mechanism, is used to capture the temporal dependencies within a single hypothesis. Z mix is the mixed feature after concatenating multiple hypothesized features. HM - MLP, namely the hypothesized mixture multi - layer perceptron, is used for cross - hypothesis information interaction. Z” mhi represents the mhi - th hypothesized feature refined by HM - MLP.

[0166] Step S43: Perform cross - hypothesis interaction. Specifically, introduce a cross - hypothesis attention layer to promote information interaction between different hypotheses and enhance pose diversity. Retain multi - scale spatio - temporal features through hierarchical feature transmission.

[0167] Q = Z″ qi , K = Z″ qj , V = Z" qk (qi ≠ qj ≠ qk)

[0168]

[0169] where Q, K, V represent Query, Key, and Value respectively, which are features from different hypotheses. Z″ qi represents the qi - th hypothesized feature refined by HM - MLP. Z″ qj represents the qj - th hypothesized feature refined by HM - MLP. Z″ qk represents the qk - th hypothesized feature refined by HM - MLP. D represents the feature dimension, which is used to scale the dot - product attention scores. MH - CA, namely the multi - head cross - attention mechanism, is used to model the complementary relationships between different hypotheses. K T is the transposed matrix of Key (key matrix).

[0170] Step S44: Perform cross - hypothesis distillation and geometric consistency constraint. Perform adaptive weighted fusion on the prediction results of multiple hypotheses, calculate the cross - hypothesis KL - divergence distillation loss to maintain output consistency. At the same time, add human skeleton geometric constraint losses, such as joint length constraint, angle smoothness constraint, etc., to improve the physical rationality of the prediction results. The relevant error calculation formulas are as follows:

[0171]

[0172] where L MPJPE, i.e., the average joint position error, is used to constrain the global position accuracy. N is the total number of joints. represents the 3D coordinates of the predicted mk-th joint. t mk represents the 3D coordinates of the actual mk-th joint. ||·|| 2 is the L2 norm, i.e., the Euclidean distance, which calculates the distance between two 3D points.

[0173]

[0174] where L length , i.e., the joint length constraint loss, is used to ensure that the joint lengths conform to human priors. ε is the edge set of the human skeleton, representing the connection relationship between joints. (o, u) is a pair of connected joints, where o and u represent the joint numbers at both ends of the bone. |·| is the absolute value operation. represents the 3D coordinates of the predicted o, u-th joints. t o , t u represents the 3D coordinates of the actual o, u-th joints.

[0175]

[0176] where L angle is the joint angle constraint loss, which is used to ensure that the joint angles of the predicted pose conform to human kinematic laws. (ao, ap, aq) is the definition of a joint angle, where ap is the central joint, and ao and aq are the two joints connected to it. is the set of joint angles, representing the human joint angles that need to be constrained, such as the angles of the knee joint or elbow joint. represents the 3D coordinates of the predicted ao, ap, aq-th joints. θ ao-ap-aq is the actual joint angle, used as the ground truth. · is the vector dot product operation. cos -1 () is the inverse cosine function. ||·|| 2 is the L2 norm, which is used to calculate the vector length.

[0177] L geo = L length + L angle

[0178] where L geo is the geometric constraint loss. L length is the joint length constraint loss. L angle is the joint angle constraint loss.

[0179] Step S45: Perform multi-stage joint optimization and multi-hypothesis fusion output. End-to-end train the combination reconstruction loss, distillation loss, and geometric constraint loss, and gradually adjust the multi-hypothesis weights through a learning strategy. In the inference stage, use learnable attention weights to fuse the multi-hypothesis outputs to generate the final optimized 3D human pose sequence. The calculation formula is as follows:

[0180]

[0181] where represents the 3D keypoint prediction result of the yk-th hypothesis. w wk represents the weight of the wk-th hypothesis, calculated through the Softmax function. sc wk represents the confidence score of the wk-th hypothesis. represents the final fused 3D keypoint prediction result.

[0182] In this embodiment, the implementation of step S5 specifically includes the following steps:

[0183] Step S51: According to the keypoint coordinates, judge the operator's action posture and give real-time evaluations and suggestions.

[0184] Step S52: Perform temporal weighted scoring on each action to obtain the final score sc final , and the formula is as follows:

[0185]

[0186] where w fi is the set weight of the fi-th partial score, and sc fi is the score of the fi-th part.

[0187] Step S53: Give training evaluations and propose training improvement suggestions based on the temporal scores and the total score.

[0188] Specifically, the system provided by the embodiments of the present invention uses monocular vision technology to achieve contactless motion assessment, and can complete the full-process analysis through an ordinary camera, which breakthroughly solves the high-cost restriction brought by depth sensors such as Kinect in traditional solutions. The present invention breaks through the limitations of traditional manual assessment in the judgment of the standardization of CPR actions, quantifies and analyzes biomechanical parameters such as chest compression depth, frequency, and torso inclination angle through a high-precision motion capture algorithm, and realizes millimeter-level motion error detection. Based on the heterogeneous illumination compensation algorithm, the present invention adopts a local-global hierarchical optimization strategy. First, the local texture features are enhanced by adaptive histogram equalization, and then the global illumination reconstruction is combined with the Retinex theory, effectively overcoming the feature loss problem in a single optimization mode. The present invention innovatively uses a spatio-temporal constrained 3D reconstruction model to map 2D joint points into three-dimensional space vectors, and accurately calculates the spatial angle between the arm compression axis and the ground by establishing the geometric relationship between the normal vector of the torso plane and the ground coordinate system, ensuring the effectiveness of the mechanical conduction of chest compressions. The intelligent medical emergency teaching plan based on the key point detection algorithm constructed by the present invention can accurately and effectively score CPR training and improve the training effect.

[0189] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is used to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions, specifically used to load and execute one or more instructions in the computer storage medium to implement the above method.

[0190] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above-mentioned method. The storage medium can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electrical, magnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device.

[0191] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those with ordinary skills in the field to which the present invention belongs. The "first", "second", and similar terms used in the present invention do not indicate any order, quantity, or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up", "down", "left", "right", etc. are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0192] As described above, these are only the preferred embodiments of the present invention, and the present invention is not limited to other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

[0193] The present invention is not limited to the above-mentioned optimal implementation manner. Anyone can obtain various other forms of intelligent medical emergency teaching systems based on key point detection under the inspiration of the present invention. All equivalent changes and modifications made within the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. An intelligent medical emergency teaching system based on key point detection, characterized in that: include: Adaptive image enhancement module for local-global hierarchical optimization of input video frames, including: Simulate RAW-RGB data through inverse mapping function and Gaussian noise; Cascading lightweight PEM modules to extract local features and enhance illumination, the PEM module consists of a depthwise separable convolution and a lightweight Transformer; Generate global color matrix and gamma value based on dynamic query mechanism, and optimize end-to-end through orthogonal constraints and joint loss; The key point detection module is used to extract the operator's body key points from the enhanced video frame, including: Spatial attention mechanism enhances keypoint heatmap regression; Optimize key point localization based on structural consistency loss with predefined anatomical length and angle constraints; Filter the target operator's key points according to the aspect ratio of the candidate frame; A three-dimensional posture reconstruction module, used to map the key points to three-dimensional space, includes: Multiple hypothesis generators generate initial 3D pose hypotheses in parallel; The multi-head cross-attention mechanism across hypothesis interactions fuses spatiotemporal features; Optimize the pose sequence based on the geometric constraint loss of joint length error and angle error; Multimodal scoring module, used to quantitatively evaluate operational actions, including: The dynamic time warping algorithm matches the temporal and spatial similarities of the operator's trajectory with the standard template; A multi-dimensional weighted score integrating compression depth, frequency, trunk inclination and arm extension; Generate real-time feedback instructions based on score differences.

2. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: The cascade lightweight PEM module comprises: Depthwise separable convolutional layer to extract local illumination features; Lightweight Transformer layer that models global illumination dependencies through multi-head self-attention mechanism; The feature fusion unit spatially aligns the local features output by the depthwise separable convolutional layer with the global features output by the Transformer layer to generate a multiplication graph M and an addition graph A, where: The multiplication map M is multiplied element-by-element with the original input image X to enhance the local illumination contrast; The addition map A is added element by element with the multiplication result to compensate for the illumination deviation; The final output is the local enhanced image X local ; The depthwise separable convolutional layer extracts local illumination features through a 3×3 convolution kernel.

3. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: In the dynamic query mechanism, the generation of the global Gamma value includes the following steps: Feature extraction: Extract the illumination-sensitive channel w of the multi-head attention intermediate feature vector from the intermediate feature vector output by the multi-head attention mechanism γ ; Nonlinear mapping: w γ Input the Sigmoid function and constrain its value range to the interval (0,1); Dynamic adjustment: Linearly transform the Sigmoid output and map the Gamma value range to [0.5, 2.0]. The formula is: Gamma value = 1.5 × Sigmoid(w γ ) + 0.

5.

4. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: The anatomical structure length constraint is implemented based on the predefined anatomical length, including: Setting predefined anatomical lengths of adjacent key points based on standard human anatomical data; Calculating the square of the difference between the predicted key point spacing and the predefined anatomical length to generate a structural consistency loss; The optimization method of the anatomical structure length constraint comprises the following steps: Definition of adjacent key point pairs: Based on human anatomy priors, a set of key point pairs to be constrained is predefined, including connected joints or bone endpoints; Standard length setting: Set the corresponding standard anatomical length lc for each key point pair (ti, tj) ti-tj ; Distance error calculation: Extract the predicted key point coordinates p ti With p tj , calculate the square of the difference between its Euclidean distance and the standard length; Structural consistency loss: sum the squared differences of all key point pairs, the formula is: Structural consistency loss = Σ(prediction distance - standard anatomical length)²; Model optimization: By minimizing the structural consistency loss, the key point positioning is constrained to conform to the human anatomy.

5. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: The geometric constraint loss optimizes the physical plausibility of the 3D pose reconstruction by the following steps: Joint length constraints: Predefine the medical standard lengths of the connected joint pairs in the human skeleton; Calculating the Euclidean distance between the predicted joint coordinates and obtaining the absolute difference with the medical standard length; Sum the differences of all connected joint pairs to generate the joint length error; Joint angle constraints: According to human kinematics standards, the medical standard angle ranges of key human joints are predefined; Calculate the angle between adjacent bone vectors based on the predicted joint coordinates, and obtain the cosine similarity difference with the medical standard angle range; Sum the squares of the differences of all key joints to generate the joint angle error; Global position constraints: Calculate the square of the Euclidean distance between the predicted joint 3D coordinates and the true coordinates; The squared distances of all joints are summed and averaged to generate the global position error; Multi-loss collaborative optimization: The joint length error, joint angle error and global position error are weightedly summed to drive the model to generate a three-dimensional posture sequence that conforms to human biomechanics.

6. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: The weighted score is generated by the following steps: Multimodal parameter extraction: extract clinical indicators including compression depth and frequency and biomechanical parameters including trunk inclination and arm extension from the three-dimensional posture sequence; Dynamic weight allocation: According to the degree of influence of the parameters on the operational norms, the weights w of each parameter are automatically generated through the pre-trained model. fi ; The pre-trained model optimizes weight parameters through expert scoring annotation in the CPR operation data set; Weighted score calculation: Score each parameter as sc fi With the corresponding weight w fi After multiplication, sum it up and output the final score sc final ; Feedback instruction generation: Based on the final score sc final ; Deviations from standard thresholds generate real-time correction instructions for pressure intensity, frequency or posture.

7. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: In the key point detection module, candidate box screening achieves attention focusing through the following steps: Candidate box aspect ratio calculation: For all detected candidate boxes, calculate the ratio of their width to height; Target operator screening: Select the candidate box with the aspect ratio closest to 1:1 as the target operator. The calculation formula is: Screening score = min(candidate box width, candidate box height) / max(candidate box width, candidate box height); False detection suppression: The interference targets with abnormal aspect ratios are eliminated through the screening score to improve the key point detection accuracy; the aspect ratio screening threshold is set to 0.8-1.

2.

8. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized by: The cross-hypothesis interaction fuses multi-branch features through the following steps: Hypothesis feature extraction: Extract spatiotemporal features z containing joint velocity and temporal displacement from multiple parallel hypothesis branches qi 、z qj 、z qk ; Interactive Attention Construction: The features of different hypothesis branches are input into the multi-head cross attention layer as query, key, and value respectively; Modeling the complementary dependencies between different hypotheses through attention weight calculation; Feature fusion output: The cross-hypothesis attention output is weightedly fused with the original features of each hypothesis branch to generate a physically reasonable 3D pose sequence.

9. The intelligent medical emergency teaching system based on key point detection according to claim 1 is characterized in that: The system achieves contactless 3D motion tracking through a monocular RGB camera, and outputs the angle between the pressing axis and the ground space with millimeter-level accuracy.

10. The intelligent medical emergency teaching system based on key point detection according to claim 9 is characterized in that: The method for calculating the angle between the pressing axis and the ground space includes: Trunk plane construction: Based on the coordinates of the shoulder and hip key points in the three-dimensional posture sequence, the trunk plane equation is fitted by the least squares method; Normal vector generation: extracting a normal vector from the trunk plane equation to represent the spatial direction of the pressing axis; Ground coordinate system mapping: pre-define the standard normal vector direction of the ground coordinate system; Spatial angle calculation: According to the angle between the trunk normal vector and the standard ground normal vector, the vertical deviation between the pressing axis and the ground is determined. The calculation formula is: Angle = arccos(torso normal vector × ground normal vector) / (||torso normal vector|| × ||ground normal vector||); Real-time feedback: the angle is compared with a preset verticality threshold to generate a pressing posture correction instruction.

Citation Information

Cited By

  • Digital human-driven key point generation method and device, and medium

    CN120876883A