Multi-angle CPR training video detection method based on machine vision
By applying machine vision technology in CPR evaluation, combined with YOLOv11 and Kalman-Hungarian algorithms, the LCSG-YOLO and KCHT algorithms are designed, and the existing CPR evaluation depends on sensors is solved, achieving efficient and accurate CPR evaluation.
Patent Information
- Application Number
- CN202510107285.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing CPR skill assessment methods rely on expensive and bulky 3D capture devices or sensors, resulting in inaccurate, inconsistent and inefficient assessments.
A multi-angle CPR training video detection method based on machine vision is proposed, combining the YOLOv11 model and the Kalman-Hungarian matching algorithm, the LCSG-YOLO model and the KCHT matching algorithm are designed to achieve efficient CPR evaluation without sensors.
Through pure machine vision technology, the precise detection of key objectives in cardiopulmonary resuscitation is improved, the accuracy and efficiency of evaluation is reduced, missed detection and false alarms are reduced, and reliable identification of small objects or obstructed objects is ensured.
Smart Images

Figure CN120014518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of CPR evaluation systems, and in particular to a multi-angle CPR training video detection method based on machine vision. Background Art
[0002] Cardiac arrest (CA) is a life-threatening emergency, and the survival rate of out-of-hospital patients is highly dependent on the timely and effective implementation of cardiopulmonary resuscitation (CPR). However, current CPR skill assessment mainly relies on simulators and manual scoring, which often leads to inconsistent, inaccurate and inefficient assessment results.
[0003] In order to improve the accuracy and efficiency of CPR skill assessment, the medical community and the field of artificial intelligence have begun to work together. Existing studies have explored a variety of AI-based methods, including object detection for scoring, motion capture, and even systems like ChatGPT. For example, Zhang et al. developed a binocular vision-based wearable device to monitor wrist movement to assess CPR quality in real time. Similarly, Tang et al. used a 3D motion capture system to assess CPR quality, highlighting the potential of accurate motion data. In addition, Wang et al. applied AI and the ChatGPT-4 model to analyze CPR training videos and achieved automatic scoring comparable to expert evaluation. Despite some progress, many methods still rely on expensive and bulky 3D capture devices or sensors. This study aims to overcome these limitations and propose a sensor-free CPR assessment method.
[0004] The You Only Look Once (OLO) algorithm is widely recognized for its real-time performance and accuracy, making it a popular choice for object detection tasks. The latest version, YOLOv11, optimizes its architecture by reducing the number of parameters, thereby increasing speed while enhancing detection accuracy. This makes YOLO particularly suitable for dynamic environments without the need for additional sensors. In addition, the Kalman-Hungarian (KH) matching algorithm improves target tracking and allocation, further improving detection accuracy and system efficiency.
[0005] In order to get rid of sensor dependence and develop a CPR assessment technology that is economical, efficient, flexible and applicable to actual scenarios, this study proposed a new CPR assessment system based on machine vision. The system combines the YOLO11 model and the KH matching algorithm, designs the LCSG-YOLO model and the KCHT matching algorithm, and formulates a new CPR detection and evaluation standard based on machine vision, providing an innovative solution for economical, efficient and accurate CPR assessment, and laying the foundation for practical application. Summary of the invention
[0006] The technical problem to be solved by the present invention is to overcome the above technical defects and provide a multi-angle CPR training video detection method based on machine vision.
[0007] In order to solve the above problems, the technical solution of the present invention is: multi-angle CPR training video detection based on machine vision, characterized by: including establishing an LCGS-YOLO model, preparatory work, KCHT matching algorithm design and formulating scoring criteria;
[0008] The establishment of the LCGS-YOLO model includes the design of a local conditional shared feature shared layer, a GCSCA module design, and a GCSCA optical module design;
[0009] The preparation work includes the creation of a dedicated data set and the establishment of grayscale axis standards for CPR training videos;
[0010] The KCHT matching algorithm design includes target dynamic prediction, observation update, Hungarian algorithm application and confidence correction mechanism;
[0011] The scoring criteria developed include evaluation of Cohen's Kappa coefficient and classification accuracy.
[0012] Furthermore, the local conditional shared feature sharing layer is designed to enhance the local and global information in the input feature map by integrating multi-scale context features. This module uses adaptive pooling technology to extract global and local context and adopts learnable weights to improve the target focus and detection accuracy at different spatial scales. Figure X After two scales of adaptive pooling;
[0013] C1=AdaptiveAvgPool2d(1)(x), C2=AdaptiveAvgPool2d(2)(x)
[0014] Channel reduction is performed through a 1x1 convolutional layer to generate a shared feature representation F_shared. Then, context information of different scales is fused and normalized through a Sigmoid activation function to generate a channel attention map, F shared =σ(W1·C1+W2·C2);
[0015] The shared feature map is upsampled using bilinear interpolation to make its spatial dimension consistent with the input feature Figure X Match F shared =Interpolate(F shared ,size=x size ,mode=bilinear)
[0016] The input feature map x is multiplied element-by-element with the shared feature map Fshared and weighted by learnable weights to generate the output x of the module. out = x·F shared α;
[0017] The GCSCA (Global Context and Spatial Context Attention) module in the CSCA module design adopts a self-attention mechanism to capture global and spatial context information, thereby enhancing the feature representation ability and robustness of the model;
[0018] In order to improve the computational efficiency and reduce the complexity of the model in the design of the GCSCA light module, the GCSCA module was simplified and enhanced, and the GCSCA Light module was created.
[0019] Furthermore, the creation of the dedicated dataset was done by selecting 40% of the collected videos, a total of 105 clips, and extracting 20 to 30 key frames from each video based on the complexity of the action and key events, such as the start and end of chest compressions and artificial respiration. This process ultimately produced 2,156 annotated frames, which were divided into a training set of 1,946 frames and a validation set of 210 frames;
[0020] The grayscale values of chest compression and artificial respiration corresponding to key labels were extracted from the dataset for establishing the grayscale axis standard of the CPR training video. These values were used to create a 10-level grayscale scale to facilitate the identification and counting of chest compressions and the detection of compression depth.
[0021] Furthermore, the target dynamic prediction is that the KCHT algorithm uses the Kalman filter to dynamically predict the position, speed and other state information of the target to generate a reasonable prediction for each target;
[0022] Observation update: The KC filter updates the target state using the YOLO detection results. By comparing the predicted position with the detected position, the state is adjusted to make the prediction consistent with the actual target state;
[0023] Hungarian algorithm application: Use the Hungarian algorithm for temporal target matching, and associate targets in consecutive frames through optimal matching;
[0024] Confidence correction mechanism: The KCHT algorithm applies weights to correct predictions when confusion or inconsistency occurs.
[0025] Furthermore, the Cohen's Kappa coefficient is a statistical indicator for evaluating the difference between classification prediction accuracy and randomness, and is suitable for evaluating the consistency of two evaluators or classifiers on classification tasks, even if there is randomness in these classification tasks;
[0026]
[0027] The evaluation of classification accuracy represents the percentage of samples that pass classification and fail classification among all samples. The accuracy is calculated by the following formula:
[0028]
[0029] The multi-angle CPR training video detection method based on machine vision is characterized in that the method design includes the following components:
[0030] Small object optimization: The LCSG-YOLO model improves the detection sensitivity of key CPR areas, such as chest compressions and artificial respiration, by improving the convolutional layers and integrating local contextual information. These improvements reduce missed detections and false positives and ensure reliable recognition of small or occluded objects.
[0031] KCHT matching algorithm: The KCHT algorithm combines the Kalman filter with the Hungarian matching algorithm and introduces confidence indicators and timing features to ensure efficient target pairing and tracking in multi-angle environments. The Kalman filter can refine the target trajectory, reduce noise, and improve matching accuracy in dynamic scenes.
[0032] Cardiopulmonary resuscitation frequency detection: Accurately detect the cardiopulmonary resuscitation frequency through time-based average calculation, grayscale and jitter feature analysis, and fast Fourier transform frequency domain analysis. These technologies can accurately identify and calculate the frequency of chest compressions and artificial respiration;
[0033] CPR depth detection: Determines compression depth using signal smoothing, cycle detection, peak-to-valley difference measurement, and average depth analysis. These methods ensure compliance with CPR standards and provide reliable depth assessment.
[0034] Multi-angle fusion: In multi-angle scenes, the dynamic weight allocation mechanism adjusts the weights according to the viewing angle duration and video frame rate, ensuring accurate recognition of CPR actions even with limited angle information;
[0035] CPR training video module: Integrate and analyze results, evaluate indicators such as the number of compressions, frequency and depth, and the system automatically determines whether each indicator meets international CPR standards, providing a comprehensive assessment of operational compliance
[0036] Compared with the existing technology, the advantages of the present invention are: 1. Proposing the LCSG-YOLO model: Through pure machine vision technology, the key goals in cardiopulmonary resuscitation - appropriate frequency of cardiopulmonary compression and artificial respiration are accurately detected, especially in complex dynamic environments, to ensure efficient and accurate evaluation.
[0037] 2. Create a dedicated dataset: A custom dataset for CPR evaluation was built, and the LCSG-YOLO model was trained based on this dataset.
[0038] 3. Improved KCHT matching algorithm: Based on the traditional KH matching algorithm, the target matching process is optimized by combining confidence and timing information, which significantly improves the matching accuracy and robustness of the LCSG-YOLO model.
[0039] 4. Establish scoring criteria: Based on the AHA CPR and ECC guidelines, CPR assessment criteria suitable for automatic scoring were developed, providing a scientific basis for the intelligent assessment system.
[0040] 5. Design a complete CPR detection process: By uploading the video, the system automatically evaluates the CPR operation according to the preset standards and determines whether it meets the standards without relying on sensors. This provides scientific support for CPR quality improvement and training, and promotes the popularization of intelligent evaluation applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a multi-view perspective device placement diagram for cardiopulmonary resuscitation of the present invention.
[0042] Figure 2 It is a multi-angle CPR training diagram of the present invention.
[0043] Figure 3 It is the gray value scale of the present invention.
[0044] Figure 4 It is a multi-angle detection flow chart of the present invention.
[0045] Figure 5 Schematic diagram of the LCSG-YOLO mode of the present invention.
[0046] Figure 6 is the input feature of the present invention Figure X .
[0047] Figure 7 It is the GCSCA module design diagram of the present invention.
[0048] Figure 8 It is the CPR training video process of the present invention.
[0049] Fig. 9 It is the dithered grayscale data graph of the present invention.
[0050] Fig.10 It is a depth detection data diagram of the present invention.
[0051] Fig.11 It is a CPR training video result diagram of the present invention.
[0052] Fig.12 This is a result diagram of the multi-view CPR training video testing technology of the present invention.
[0053] Fig.13 It is the evaluation confusion matrix diagram of the present invention.
[0054] Fig.14 It is a frequency error distribution diagram of the present invention.
[0055] Fig.15 It is the depth detection distribution diagram of the present invention. DETAILED DESCRIPTION
[0056] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings, wherein the same components are represented by the same reference numerals.
[0057] It should be noted that the words "front", "rear", "left", "right", "up" and "down" used in the following description refer to directions in the drawings, and the words "inside" and "outside" refer to directions toward or away from the geometric center of a specific component, respectively.
[0058] In order to make the contents of the present invention more clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0059] like Figure 6 Shown
[0060] Building the LCGS-YOLO model includes the design of local conditional shared feature sharing layer, GCSCA module design, and GCSCA optical module design;
[0061] The local conditional shared feature sharing layer enhances the local and global information in the input feature map by integrating multi-scale context features. This module uses adaptive pooling technology to extract global and local context and adopts learnable weights to improve the target focus and detection accuracy at different spatial scales. The input feature Figure X After two scales of adaptive pooling: global pooling and mid-scale pooling, the global pooling feature map C1 and the mid-scale pooling feature map C2 are obtained respectively.
[0062] C1=AdaptiveAvgPool2d(1)(x), C2=AdaptiveAvgPool2d(2)(x)
[0063] Channel reduction is performed through a 1x1 convolutional layer to generate a shared feature representation F_shared. Then, context information of different scales is fused and normalized through a Sigmoid activation function to generate a channel attention map;
[0064] F shared =σ(W1·C1+W2·C2)
[0065] The shared feature map is upsampled using bilinear interpolation to make its spatial dimension consistent with the input feature Figure X Match;
[0066] F shared =Interpolate(F shared ,size=x size ,mode=bilinear)
[0067] The input feature map x is multiplied element-by-element with the shared feature map Fshared and weighted by learnable weights to generate the output of the module.
[0068] x out = x·F shared ·α.
[0069] like Figure 7 Shown
[0070] The GCSCA (Global Context and Spatial Context Attention) module uses a self-attention mechanism to capture global and spatial context information, thereby enhancing the model's feature representation and robustness. The module can extract global information and spatial details from local areas, improve the ability to identify key features in complex scenes, and thus improve the accuracy of target detection in multi-target and challenging environments;
[0071] The technical design consists of the following components:
[0072] Channel Attention: By adaptively adjusting the weight of each channel, the model is able to highlight the features of key channels and enhance the attention to local regions and important objects, especially in complex backgrounds.
[0073] C att =σ(W2·ReLU(W1·(P avg +P max )))
[0074] X c = x·(C att ·w c )
[0075] Spatial attention points: By weighting spatial features, the model can accurately locate the target area, isolate noise, and improve detection accuracy
[0076]
[0077] x s =x c ·(S att ·w s )
[0078] Global context convolution layer: aggregates global information, enhances the model's global perception ability, helps the model maintain consistency under multiple targets and complex backgrounds, and improves target recognition accuracy.
[0079] x g =x s +P global
[0080] Feature fusion and residual link: feature fusion and residual link strategies are used to ensure effective information propagation, avoid gradient vanishing, and improve the stability and robustness of the model.
[0081] x out =Conv 1x1 (x s +x g )+R(x)
[0082] In order to improve computational efficiency and reduce model complexity, the GCSCA module was simplified and enhanced to create the GCSCA Light module. This module is particularly suitable for the neck part of the object detection network and provides effective feature enhancement under limited computing resources.
[0083] like Figures 1 to 3 Shown
[0084] The preparation work includes the creation of a dedicated dataset and the establishment of grayscale axis standards for CPR training videos;
[0085] Due to the limited publicly accessible CPR datasets, this study collected data from 265 participants in a controlled CPR training environment. Each training room was equipped with advanced CPR teaching equipment and three cameras were set up, located in the upper left corner, bottom and upper right corner to capture CPR operations from multiple angles. The data was collected in four CPR training rooms, and each training room maintained a roughly consistent camera configuration with only minor position adjustments. Since there is no public CPR dataset, we used a video parser to parse and integrate the key frame images extracted from the CPR video, and manually annotated the data, with five labels, as shown in Table 1;
[0086]
[0087] To create a dedicated dataset, 40% of the collected videos, a total of 105 clips, were selected. 20 to 30 key frames were extracted from each video based on the complexity of the action and key events, such as the start and end of chest compressions and artificial respiration. This process ultimately produced 2,156 annotated frames, which were divided into a training set of 1,946 frames and a validation set of 210 frames.
[0088] Grayscale values of chest compression and artificial respiration corresponding to key labels were extracted from the dataset, and these values were used to create a 10-level grayscale scale to facilitate the identification and counting of chest compressions and the detection of compression depth;
[0089] In the action recognition stage, the grayscale value of the target area is compared with the constructed grayscale scale. When the grayscale value falls within the range of 4 to 7 on the scale, the action is classified as effective chest compression or artificial respiration. This method improves the accuracy of action recognition and counting, provides a reliable mechanism for assessing compression depth, and therefore ensures compliance with CPR standards and improves the accuracy of the overall assessment.
[0090] like Figure 8 Shown
[0091] The KCHT matching algorithm combines the Kalman filter with the Hungarian algorithm (KH matching) and introduces a confidence mechanism that uses temporal features to dynamically predict the change of the target position in consecutive video frames. By applying the recursive principle of the Kalman filter, the algorithm fuses the target state estimate (such as position and velocity) with the observed data, and improves the matching accuracy through prediction update and confidence correction.
[0092] The Kalman filter is extended to a Kalman confidence filter (KC filter), which combines label detection and updates the target state equation based on the Kalman prediction, thereby further improving the accuracy of state estimation. The main steps of the algorithm are as follows:
[0093] Target dynamic prediction is that the KCHT algorithm uses the Kalman filter to dynamically predict the position, speed and other state information of the target and generate reasonable predictions for each target;
[0094] Observation update: The KC filter updates the target state using the YOLO detection results. By comparing the predicted position with the detected position, the state is adjusted to make the prediction consistent with the actual target state;
[0095] Hungarian algorithm application: Use the Hungarian algorithm for temporal target matching, and associate targets in consecutive frames through optimal matching;
[0096] Confidence correction mechanism: When confusion or inconsistency occurs, the KCHT algorithm applies weights to correct predictions;
[0097] In the CPR training video module, CPR quality assessment is enhanced by analyzing features such as grayscale value, jitter and duration. The grayscale value reflects the depth of compression and the stability of the action; jitter detection is used to identify non-compliant actions, such as unstable or too fast compression; and duration ensures compliance with CPR standards, which not only confirms the standardization of the operation process, but also ensures the integrity of the operation. This multi-dimensional analysis makes the assessment more detailed and accurate.
[0098] like Figures 9 and 10 Shown
[0099] The effectiveness of cardiopulmonary compression and artificial respiration during cardiopulmonary resuscitation is evaluated through three analyses: time analysis, grayscale change and jitter analysis, and frequency domain analysis.
[0100] First, the time features are detected by calculating the standardized average time of cardiopulmonary compression and artificial respiration. For each label (i.e., cardiopulmonary compression and artificial respiration) in each view, the action frequency is determined by calculating the duration of each label and comparing it with the standard average time;
[0101]
[0102] Ttotal represents the total duration in seconds or minutes, Tavg represents the average duration of a set of standard CPR actions, Ntotal represents the total number of times, and BPM represents the number of actions per minute.
[0103] Grayscale analysis: First, the LCSG-YOLO and KCHT matching algorithms are used to accurately locate the cardiopulmonary compression area and extract the grayscale value of the area. By comparing this value with the pre-set standard grayscale axis, the effectiveness of the compression action can be evaluated. If the detected grayscale value falls within the standard range, the action is considered effective, whether it is compression or artificial respiration; otherwise, the action is considered invalid and excluded from the statistics;
[0104] Jitter analysis: Pressing actions often produce small motion jitters. To assess this, we analyzed the jitter values of a “marker point” located at the center of the pressing area to detect vibrations and jitters generated during the action.
[0105] Fast Fourier transform analysis: In order to further study the frequency characteristics of CPR actions, we use fast Fourier transform to perform frequency domain analysis on the grayscale and jitter signals of cardiopulmonary compression and artificial respiration;
[0106] Time domain signal: First, extract the time domain signal of cardiopulmonary compression and artificial respiration as input signal;
[0107] Frequency domain conversion: Through fast Fourier transform, the time domain signal is converted into a frequency domain signal and its frequency characteristics are calculated;
[0108] Frequency characteristic graph: Based on the results of FFT, a frequency characteristic graph is generated to capture the peaks and valleys in the frequency domain signal, which represent each cycle of cardiopulmonary compression and artificial respiration respectively;
[0109] Cycle analysis: The number of peaks and valleys in the frequency domain can accurately determine the timing and frequency of each cardiopulmonary compression and artificial respiration cycle;
[0110]
[0111] x(n) is the time domain signal, X(f) is the frequency domain signal, f is the frequency band, and N is the length of the signal.
[0112] Depth detection calculates the compression depth by analyzing the amplitude of the jitter signal generated during each compression. The CPR depth detection module determines the depth of each compression based on the jitter signal at the center of the target area;
[0113] Signal smoothing and cycle detection: When detecting compression depth, the jitter signal is smoothed. Then, a peak detection algorithm is used to identify the peaks and valleys within each compression cycle. By locating these peaks and valleys, the start and end points of each compression cycle can be determined, providing key data points for calculating compression depth.
[0114] Compression depth calculation: Compression depth refers to the difference between the peak and valley values within a compression cycle;
[0115] After YOLO11 n recognition and filter and matching algorithm, the recognition confidence of YOLO11 n for the label and the confidence after filter and matching algorithm are extracted;
[0116] Δd i =Peak i -Trough i ,
[0117] Peaki represents the peak value of the i-th compression cycle, and Trougi represents the valley value of the cycle. In this way, the depth of each compression cycle can be accurately calculated;
[0118] The average value of compression depth is calculated as:
[0119] The depth of all compression cycles can be calculated as the average value using the following formula:
[0120]
[0121] n represents the total number of compression cycles, and Indicates the calculated value of the total compression depth
[0122] By averaging the depths of multiple cycles, the effects of noise that may occur in a single cycle can be minimized, resulting in a more stable and accurate compression depth calculation result.
[0123] like Fig.11 Shown
[0124] In order to address challenges such as illumination changes, complex experimental environments, potential occlusions or misaligned viewpoint information, a dynamic weight allocation mechanism is introduced. This mechanism ensures that the contribution of information from different viewpoints is appropriately weighted, thereby improving the accuracy of CPR analysis;
[0125] In cardiopulmonary resuscitation video analysis, qualification determination is based on the multi-perspective dynamic weighted fusion results of the LCSG-YOLO and KCHT matching algorithms, as well as a comprehensive consideration of parameters such as time, frequency, and compression depth. Due to differences in lighting, camera angles, and distances in the experimental setting, this may lead to errors in detecting chest compressions and artificial respiration. In order to mitigate this effect, this study designed a qualification determination process that combines compression frequency and depth and uses machine learning criteria to evaluate CPR training videos. The following is an overview of the process and criteria for determining whether the CPR test has been passed in this study:
[0126] CPR frequency: The compression frequency of each view is calculated by combining the time count weight 0.5 and the jitter grayscale weight 0.5 to obtain the final CPR frequency.
[0127] CPR compression count: Similar to frequency, the number of compressions per view angle is determined by combining the time count weight of 0.5 with the jitter grayscale weight of 0.5 to calculate the final number of compressions;
[0128] Press depth: The press depths from three perspectives are integrated through a dynamic weight distribution mechanism;
[0129] Effective cardiopulmonary compression rate: The compression rate must be between 100 and 120 times per minute;
[0130] Minimum compression depth: The compression depth must be at least 5 cm;
[0131] Compliance Determination: Based on the above criteria, the CPR operation is evaluated to determine whether it meets the established standards.
[0132] like Figures 12 to 14 Shown
[0133] The scoring criteria were developed including the evaluation of Cohen's Kappa coefficient and classification accuracy;
[0134] Cohen's Kappa coefficient is a statistical indicator used to assess the difference between the accuracy of classification predictions and randomness. It is particularly suitable for assessing the degree of agreement between two evaluators or classifiers on classification tasks, even if there is randomness in these classification tasks;
[0135]
[0136] The evaluation of classification accuracy represents the percentage of samples that pass classification and fail classification among all samples. The accuracy is calculated by the following formula:
[0137]
[0138] This study compares the detection performance of LCSG-YOLO with the YOLO series models YOLO11 and YOLOv8, see Table 2-A. The results show that LCSG-YOLO outperforms other models in key indicators, especially when dealing with complex occlusion situations. Although its mAP50 performance is comparable to YOLO11, LCSG-YOLO achieves a slight improvement in mAP50-95, reaching 0.681. In terms of inference speed, LCSG-YOLO runs at a speed of 1.3 milliseconds, see Fig.12 , significantly faster than YOLOv8 and YOLO11, and therefore more suitable for real-time applications. Specifically, LCSG-YOLO shows higher accuracy in detecting cardiopulmonary compression and artificial respiration areas, DC occlusion and DH occlusion, significantly improving the detection capability of complex occlusions and the precise positioning of compression areas. The performance of the enhanced KCHT matching algorithm in all categories has been significantly improved, with the accuracy of each category exceeding 97%. In particular, the accuracy of the key indicators - cardiopulmonary compression and artificial respiration has increased from 97.8% and 90.8% to 98.5% and 98.7%, respectively, as shown in Table 2-B. These improvements have laid a solid foundation for accurately identifying compression frequency and depth, thereby ensuring the quality and accuracy of cardiopulmonary resuscitation operations in complex scenarios;
[0139] Table 2. Technical results of multi-view CPR practical video test Table A. Model comparison table
[0140]
[0141]
[0142] B. LCSG-YOLO+KCHT Results Table
[0143]
[0144] C. Multi-view CPR practical video detection technology performance test table
[0145]
[0146]
[0147] The final data were compared with the data of the CPR teaching auxiliary sensor to evaluate the performance of the CPR training video under the dual-view and three-view configurations, see Table 2-C, Fig.13 and Fig.14 ,The results show that in the regression analysis, the three-view configuration ,performs better than the two-view setting, with the mean square error and ,mean absolute error values of 276.08 and 8.59, respectively.,In addition, the three-view configuration has a higher Kappa coefficient of 0.63, ,accuracy of 0.82, and F1 score of 0.84.,Also, the area under the receiver operating characteristic ,curve ROC-AUC,0.82, also exceeds that of the two-view configuration, 285.42 and 11.14.,Although the recall rate of the two-view configuration is slightly higher, 0.86, ,the three-view configuration is superior in overall performance.
[0148] like Fig.15 Shown
[0149] There is a negative correlation between compression frequency and compression depth. In the low-frequency range (0-100 times / minute), most compressions meet the standard and are more effective, while in the high-frequency range of 120 times / minute and above, the compression depth is reduced, although some compressions still meet the standard. This change is affected by individual differences and proficiency of rescuers, as well as environmental factors such as light and lens. In addition, changes in viewing angles have little effect on compression depth, and there is no significant difference between the dual-view group and the triple-view group. In general, lower compression frequencies tend to produce deeper compressions, while higher frequencies may lead to insufficient compression depth due to time constraints.
[0150] like Figure 4 Shown
[0151] A multi-angle CPR training video detection method based on machine vision. The design of this method includes the following components:
[0152] Small object optimization: The LCSG-YOLO model improves the detection sensitivity of key CPR areas, such as chest compressions and artificial respiration, by improving the convolutional layers and integrating local contextual information. These improvements reduce missed detections and false positives and ensure reliable recognition of small or occluded objects.
[0153] KCHT matching algorithm: The KCHT algorithm combines the Kalman filter with the Hungarian matching algorithm and introduces confidence indicators and timing features to ensure efficient target pairing and tracking in multi-angle environments. The Kalman filter can refine the target trajectory, reduce noise, and improve matching accuracy in dynamic scenes.
[0154] Cardiopulmonary resuscitation frequency detection: Accurately detect the cardiopulmonary resuscitation frequency through time-based average calculation, grayscale and jitter feature analysis, and fast Fourier transform frequency domain analysis. These technologies can accurately identify and calculate the frequency of chest compressions and artificial respiration;
[0155] CPR depth detection: Determines compression depth using signal smoothing, cycle detection, peak-to-valley difference measurement, and average depth analysis. These methods ensure compliance with CPR standards and provide reliable depth assessment.
[0156] Multi-angle fusion: In multi-angle scenes, the dynamic weight allocation mechanism adjusts the weights according to the viewing angle duration and video frame rate, ensuring accurate recognition of CPR actions even with limited angle information;
[0157] CPR training video module: Integrates and analyzes results, evaluates indicators such as the number of compressions, frequency and depth, and the system automatically determines whether each indicator meets international CPR standards, providing a comprehensive assessment of operational compliance.
[0158] By combining the YOLO11 model with the KH matching algorithm, we proposed the LCSG-YOLO model and the KCHT matching algorithm;
[0159] The LCSG-YOLO model is an improvement on YOLO11, which is optimized for key action and target detection in CPR training videos. It has local context sensitivity, enhanced feature extraction, and improved detection accuracy of key CPR areas.
[0160] The KCHT matching algorithm is an improvement on the KH matching algorithm, with features such as key point detection and contextual hierarchical tracking, aiming to track and analyze the motion trajectory in CPR training videos more accurately.
[0161] This approach provides a cost-effective and flexible standard for CPR training videos and assessments, offering a practical and innovative solution for CPR training and assessments.
[0162] The present invention and its embodiments are described above, and such description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if ordinary technicians in the field are inspired by it, without departing from the purpose of the invention, they can design a structure and embodiment similar to the technical solution without creativity, which should belong to the protection scope of the present invention.
Claims
1. Multi-angle CPR training video detection based on machine vision, characterized by: Including building the LCGS-YOLO model, preparation work, KCHT matching algorithm design and developing scoring criteria; The establishment of the LCGS-YOLO model includes the design of a local conditional shared feature shared layer, a GCSCA module design, and a GCSCA optical module design; The preparation work includes the creation of a dedicated data set and the establishment of grayscale axis standards for CPR training videos; The KCHT matching algorithm design includes target dynamic prediction, observation update, Hungarian algorithm application and confidence correction mechanism; The scoring criteria developed include evaluation of Cohen's Kappa coefficient and classification accuracy.
2. The multi-angle CPR training video detection based on machine vision according to claim 1 is characterized in that: The local condition sharing feature sharing layer is designed for the local condition sharing feature sharing layer to enhance the local and global information in the input feature map by integrating multi-scale context features. The module uses adaptive pooling technology to extract global and local contexts and adopts learnable weights to improve the target focus and detection accuracy at different spatial scales. The input feature map X undergoes adaptive pooling at two scales; C1=AdaptiveAvgPool2d(1)(x), C2=AdaptiveAvgPool2d(2)(x) Channel reduction is performed through a 1x1 convolutional layer to generate a shared feature representation F_shared, and then context information of different scales is fused and normalized through a Sigmoid activation function to generate a channel attention map, F shared =σ(W1·C1+W2·C2); Upsample the shared feature map using bilinear interpolation to match its spatial dimensions with the input feature map XF shared =Interpolate(F shared ,size=x size ,mode=bilinear) The input feature map x is multiplied element-by-element with the shared feature map Fshared and weighted by learnable weights to generate the output x of the module. out = x·F shared α; The GCSCA (Global Context and Spatial Context Attention) module in the CSCA module design adopts a self-attention mechanism to capture global and spatial context information, thereby enhancing the feature representation ability and robustness of the model; In order to improve the computational efficiency and reduce the complexity of the model in the design of the GCSCA light module, the GCSCA module was simplified and enhanced, and the GCSCA Light module was created.
3. The multi-angle CPR training video detection based on machine vision according to claim 1 is characterized in that: The creation of the dedicated dataset was done by selecting 40% of the collected videos, a total of 105 clips, and extracting 20 to 30 key frames from each video based on the complexity of the action and key events, such as the start and end of chest compressions and artificial respiration. This process ultimately produced 2,156 annotated frames, which were divided into a training set of 1,946 frames and a validation set of 210 frames; The grayscale values of chest compression and artificial respiration corresponding to key labels were extracted from the dataset for establishing the grayscale axis standard of the CPR training video. These values were used to create a 10-level grayscale scale to facilitate the identification and counting of chest compressions and the detection of compression depth.
4. The multi-angle CPR training video detection based on machine vision according to claim 1, characterized in that: The target dynamic prediction is that the KCHT algorithm uses the Kalman filter to dynamically predict the position, speed and other state information of the target, and generates a reasonable prediction for each target; Observation update: The KC filter updates the target state using the YOLO detection results. By comparing the predicted position with the detected position, the state is adjusted to make the prediction consistent with the actual target state; Hungarian algorithm application: Use the Hungarian algorithm for temporal target matching, and associate targets in consecutive frames through optimal matching; Confidence correction mechanism: The KCHT algorithm applies weights to correct predictions when confusion or inconsistency occurs.
5. The multi-angle CPR training video detection based on machine vision according to claim 1, characterized in that: The Cohen's Kappa coefficient is a statistical indicator for evaluating the difference between classification prediction accuracy and randomness, and is suitable for evaluating the consistency of two evaluators or classifiers on classification tasks, even if there is randomness in these classification tasks; The evaluation of classification accuracy represents the percentage of samples that pass classification and fail classification among all samples. The accuracy is calculated by the following formula:
6. A multi-angle CPR training video detection method based on machine vision, characterized in that: The method design includes the following components: Small object optimization: The LCSG-YOLO model improves the detection sensitivity of key CPR areas, such as chest compressions and artificial respiration, by improving the convolutional layers and integrating local contextual information. These improvements reduce missed detections and false positives and ensure reliable recognition of small or occluded objects. KCHT matching algorithm: The KCHT algorithm combines the Kalman filter with the Hungarian matching algorithm and introduces confidence indicators and timing features to ensure efficient target pairing and tracking in multi-angle environments. The Kalman filter can refine the target trajectory, reduce noise, and improve matching accuracy in dynamic scenes. Cardiopulmonary resuscitation frequency detection: Accurately detect the cardiopulmonary resuscitation frequency through time-based average calculation, grayscale and jitter feature analysis, and fast Fourier transform frequency domain analysis. These technologies can accurately identify and calculate the frequency of chest compressions and artificial respiration; CPR depth detection: Determines compression depth using signal smoothing, cycle detection, peak-to-valley difference measurement, and average depth analysis. These methods ensure compliance with CPR standards and provide reliable depth assessment. Multi-angle fusion: In multi-angle scenes, the dynamic weight allocation mechanism adjusts the weights according to the viewing angle duration and video frame rate, ensuring accurate recognition of CPR actions even with limited angle information; CPR training video module: Integrates and analyzes results, evaluates indicators such as the number of compressions, frequency and depth, and the system automatically determines whether each indicator meets international CPR standards, providing a comprehensive assessment of operational compliance.
Citation Information
Patent Citations
Nail fold leukocyte monitoring method, device, equipment and medium
CN116385447A
Construction method and application of post-cardiac-arrest brain injury adverse nerve prognosis prediction model
CN118762848A
Systems, devices, and methods for vital sign monitoring
US20230196567A1