Clinical skill training level evaluation method and device, equipment and storage medium
By collecting multimodal data and using dynamic spatiotemporal map attention network model for evaluation, the subjectivity and time-consuming problems of evaluation in traditional clinical skills training are solved, and efficient and accurate training level assessment and real-time feedback are achieved.
Patent Information
- Application Number
- CN202510701130.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-08-26
AI Technical Summary
In traditional clinical skills training, there are subjective biases caused by experience-dependent experience, time-consuming and labor-intensive, lack of real-time feedback, and the inability of a single data source to fully reflect operational details.
Multimodal data (video, audio, sensor data) is collected, pre-processed and feature extraction is performed through the multimodal large model of the dynamic spatio-temporal graph attention network, and combined with speech synthesis technology and cognitive map analysis to generate a real-time evaluation report.
Improve the accuracy and efficiency of evaluation, provide personalized real-time feedback, reduce labor costs, and improve training quality and efficiency.
Smart Images

Figure CN120543032A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of clinical medical education technology, and in particular to a method, apparatus, device and storage medium for evaluating clinical skills training level. Background Art
[0002] Currently, skills training and assessment (i.e., clinical skills training level assessment) are key components of clinical medical education. However, traditional skills training primarily relies on manual guidance, and the assessment of clinical skills training levels presents the following challenges: manual assessment relies on the instructor's experience and is subject to subjective bias; manual assessment is time-consuming and labor-intensive, making it difficult to cover a large number of trainees; the lack of a real-time feedback mechanism prevents trainees from receiving timely improvement suggestions; and current automated assessment systems typically rely on a single data source and fail to fully reflect the details of trainees' clinical skills training. Summary of the Invention
[0003] In view of this, the purpose of this application is to provide a clinical skills training level assessment method, device, equipment and storage medium, which can improve the accuracy of clinical skills training level assessment and enhance the efficiency of assessment. The specific scheme is as follows:
[0004] In a first aspect, the present application discloses a method for evaluating clinical skills training level, comprising:
[0005] Collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data;
[0006] Preprocessing the current multimodal data to obtain preprocessed data, and extracting features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result;
[0007] Inputting the operation feature extraction results into a post-training skill training level assessment model to assess the clinical skill operation level of the target trainee based on the operation feature extraction results to obtain a skill training level assessment result;
[0008] Among them, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data on clinical skill operations of different trainees during historical clinical skills training.
[0009] Optionally, the collecting of multimodal data on clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data includes:
[0010] According to the preset time synchronization mechanism, the video data, audio data and sensor data of the clinical skill operation of the target trainees in the current clinical skill training process are synchronously collected to obtain the current multimodal data;
[0011] The video data is captured by an event camera; the audio data is collected by a multi-physics field audio acquisition device based on a beamforming matrix; and the sensor data is collected by a sensor device including a pressure sensor and an acceleration sensor.
[0012] Optionally, preprocessing the current multimodal data to obtain preprocessed data includes:
[0013] Performing filtering and noise reduction on the audio data in the current multimodal data to obtain audio processed data;
[0014] Filtering is performed on the sensor data in the current multimodal data to obtain sensor-processed data.
[0015] Optionally, extracting features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result includes:
[0016] Detecting a key operation object in current video data of the current multimodal data using an object detection algorithm to obtain current operation object detection information;
[0017] Using a posture estimation algorithm to estimate the operation posture of the target student in the current video data to extract position and angle information of key joints of the target student to obtain operation posture information;
[0018] Using a dynamic spatiotemporal attention fusion network pre-created based on a dynamic spatiotemporal attention mechanism, feature extraction is performed on each video frame in the current video data to extract features that can characterize the target student's operation action, thereby obtaining a video feature extraction result;
[0019] Extracting features that can reflect the target student's voice instructions and operation status from the audio processed data to obtain an audio feature extraction result;
[0020] Features that can reflect the physical characteristics of the target trainee's operating actions in the sensor-processed data are extracted to obtain sensor feature extraction results.
[0021] Optionally, the clinical skills training level assessment method further includes:
[0022] Collect multimodal data on clinical skill operations of different trainees during historical clinical skill training, as well as corresponding operation evaluation information and operation suggestion information, to obtain historical multimodal data;
[0023] The historical multimodal data is input into a multimodal large model based on a dynamic spatiotemporal graph attention network using a supervised learning method for model training to obtain the post-training skill training level evaluation model.
[0024] Optionally, the step of inputting the operation feature extraction result into a post-training skill training level assessment model to assess the target trainee's clinical skill operation level based on the operation feature extraction result to obtain a skill training level assessment result includes:
[0025] Performing feature fusion on the video feature extraction result, the operation posture information, and the current operation object detection information to obtain fused video features;
[0026] The fused features, the audio feature extraction results, and the sensor feature extraction results are added to a post-training skill training level assessment model to semantically align the fused video features, the audio feature extraction results, and the sensor feature extraction results to obtain corresponding aligned video features, aligned audio features, and aligned sensor features, and feature fusion is performed on the aligned video features, the aligned audio features, and the aligned sensor features to obtain fused features, and then the level of clinical skill operation of the target trainee is assessed based on the fused features to obtain a skill training level assessment result;
[0027] Accordingly, after evaluating the clinical skill operation level of the target trainee based on the operation feature extraction result and obtaining the skill training level evaluation result, the method further includes:
[0028] Generate operation evaluation information and operation suggestion information for the target trainee through the post-training skill training level evaluation model, and obtain skill training level evaluation results, current operation evaluation results and current operation suggestions;
[0029] The current operation suggestion is output in voice using speech synthesis technology, and the projection direction of the output voice is adjusted based on the current head posture of the target trainee using a three-dimensional spatial sound field distribution model based on the NeRF architecture to guide the clinical skill operation of the target trainee in real time.
[0030] Optionally, the clinical skills training level assessment method further includes:
[0031] Using a cognitive map analysis engine and based on the skill training level assessment result, the current operation evaluation result, and the current operation suggestion, a clinical skill training level assessment report for the target trainee is created, and the clinical skill training level assessment report is visually displayed;
[0032] The clinical skills training level assessment report is encrypted using an encryption algorithm to obtain an encrypted assessment report, and the encrypted assessment report is stored in a preset storage space.
[0033] In a second aspect, the present application discloses a clinical skills training level assessment device, comprising:
[0034] The data acquisition module is used to collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data;
[0035] A preprocessing module, configured to preprocess the current multimodal data to obtain preprocessed data;
[0036] A feature extraction module is used to extract features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result;
[0037] A training level assessment module is used to input the operation feature extraction results into a post-training skill training level assessment model to assess the clinical skill operation level of the target trainee based on the operation feature extraction results to obtain a skill training level assessment result;
[0038] Among them, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data on clinical skill operations of different trainees during historical clinical skills training.
[0039] In a third aspect, the present application discloses an electronic device comprising a processor and a memory; wherein, when the processor executes a computer program stored in the memory, the aforementioned clinical skills training level assessment method is implemented.
[0040] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned clinical skills training level assessment method is implemented.
[0041] It can be seen that the present application first collects multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data, and then preprocesses the current multimodal data to obtain preprocessed data, and extracts features related to the clinical skill operations of the target trainees in the preprocessed data to obtain operation feature extraction results; then, the operation feature extraction results are input into the post-training skill training level evaluation model to evaluate the level of clinical skill operations of the target trainees based on the operation feature extraction results to obtain skill training level evaluation results; wherein, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data of clinical skill operations of different trainees during historical clinical skills training. This application collects multimodal data of trainees' clinical skills operations during clinical skills training, and uses a multimodal large model based on a dynamic spatiotemporal graph attention network to evaluate the level of trainees' clinical skills operations. Compared with a single data source, this application can improve the accuracy of clinical skills training level evaluation by collecting multimodal clinical skills operation data of trainees during clinical skills training. In addition, since the dynamic spatiotemporal graph attention network can improve the correlation and time series alignment accuracy between cross-modal features, the use of a multimodal large model based on a dynamic spatiotemporal graph attention network to evaluate the skills training level can further improve the accuracy of the evaluation. At the same time, compared with manual evaluation methods, it saves labor costs and time costs, improves the efficiency of evaluation, and thus improves the efficiency and quality of clinical skills training. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0043] Figure 1 A flow chart of a clinical skills training level assessment method disclosed in this application;
[0044] Figure 2 A flowchart of a specific clinical skills training level assessment method disclosed in this application;
[0045] Figure 3 This is a block diagram of a specific clinical skills training level assessment system disclosed in this application;
[0046] Figure 4This is a schematic structural diagram of a clinical skills training level assessment device disclosed in this application;
[0047] Figure 5 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0049] This application embodiment discloses a method for evaluating clinical skills training level. Figure 1 As shown, the method includes:
[0050] Step S11: Collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data.
[0051] In this embodiment, when it is necessary to evaluate the training level of trainees in a clinical skills training process (such as a minimally invasive surgery skills training process), the multimodal data of the clinical skills operations of the target trainees in the current clinical skills training process can be collected first to obtain the current multimodal data of the target trainees; wherein the current multimodal data can be data obtained after being collected by multiple collection devices such as cameras, microphones and sensors.
[0052] Specifically, the multimodal data of the clinical skill operation of the target trainee during the current clinical skill training process is collected to obtain the current multimodal data, which may include: synchronously collecting the video data, audio data and sensor data of the clinical skill operation of the target trainee during the current clinical skill training process according to a preset time synchronization mechanism to obtain the current multimodal data; the video data is obtained by shooting with an event camera; the audio data is obtained by collecting with a multi-physics field audio acquisition device based on a beamforming matrix; the sensor data is obtained by collecting with a sensor device including a pressure sensor and an acceleration sensor. It should be pointed out that when the present application collects the multimodal operation data of the target trainee during the clinical skill training process through multiple acquisition devices such as cameras, microphones and sensors, it is necessary to ensure that the data of different modalities are time-synchronized during the collection. Specifically, the collection can be carried out according to the preset time synchronization mechanism, such as matching the acquisition frequency of the sensor data with the acquisition frequency of the video and audio data. Specifically, clinical skill performance data of target trainees during the current clinical skills training process can be collected via video, audio, and sensor acquisition devices. Then, the captured video, audio, and sensor data are time-aligned. To achieve millisecond-level alignment of multimodal data, dynamic interpolation compensation can be employed, combined with a pre-built generator-discriminator network. This dynamic interpolation compensation employs B-spline interpolation combined with an LSTM (Long Short-Term Memory) prediction model. This approach enables continuous timeline mapping at heterogeneous sensor sampling rates (e.g., 10Hz-10kHz). The generator-discriminator network uses adversarial learning to eliminate cross-modal temporal errors, reducing the cumulative error to <1ms over a short period of continuous acquisition (e.g., one hour). Temporal alignment across modal data ensures temporal consistency, accurately reflecting dynamic changes during the procedure.
[0053] Furthermore, different modality acquisition devices should be selected based on specific application requirements. For example, for video data acquisition, a high-resolution, high-frame-rate camera can be used to ensure that every detail of the trainee's movements can be clearly captured. Furthermore, the camera must have excellent low-light performance to accommodate training scenarios in various environments. In one specific embodiment, the video data can be captured by an event camera and stored in a compressed format (such as H.264, H.265, etc.), thereby balancing data quality and storage space requirements.
[0054] For audio data acquisition devices, high-fidelity microphones can be used to accurately record the trainee's voice commands and ambient sounds. Furthermore, the microphones must have excellent noise reduction performance to reduce background noise interference. Specifically, the audio data can be collected using a multi-physics audio acquisition device based on a beamforming matrix. Furthermore, the audio data can be stored in audio encoding formats such as MP3 (MPEG-1 Audio Layer III, a lossy compressed audio format) or WAV (Waveform Audio File Format) to ensure a reasonable balance between audio quality and data size. In one specific embodiment, the audio data can be collected using a multi-physics audio acquisition device based on a beamforming matrix. For example, audio data collected by a 64-channel MEMS (Micro Electro Mechanical System) microphone can be used to focus on the target trainee's voice commands during clinical skills training. Furthermore, a generative adversarial network (GAN) can be deployed in the microphone device to predict the optimal noise reduction mask in real time, thereby effectively suppressing ambient noise and retaining the trainee's valid voice information. In addition, the Wave encoding format (uncompressed pulse code modulation format) can also be used to synchronously record the sound pressure amplitude and phase information, thereby supporting three-dimensional sound field reconstruction.
[0055] For sensor data acquisition equipment, appropriate sensors can be selected based on the specific clinical skill requirements, such as pressure sensors for measuring pressure and accelerometers for detecting the acceleration of the operating action. Furthermore, the collected sensor data can be recorded in a time series format, with each time point corresponding to the measurement value of one or more sensors.
[0056] Step S12: preprocessing the current multimodal data to obtain preprocessed data, and extracting features related to the clinical skill operation of the target trainee in the preprocessed data to obtain an operation feature extraction result.
[0057] In this embodiment, after collecting the current multimodal data of the clinical skill operations of the target trainees during the current clinical skill training process, further, corresponding preprocessing operations (such as filtering, noise reduction, etc.) can be performed on the above-mentioned current multimodal data, and then the features related to the clinical skill operations of the target trainees in the preprocessed data are extracted to obtain corresponding feature extraction results.
[0058] In this embodiment, preprocessing the current multimodal data to obtain preprocessed data may specifically include filtering and performing noise reduction processing on audio data within the current multimodal data to obtain audio-processed data; and filtering and performing noise reduction processing on sensor data within the current multimodal data to obtain sensor-processed data. In this embodiment, the audio data within the current multimodal data may be filtered and noise-reduced using audio signal processing techniques to remove background noise and interference signals. The sensor data within the current multimodal data may also be filtered (e.g., low-pass filtering, high-pass filtering, etc.) to remove noise and interference signals and improve data accuracy.
[0059] In this embodiment, the features related to the clinical skill operation of the target trainee in the pre-processed data are extracted to obtain the operation feature extraction result, which can specifically include: using a target detection algorithm to detect the key operation object in the current video data of the current multimodal data to obtain current operation object detection information; using a posture estimation algorithm to estimate the operation posture of the target trainee in the current video data to extract the position and angle information of the key joints of the target trainee to obtain operation posture information; using a dynamic spatiotemporal attention fusion network pre-created based on a dynamic spatiotemporal attention mechanism to extract features of each video frame in the current video data to extract features that can characterize the operation actions of the target trainee to obtain video feature extraction results; extracting features that can reflect the voice instructions and operation status of the target trainee in the audio processed data to obtain audio feature extraction results; extracting features that can reflect the physical characteristics of the operation actions of the target trainee in the sensor processed data to obtain sensor feature extraction results. In this embodiment, after preprocessing the current multimodal data, feature extraction can be performed on the preprocessed data to obtain features related to the training process and the target trainee, such as key action features, features that can reflect the trainee's voice instructions and operation status, features that can reflect the physical characteristics of the operation action, etc. For details, see Figure 2As shown, for the current video data in the current multimodal data, the target detection algorithm (such as YOLO, Faster R-CNN and other algorithms) can be used to detect the key operation objects in the current video data, such as human body parts, medical equipment, etc., to obtain the current operation object detection information; then, the posture estimation algorithm (such as OpenPose, HRNet and other algorithms) is used to estimate the operation posture of the target student in the current video data to extract the position and angle information of the key joints of the target student to obtain the operation posture information; in addition, the dynamic spatiotemporal attention fusion network pre-created based on the dynamic spatiotemporal attention mechanism can be used to perform feature extraction on each video frame in the current video data to extract features that can characterize the operation actions of the target student and obtain the corresponding video feature extraction results.
[0060] It is important to note that the dynamic spatiotemporal attention fusion network, built based on the dynamic spatiotemporal attention mechanism, can also be combined with a feature pyramid network based on multi-scale fusion to achieve more comprehensive feature extraction from video data. The backbone network includes deformable 3D convolution kernels for accurately capturing non-rigid deformation features of medical device rotation trajectories, an HRNet network for lightweight, real-time motion trajectory capture (with an optical flow-guided attention module inserted in the third stage of the HRNet network), and an adversarial feature distillation mechanism to enhance feature expressiveness and make the network more robust to occlusion. The multi-scale fusion feature pyramid network specifically uses a five-layer wavelet decomposition to extract tremor energy features at different scales and utilizes a differentiable feature selection gate to remove redundant features. Furthermore, to emphasize the characteristics of human behavior and specific medical devices in the video data, an object detection model is used to detect medical devices related to skill training in the video. A skeleton point detection model is also used to extract the skeleton points of the target trainees in the video. Furthermore, all the extracted information can be fused, and the fused features can be vectorized and used as the input of the model in step S13.
[0061] Specifically, the dynamic spatiotemporal attention fusion network based on the dynamic spatiotemporal attention mechanism can be constructed based on a deformable 3D convolution kernel, wherein the deformable 3D convolution kernel can adopt a parameterized offset generation mechanism and introduce a dynamic offset matrix Δp into the conventional 3D convolution layer. The deformable 3D convolution kernel can be expressed by the following formula:
[0062] ;
[0063] Where p_n is the convolution kernel sampling point, and Δp_n can be predicted in real time using a lightweight fully connected network. y(p) represents the eigenvalue output at position p after the deformable 3D convolution kernel is applied, representing the feature obtained by aggregating information through the convolution operation. p is a parameter defining the location of the output feature y(p), which is used to determine its specific position within the feature map. x represents the input feature map, and the eigenvalues obtained from the input feature map x serve as the input data source for the convolution operation. The configured dynamic convolution kernel size is 5×5×3 (H×W×T), supporting the modeling of non-rigid deformations of medical device rotational trajectories.
[0064] In addition, the dynamic spatiotemporal attention fusion network can also embed the optical flow attention mechanism and the adversarial feature distillation mechanism. The optical flow attention mechanism refers to inserting the optical flow guided attention module in the third stage of the HRNet network, which can be specifically expressed as:
[0065] ;
[0066] Where, is the optical flow field feature of adjacent frames, is the RGB (Red, green, blue, three primary colors) feature map, Represents channel splicing. By adopting a lightweight optical flow estimation network, real-time motion trajectory analysis at 30 frames per second can be achieved. Conv1×1 represents a 1×1 convolution (i.e., 1-by-1 convolution), which refers to a convolution operation with a convolution kernel size of 1×1. In neural networks, this operation does not involve feature fusion in the spatial dimensions (height and width) but mainly acts on the channel dimension. A represents the output of the optical flow-guided attention module. It is an attention weight matrix, which is obtained by performing a series of operations on the input optical flow field features and RGB feature maps. It is used to measure the importance of different positions in the feature map.
[0067] In addition, the adversarial feature distillation mechanism requires the pre-construction of a dual-path feature extraction architecture, which includes a main path (for generating operational feature vectors) and a discriminative path (for generating 128-dimensional attention masks through the discriminator). The loss function L corresponding to the adversarial feature distillation mechanism is:
[0068] ;
[0069] Where, To combat losses, is the feature reconstruction loss, is the offset regularization term, represents the weight coefficient of the adversarial loss, represents the feature loss weight coefficient, Represents the weight coefficient of the offset regularization term.
[0070] Specifically, the feature pyramid network based on multi-scale fusion includes a complex Morlet (a single-frequency complex sine modulated Gaussian wave) wavelet transform layer and a differentiable feature selection gate. Among them, a complex domain wavelet filter bank is pre-constructed in the complex Morlet wavelet transform layer, which can be specifically expressed as:
[0071] ;
[0072] Where ω_0 = 6 rad / s, covering the 0.5-4 Hz micro-motion characteristic frequency band; the feature pyramid network uses 5-layer wavelet decomposition to extract tremor energy features of different scales. Represents the complex Morlet wavelet function, a function of time t. It is used to construct complex-domain wavelet filter banks. In multiscale fusion feature pyramid networks, it performs the wavelet transform function to extract signal features. t is the time variable, which, in the complex Morlet wavelet function, characterizes the characteristics of the signal at different times.
[0073] Specifically, the differentiable feature selection gate can be dynamically generated through feature fusion weights. The specific formula is:
[0074] ;
[0075] Where, is a Sigmoid function, and MLP is a 3-layer fully connected network. In addition, an adaptive threshold mechanism can be set: when the frequency domain feature energy is lower than θ=0.15, the redundant components are automatically shielded. Represents wavelet transform features, which here means processing audio data to extract audio-related features. For example, it can capture detailed information of audio signals at different scales, which helps analyze audio characteristics such as pitch and timbre. Represents the Fourier transform feature. In audio data processing, Fourier transform can help analyze the frequency distribution of audio signals. For example, different frequency components in voice commands correspond to different pronunciations and semantic information.
[0076] Feature extraction can be performed on the audio processed data in the current multimodal data. Specifically, features that can reflect the voice instructions and operation status of the target student, such as pitch, timbre, speaking speed, etc., are extracted to obtain audio feature extraction results.
[0077] Feature extraction can be performed on the sensor-processed data in the current multimodal data. Specifically, features that can reflect the physical characteristics of the target trainees' operating actions, such as the changing trend of pressure values, peak acceleration values, etc., are extracted to obtain sensor feature extraction results.
[0078] Furthermore, the multimodal feature data after feature extraction can be vectorized separately to obtain vectorized features, and then these features can be used as the input part of the subsequent model processing. For example, the feature extraction results corresponding to each modal data are encoded, and the encoded data is used as the input of the model in step S13. For example, for the audio feature extraction results, in order to facilitate subsequent semantic analysis, the voice content in the audio can be converted into text using speech recognition technology (such as DeepSpeech, Wav2Vec, etc.) to obtain the converted audio text, and then feature extraction is performed on the converted audio text, and the audio feature extraction results are encoded and used as the input of the model in step S13. For example, for the sensor feature extraction results, the sensor value and corresponding time in the sensor feature extraction results can be encoded. For example, the pressure value and time output by the pressure sensor are encoded to form a vector for use as the input of the model in step S13.
[0079] Step S13: Input the operation feature extraction result into the post-training skill training level evaluation model to evaluate the level of clinical skill operation of the target trainee based on the operation feature extraction result to obtain a skill training level evaluation result; wherein, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data of clinical skill operations for different trainees during historical clinical skills training.
[0080] In this embodiment, after extracting the features related to the clinical skill operation of the target trainee in the preprocessed data, the extracted operation feature extraction results can be input into the skill training level evaluation model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, so as to evaluate the level of clinical skill operation of the target trainee based on the above-mentioned operation feature extraction results and obtain corresponding skill training level evaluation results; wherein, the historical multimodal data includes multimodal data of clinical skill operations for different trainees during the historical clinical skills training process.
[0081] Specifically, the process of acquiring a skill training level assessment model includes first collecting multimodal data on clinical skill operations performed by different trainees during historical clinical skill training, as well as corresponding operation evaluation information and operation suggestion information, to obtain historical multimodal data. This historical multimodal data is then input into a large multimodal model based on a dynamic spatiotemporal graph attention network using a supervised learning approach for model training, thereby obtaining the trained skill training level assessment model. Specifically, multimodal data on clinical skill operations performed by different trainees during historical clinical skill training, such as video data, audio data, and sensor data from different trainees' operations, is first collected. The corresponding operation evaluation and operation suggestion information is also collected. All of this collected data is then input into a large multimodal model based on a dynamic spatiotemporal graph attention network as a dataset for model training. Furthermore, during model training, a corresponding loss function, such as a cross-entropy loss function or a mean squared error loss function, can be selected based on actual application requirements to improve the model's prediction accuracy.
[0082] In addition, in the process of training the multimodal large model based on the dynamic spatiotemporal graph attention network, it is necessary to perform labeling operations on the collected historical multimodal data. Specifically, all clinical skills can be broken down into several operation steps, and then each step can be labeled according to the scoring rules of the actual examination. The specific labeling content is which modal data determines the evaluation of the operation. For example: the chest compression operation of cardiopulmonary resuscitation needs to be scored by combining video data and sensor data to score the trainees; the confirmation of environmental safety operation of cardiopulmonary resuscitation needs to be scored by combining video data and audio data. The reason for labeling is to prevent irrelevant modal data from interfering with normal scoring when evaluating specific operations. Furthermore, after determining the current operation, it is necessary to select the required modal data according to the labeling and send it to the operation scoring module to score the trainee's operation.
[0083] It should be pointed out that each student's score, including deductions, will be recorded in detail in the database. Before generating an evaluation report, the corresponding student's historical operation test data will be reviewed first, and combined with the current test data, a personalized evaluation report will be generated to help students improve their training level. For example: in the last test, when the student was performing cardiopulmonary resuscitation chest compressions, points were deducted because the compression depth was not enough. In this test, the student was deducted for the same reason, but according to the sensor data records, the student's compression depth this time has greatly improved compared to the previous time. In this case, the following words will be generated: "Although this chest compression operation still deducted points due to insufficient compression depth, it is a great improvement compared to the previous time. Keep working hard next time." Compared with simple result feedback, it can give students more specific positive feedback and point out the direction for improvement.
[0084] Specifically, for multimodal data, multiple sub-models can be used to train the models separately. For example, when the dataset includes video data, audio data, and sensor data, they can be input into the corresponding three multimodal large models based on the dynamic spatiotemporal graph attention network for training, that is, the three models are trained separately. In addition, to improve the accuracy and robustness of model evaluation, new training set data can be collected regularly to continuously update and optimize the model. At the same time, the model can be adjusted and improved based on feedback from students and teachers to make it more in line with actual needs, thereby ensuring the reliability and effectiveness of the model.
[0085] In this embodiment, the skill training level assessment model of the dynamic spatiotemporal graph attention network can not only achieve high-precision alignment of multimodal data at the feature layer, but also use adversarial cross-modal contrastive learning to highly distinguish target detection operations from easily confused operations. Specifically, when performing skill training level assessment, the dynamic spatiotemporal graph attention network constructs a spatiotemporal graph, specifically mapping video key frames, text clips, and sensor data into heterogeneous temporal graph nodes, and constructing a dynamic adjacency matrix. The specific formula is:
[0086] ;
[0087] Where, is the time difference between modalities, i and j are used to identify the index of the graph node. In the constructed spatiotemporal graph, different nodes represent different information (such as heterogeneous time graph nodes mapped by video key frames, text fragments, sensor data, etc.). i and j can distinguish different nodes so as to calculate the relationship between nodes. and are feature vectors, representing the feature representations corresponding to nodes i and j in the spatiotemporal graph respectively. These features can be extracted from multimodal information such as video data, text data, and sensor data to characterize the characteristic attributes of the information represented by the nodes. and It is a motion vector, which reflects the motion characteristics of the information corresponding to node i and node j in the space-time graph. For example, in video data, it can represent the object's motion direction, speed and other information. Represents an element in the dynamic adjacency matrix, which is used to describe the connection relationship or association degree between node i and node j at time t. Its value is obtained through specific calculation, and the size of the value reflects the closeness of the relationship between the two nodes.
[0088] In addition, the dynamic spatiotemporal graph attention network adopts a multi-head dynamic attention mechanism, which has four-dimensional attention heads, focusing on spatiotemporal motion (video), spectral mutation (audio), mechanical changes (sensors) and cross-modal correlation features, and each head dimension , using a dynamic head weight adjustment mechanism. Compared to the Transformer architecture, this mechanism can improve the accuracy of data alignment between different modalities. Experimental results show that alignment accuracy can be improved by 37% and the mean absolute error (MAE) is reduced from 8.3ms to 5.2ms. Furthermore, the correlation between cross-modal features is improved, for example, to 0.91 (Pearson coefficient).
[0089] In a specific embodiment, a triplet contrast loss function can be used during model training. , the specific formula is:
[0090] ;
[0091] Where, the temperature coefficient , s represents the similarity measurement function, which is used to measure the similarity between two feature vectors, such as cosine similarity; and Represents the feature vector of the positive sample, i and j are used to distinguish different positive sample features; Represents the feature vector of negative samples, and k is used to distinguish different negative sample features. Positive samples come from the same operation stage, and negative samples are extracted across modalities.
[0092] In this embodiment, the operation feature extraction result is input into the post-training skill training level assessment model to evaluate the target trainee's clinical skill operation level based on the operation feature extraction result to obtain a skill training level assessment result. Specifically, it may include: performing feature fusion on the video feature extraction result, the operation posture information and the current operation object detection information to obtain fused video features; inputting the fused features, the audio feature extraction result and the sensor feature extraction result into the post-training skill training level assessment model to semantically align the fused video features, the audio feature extraction result and the sensor feature extraction result to obtain corresponding aligned video features, aligned audio features and aligned sensor features, and performing feature fusion on the aligned video features, the aligned audio features and the aligned sensor features to obtain fused features, and then evaluating the target trainee's clinical skill operation level based on the fused features to obtain a skill training level assessment result. In this embodiment, see Figure 2As shown, the video feature extraction results obtained after feature extraction of each frame image in the video data, the operation posture information obtained after posture recognition of the target trainee, and the current operation object detection information obtained after identification of key operation objects such as medical devices can be fused to obtain fused video features; then, the fused video features are semantically aligned with the audio feature extraction results obtained after noise reduction, text conversion and text feature extraction, as well as the sensor feature extraction results after filtering and feature extraction, and then the aligned video features, aligned audio features and aligned sensor features are feature fused to obtain fused features. Finally, the clinical skill operation level of the target trainee is evaluated based on the fused features to obtain corresponding evaluation results.
[0093] Furthermore, after evaluating the target trainee's clinical skill level based on the operation feature extraction results and obtaining a skill training level evaluation result, the process may further include: generating operation evaluation information and operation suggestion information for the target trainee using the post-training skill training level evaluation model, thereby obtaining a skill training level evaluation result, a current operation evaluation result, and a current operation suggestion; outputting the current operation suggestion in speech using speech synthesis technology, and adjusting the projection direction of the output speech based on the target trainee's current head posture using a three-dimensional spatial sound field distribution model based on the NeRF architecture, thereby providing real-time guidance on the target trainee's clinical skill operation. In this embodiment, in addition to performing a skill training level evaluation, the post-training skill training level evaluation model may also generate operation evaluation information and operation suggestion information (i.e., feedback information) for the target trainee, thereby obtaining a skill training level evaluation result, a current operation evaluation result, and a current operation suggestion. For example, the post-training skill training level evaluation model may perform a detailed analysis of the target trainee's operation process, identify strengths and weaknesses in the operation, and provide specific improvement suggestions (i.e., current operation suggestions), such as missing operation steps or inadequate operation movements. In addition, detailed scores can be generated based on the target trainees' operational performance and the evaluation results output by the model, including scores on the accuracy, standardization, efficiency, etc. of the operations.
[0094] In addition, this application also sets up a real-time feedback mechanism, which can use speech synthesis technology (such as Tacotron, WaveNet, etc.) to output the current operation suggestions in voice. At the same time, it can also be based on the current target trainee's head posture and use the three-dimensional spatial sound field distribution model based on the NeRF (Neural Radiance Fields, a three-dimensional reconstruction technology achieved through deep learning) architecture to make real-time and dynamic adjustments to the projection direction of the output voice, so as to better guide the target trainees' clinical skills operations during the skills training process.
[0095] Furthermore, after the target trainee's clinical skill operation level is evaluated based on the operation feature extraction result and the skill training level evaluation result is obtained, the method may further include: using a cognitive map analysis engine and creating a clinical skill training level evaluation report for the target trainee based on the skill training level evaluation result, the current operation evaluation result and the current operation suggestion, and visually displaying the clinical skill training level evaluation report; encrypting the clinical skill training level evaluation report using an encryption algorithm to obtain an encrypted evaluation report, and storing the encrypted evaluation report in a preset storage space. In this embodiment, a clinical skill training level evaluation report of the target trainee during the current training process is generated using a cognitive map analysis engine and based on the skill training level evaluation result, the current operation evaluation result and the current operation suggestion output by the model. Furthermore, the target trainee's clinical skill training level evaluation report, including operation data, evaluation results, operation evaluation results, operation suggestions, etc., can be intuitively displayed through visualization tools such as charts and curves, to facilitate the understanding and analysis of the training effect by trainees and teachers. In addition, for the security of data, encryption algorithms such as the Homomorphic Encryption algorithm can be used to encrypt the clinical skills training level assessment report, and the encrypted assessment report can be saved in a preset storage space, such as a preset database.
[0096] It should be pointed out that the relevant data of each skill training of the trainees (such as the content in the clinical skill training level assessment report, etc.) will be recorded in the preset storage space, such as the preset database. Before generating a new assessment report, the historical skill training information of the corresponding trainees will be queried, such as the content and operation scores in the historical assessment report, for comparison with the current skill training data of the current trainees, and fine-grained trends will be analyzed, such as the amplitude data accurate to each operation, and the report template will be set with prompt words to generate a customized real-time assessment report for the current trainees.
[0097] In addition, to improve the security of data in the preset database, corresponding access control policies can be set to prevent malicious access to the data in the database. When accessing data in the preset database, the attribute-based encryption (ABE) algorithm can be used to combine student identity attributes, permissions, and time constraints to perform fine-grained access control, thereby ensuring data security and privacy. In addition, adaptive storage policies can be set. For example, frequently accessed data in the database can be stored in an NVMe SSD (a solid-state drive using the NVMe (Non-Volatile Memory Express) protocol) to reduce access latency, and less frequently accessed data can be stored in a mechanical hard drive to reduce storage costs.
[0098] It can be seen that the embodiment of the present application collects multimodal data of the trainees' clinical skill operations during the clinical skills training process, and uses a multimodal large model based on the dynamic spatiotemporal graph attention network to evaluate the level of the trainees' clinical skill operations. Compared with a single data source, the present application can improve the accuracy of the clinical skill training level assessment by collecting multimodal clinical skill operation data of the trainees during the clinical skills training process. In addition, since the dynamic spatiotemporal graph attention network can improve the correlation and time series alignment accuracy between cross-modal features, the use of a multimodal large model based on the dynamic spatiotemporal graph attention network to evaluate the skill training level can further improve the accuracy of the assessment. At the same time, compared with the manual evaluation method, it saves labor costs and time costs, improves the efficiency of the assessment, and thus improves the efficiency and quality of clinical skills training.
[0099] Specifically, the clinical skills training level assessment scheme proposed in this application can be implemented through a pre-created clinical skills training level assessment system, which includes multiple functional modules, see Figure 3 shown. Figure 3A specific clinical skills training level assessment system is shown, which includes five modules: a multimodal data acquisition module, a data preprocessing module, a large model analysis module, a real-time feedback and assessment module, and a data management and training module; among them, the multimodal data acquisition module is used to collect multimodal data of clinical skills operations of target trainees during the current clinical skills training process, including the collection of video data, audio data, and sensor data; the data preprocessing module is used to receive the video data, audio data, and sensor data input by the multimodal data acquisition module, and perform corresponding preprocessing operations on each modality data, for example, target detection, posture estimation, and image feature extraction on video data; speech recognition, noise filtering, and feature extraction on audio data; and filtering and feature extraction on sensor data. The large model analysis module is used to fuse features of the pre-processed multimodal data, and then perform skill operation identification and operation evaluation in sequence based on the fused features and through the pre-created skill training level assessment model, as well as generate targeted improvement suggestions based on the current student's historical training data, thereby generating an assessment report for the current student; the real-time feedback and assessment module is used to provide real-time operation guidance through voice synthesis and text prompts, and generate detailed assessment reports, including scores, operation analysis and improvement suggestions, etc.; the data management and training module is used to store the student's operation data, model evaluation results and feedback information, and can use medical experts to annotate and verify some data to provide high-quality training data, and then collect new operation data regularly to continuously train and optimize the model. In this way, the system performance is continuously improved to adapt to the needs of different students. In addition, real-time evaluation and feedback can reduce the time and cost of manual evaluation, improve the efficiency of evaluation, and thus improve the efficiency of skills training. Moreover, the model-based evaluation results are more objective and reduce human bias. Providing targeted improvement suggestions based on the specific problems of the trainees (i.e., personalized feedback information) can improve learning effects, thereby solving the problems of subjectivity, inefficiency, lack of real-time feedback, and single data in traditional clinical skills training and assessment. In addition, by collecting multimodal data and combining it with multimodal large model technology, the efficiency and quality of skills training can be further improved.
[0100] Correspondingly, the present application also discloses a clinical skills training level assessment device, see Figure 4 As shown, the device includes:
[0101] The data collection module 11 is used to collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data;
[0102] A preprocessing module 12 is used to preprocess the current multimodal data to obtain preprocessed data;
[0103] A feature extraction module 13 is used to extract features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result;
[0104] A training level assessment module 14 is configured to input the operation feature extraction result into a post-training skill training level assessment model to assess the clinical skill operation level of the target trainee based on the operation feature extraction result to obtain a skill training level assessment result;
[0105] Among them, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data on clinical skill operations of different trainees during historical clinical skills training.
[0106] Among them, the specific work processes of the above modules can refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0107] It can be seen that in the embodiment of the present application, multimodal data of the clinical skill operation of the target trainee in the current clinical skill training process is first collected to obtain current multimodal data, and then the current multimodal data is preprocessed to obtain preprocessed data, and the features related to the clinical skill operation of the target trainee in the preprocessed data are extracted to obtain operation feature extraction results; then, the operation feature extraction results are input into the post-training skill training level evaluation model to evaluate the level of the clinical skill operation of the target trainee based on the operation feature extraction results to obtain skill training level evaluation results; wherein, the post-training skill training level evaluation model is a model obtained after training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data of clinical skill operations for different trainees in the historical clinical skills training process. The embodiment of the present application collects multimodal data of trainees' clinical skill operations during clinical skill training, and uses a multimodal large model based on a dynamic spatiotemporal graph attention network to evaluate the level of trainees' clinical skill operations. Compared with a single data source, the embodiment of the present application collects multimodal clinical skill operation data of trainees during clinical skill training, which can improve the accuracy of clinical skill training level evaluation. In addition, since the dynamic spatiotemporal graph attention network can improve the correlation and time series alignment accuracy between cross-modal features, the use of a multimodal large model based on a dynamic spatiotemporal graph attention network to evaluate the skill training level can further improve the accuracy of the evaluation. At the same time, compared with manual evaluation, it saves labor costs and time costs, improves the efficiency of evaluation, and thus improves the efficiency and quality of clinical skills training.
[0108] In some specific embodiments, the data acquisition module 11 may specifically include:
[0109] The synchronous acquisition unit is used to synchronously collect video data, audio data and sensor data of the clinical skill operation of the target trainees during the current clinical skill training according to the preset time synchronization mechanism to obtain the current multimodal data;
[0110] The video data is captured by an event camera; the audio data is collected by a multi-physics field audio acquisition device based on a beamforming matrix; and the sensor data is collected by a sensor device including a pressure sensor and an acceleration sensor.
[0111] In some specific embodiments, the pre-processing module 12 may specifically include:
[0112] a first processing unit, configured to perform filtering and noise reduction processing on the audio data in the current multimodal data to obtain audio processed data;
[0113] The second processing unit is configured to filter the sensor data in the current multimodal data to obtain sensor-processed data.
[0114] In some specific embodiments, the feature extraction module 13 may specifically include:
[0115] an object detection unit, configured to detect a key operation object in the current video data of the current multimodal data using an object detection algorithm to obtain current operation object detection information;
[0116] a posture estimation unit, configured to estimate the operation posture of the target student in the current video data using a posture estimation algorithm, so as to extract position and angle information of key joints of the target student and obtain operation posture information;
[0117] a first feature extraction unit, configured to perform feature extraction on each video frame in the current video data using a dynamic spatiotemporal attention fusion network pre-created based on a dynamic spatiotemporal attention mechanism, so as to extract features capable of characterizing the operation actions of the target learner and obtain a video feature extraction result;
[0118] A second feature extraction unit is used to extract features that can reflect the voice instructions and operation status of the target student from the audio processed data to obtain an audio feature extraction result;
[0119] The third feature extraction unit is used to extract features that can reflect the physical characteristics of the target student's operation actions in the sensor-processed data to obtain sensor feature extraction results.
[0120] In some specific embodiments, the clinical skills training level assessment device may further include:
[0121] The data collection unit is used to collect multimodal data of clinical skill operations of different trainees during historical clinical skill training, as well as corresponding operation evaluation information and operation suggestion information, to obtain historical multimodal data;
[0122] The model training unit is used to input the historical multimodal data into a multimodal large model based on a dynamic spatiotemporal graph attention network in a supervised learning manner for model training to obtain the post-training skill training level evaluation model.
[0123] In some specific embodiments, the training level assessment module 14 may specifically include:
[0124] a feature fusion unit, configured to perform feature fusion on the video feature extraction result, the operation posture information, and the current operation object detection information to obtain fused video features;
[0125] a training level assessment unit, configured to incorporate the fused features, the audio feature extraction results, and the sensor feature extraction results into a post-training skill training level assessment model, to semantically align the fused video features, the audio feature extraction results, and the sensor feature extraction results to obtain corresponding aligned video features, aligned audio features, and aligned sensor features, and to perform feature fusion on the aligned video features, the aligned audio features, and the aligned sensor features to obtain fused features, and then to assess the level of clinical skill operation of the target trainee based on the fused features to obtain a skill training level assessment result;
[0126] Accordingly, the training level assessment module 14 may further include:
[0127] A first generating unit is configured to generate operation evaluation information and operation suggestion information for the target trainee through the post-training skill training level evaluation model, and obtain a skill training level evaluation result, a current operation evaluation result, and a current operation suggestion;
[0128] A speech synthesis unit, configured to output the current operation suggestion in speech using speech synthesis technology;
[0129] An adjustment unit is used to adjust the projection direction of the output voice based on the current head posture of the target trainee using a three-dimensional spatial sound field distribution model based on the NeRF architecture, so as to guide the clinical skill operation of the target trainee in real time.
[0130] In some specific embodiments, the clinical skills training level assessment device may further include:
[0131] a second generating unit, configured to create a clinical skills training level assessment report for the target trainee based on the skills training level assessment result, the current operation evaluation result, and the current operation suggestion by using a cognitive map analysis engine, and to visually display the clinical skills training level assessment report;
[0132] A report encryption unit, configured to encrypt the clinical skills training level assessment report using an encryption algorithm to obtain an encrypted assessment report;
[0133] The report storage unit is used to store the encrypted evaluation report in a preset storage space.
[0134] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram should not be considered as any limitation to the scope of application of the present application.
[0135] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the clinical skills training level assessment method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0136] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0137] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0138] The operating system 221 is used to manage and control the hardware devices and computer program 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the clinical skills training level assessment method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program capable of implementing other specific tasks.
[0139] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the aforementioned method for evaluating clinical skills training level is implemented. The specific steps of this method can be referred to the corresponding contents disclosed in the aforementioned embodiments and will not be repeated here.
[0140] Furthermore, an embodiment of the present application also discloses a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the clinical skills training level assessment method disclosed above.
[0141] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0142] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0143] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0144] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0145] The above is a detailed introduction to the clinical skills training level assessment method, device, equipment and storage medium provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea; at the same time, for general technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A method for evaluating clinical skills training level, characterized in that: include: Collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data; Preprocessing the current multimodal data to obtain preprocessed data, and extracting features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result; Inputting the operation feature extraction results into a post-training skill training level assessment model to assess the clinical skill operation level of the target trainee based on the operation feature extraction results to obtain a skill training level assessment result; Among them, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data on clinical skill operations of different trainees during historical clinical skills training.
2. The clinical skills training level assessment method according to claim 1, characterized in that: The multimodal data of clinical skill operations of target trainees during the current clinical skill training process is collected to obtain current multimodal data, including: According to the preset time synchronization mechanism, the video data, audio data and sensor data of the clinical skill operation of the target trainees in the current clinical skill training process are synchronously collected to obtain the current multimodal data; The video data is captured by an event camera; the audio data is collected by a multi-physics field audio acquisition device based on a beamforming matrix; and the sensor data is collected by a sensor device including a pressure sensor and an acceleration sensor.
3. The clinical skills training level assessment method according to claim 2, characterized in that: The preprocessing of the current multimodal data to obtain preprocessed data includes: Performing filtering and noise reduction on the audio data in the current multimodal data to obtain audio processed data; Filtering is performed on the sensor data in the current multimodal data to obtain sensor-processed data.
4. The clinical skills training level assessment method according to claim 3, characterized in that: The extracting features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result includes: Detecting a key operation object in current video data of the current multimodal data using an object detection algorithm to obtain current operation object detection information; Using a posture estimation algorithm to estimate the operation posture of the target student in the current video data to extract position and angle information of key joints of the target student to obtain operation posture information; Using a dynamic spatiotemporal attention fusion network pre-created based on a dynamic spatiotemporal attention mechanism, feature extraction is performed on each video frame in the current video data to extract features that can characterize the target student's operation action, thereby obtaining a video feature extraction result; Extracting features that can reflect the target student's voice instructions and operation status from the audio processed data to obtain an audio feature extraction result; Features that can reflect the physical characteristics of the target trainee's operating actions in the sensor-processed data are extracted to obtain sensor feature extraction results.
5. The clinical skills training level assessment method according to claim 4, characterized in that: Also includes: Collect multimodal data on clinical skill operations of different trainees during historical clinical skill training, as well as corresponding operation evaluation information and operation suggestion information, to obtain historical multimodal data; The historical multimodal data is input into a multimodal large model based on a dynamic spatiotemporal graph attention network using a supervised learning method for model training to obtain the post-training skill training level evaluation model.
6. The clinical skills training level assessment method according to claim 5, characterized in that: Inputting the operation feature extraction results into a post-training skill training level assessment model to assess the target trainee's clinical skill operation level based on the operation feature extraction results to obtain a skill training level assessment result includes: Performing feature fusion on the video feature extraction result, the operation posture information, and the current operation object detection information to obtain fused video features; The fused features, the audio feature extraction results, and the sensor feature extraction results are added to a post-training skill training level assessment model to semantically align the fused video features, the audio feature extraction results, and the sensor feature extraction results to obtain corresponding aligned video features, aligned audio features, and aligned sensor features, and feature fusion is performed on the aligned video features, the aligned audio features, and the aligned sensor features to obtain fused features, and then the level of clinical skill operation of the target trainee is assessed based on the fused features to obtain a skill training level assessment result; Accordingly, after evaluating the clinical skill operation level of the target trainee based on the operation feature extraction result and obtaining the skill training level evaluation result, the method further includes: Generate operation evaluation information and operation suggestion information for the target trainee through the post-training skill training level evaluation model, and obtain skill training level evaluation results, current operation evaluation results and current operation suggestions; The current operation suggestion is output in voice using speech synthesis technology, and the projection direction of the output voice is adjusted based on the current head posture of the target trainee using a three-dimensional spatial sound field distribution model based on the NeRF architecture to guide the clinical skill operation of the target trainee in real time.
7. The clinical skills training level assessment method according to claim 6, characterized in that: Also includes: Using a cognitive map analysis engine and based on the skill training level assessment result, the current operation evaluation result, and the current operation suggestion, a clinical skill training level assessment report for the target trainee is created, and the clinical skill training level assessment report is visually displayed; The clinical skills training level assessment report is encrypted using an encryption algorithm to obtain an encrypted assessment report, and the encrypted assessment report is stored in a preset storage space.
8. A clinical skills training level assessment device, characterized in that: include: The data acquisition module is used to collect multimodal data of clinical skill operations of target trainees during the current clinical skill training process to obtain current multimodal data; A preprocessing module, configured to preprocess the current multimodal data to obtain preprocessed data; A feature extraction module is used to extract features related to the clinical skill operation of the target trainee from the preprocessed data to obtain an operation feature extraction result; A training level assessment module is used to input the operation feature extraction results into a post-training skill training level assessment model to assess the clinical skill operation level of the target trainee based on the operation feature extraction results to obtain a skill training level assessment result; Among them, the post-training skill training level evaluation model is a model obtained by training a multimodal large model based on a dynamic spatiotemporal graph attention network using historical multimodal data, and the historical multimodal data includes multimodal data on clinical skill operations of different trainees during historical clinical skills training.
9. An electronic device, characterized in that: It comprises a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the clinical skills training level assessment method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program; wherein, when the computer program is executed by a processor, it implements the clinical skills training level assessment method as described in any one of claims 1 to 7.
Citation Information
Cited By
Data acquisition personnel skill evaluation system, method and device
CN121235550A