Multi-modal data driven caregiver training method and device, equipment and medium

Through multimodal data synchronization and cross-modal feature fusion, virtual reality training courses are generated in real time, which solves the one-sided problems of action-instruction synergistic distortion and single-modal evaluation in traditional nursing training, and improves the standardization and safety of nursing operations.

CN120580907AActive Publication Date: 2025-09-02HANGZHOU LINAN DISTRICT FIRST PEOPLES HOSPITAL (MEDICAL COMMUNITY OF HANGZHOU LINAN DISTRICT FIRST PEOPLES HOSPITAL)

Patent Information

Application Number
CN202511073049.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-09-02
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

In traditional nursing training methods, the time deviation of multi-source data leads to distortion of the correlation analysis of actions and instructions, and the single-modal evaluation system is difficult to comprehensively quantify the quality of nursing actions, and the lack of real-time feedback and personalized training, resulting in poor transfer of training effects in real clinical scenarios.

Method used

Through multimodal data synchronization and cross-modal feature fusion, a time-series neural network and probability generation model are used to generate virtual reality training courses in real time, realizing accurate evaluation and personalized training of nursing actions, and generating a personalized training task instruction set.

Benefits of technology

It realizes millisecond alignment of multimodal signals throughout the nursing operation, solves the one-sided problem of single-modal evaluation, improves the standardization and safety of nursing operation, and provides a quantifiable, traceable and reusable intelligent training solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580907A_ABST
    Figure CN120580907A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal data driven caregiver training method and device, equipment and a medium, and the method comprises the steps: eliminating the clock deviation of multi-source equipment through the synchronous processing of joint movement, a voice instruction and hand pressure data, and generating a time alignment data set; fusing the nursing action dynamic features and the pressure gradient features by using a time sequence neural network, and constructing a space-time fusion sequence; the joint angle change rate and the pressure parameter are decoded based on a probability model, and an operation continuity index is quantified in combination with time window integration; and finally, identifying weak links of skills according to the mechanical parameter sequence, and matching a clinical standard course library to generate a personalized training instruction. According to the method, the problems of action-instruction collaborative distortion, one-sided single-mode evaluation and feedback disjunction caused by multi-source data time delay are solved, real-time accurate evaluation and immersive adaptive training of nursing operation biomechanical parameters are realized, and the venipuncture stability and the pressure sore prevention force control capability are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of nursing assistance training, and in particular to a multimodal data-driven caregiver training method, apparatus, equipment and medium. Background Art

[0002] Clinical nursing skills training is a key component in improving nursing quality, especially for delicate procedures like venipuncture and pressure ulcer prevention. Traditional training methods rely primarily on manual observation and static model drills. These methods transfer skills by having instructors demonstrate standard movements in simulated scenarios, with trainees practicing repeatedly and receiving verbal corrections. However, the dynamic nature of nursing procedures, the multi-factor coupling, and the need for immediate feedback significantly limit the transferability of traditional training models to real-world clinical scenarios.

[0003] There are three core defects in the existing technology: First, because the joint motion monitoring equipment, voice command system and pressure sensor device use independent clock sources, the time deviation of multi-source data leads to distorted analysis of the correlation between movements and commands, and cannot accurately capture the synergistic relationship between changes in elbow joint stability and voice commands at the moment of intravenous puncture; second, a single-modal evaluation system that only analyzes joint angles or only monitors pressure is difficult to fully quantify the coupling effect of spinal rotation dynamics and sacral pressure distribution during turning operations, resulting in a one-sided assessment of nursing action quality; third, training feedback is disconnected from actual operations, and there is a lack of the ability to generate personalized courses based on biomechanical parameters in real time. Trainees cannot receive immersive and intensive training for defects such as needle shaking and sudden changes in turning force. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a multimodal data-driven caregiver training method, device, equipment and medium that can achieve accurate synchronization of multimodal data, integrate dynamic biomechanical assessment and instantly generate virtual reality training courses.

[0005] The purpose of the present invention is achieved by the following scheme: In a first aspect, the present invention provides a multimodal data-driven caregiver training method, comprising the following steps: S1: Time-synchronize the joint motion data, voice command data, and hand pressure data collected by the caregiver during the operation, align the timestamps of the joint motion sensor clock, voice acquisition device clock, and pressure sensor clock, and generate a time-aligned multimodal dataset; S2: Perform cross-modal feature fusion processing on the time-aligned multimodal dataset, extract the dynamic features of nursing actions through a temporal neural network, calculate the gradient features of palm pressure distribution, and generate a spatiotemporal fusion feature sequence; S3: Perform micro-motion analysis on the spatiotemporal fusion feature sequence, decode the elbow flexion and extension angle change rate and the palm key area pressure change parameters through a preset probabilistic generation model, calculate the continuity index of the nursing action fluency, and generate a mechanical parameter sequence; S4: Make adaptive feedback decisions on the mechanical parameter sequence, identify weak links in skills based on abnormal joint angle offset rate and pressure gradient mutation value, match the preset clinical nursing standardized course library, generate personalized training task instruction set, and send the training task instruction set to the virtual reality training terminal. The personalized training task instruction is used to indicate the intensive training content for the weak links in the caregiver's skills.

[0006] In one embodiment, S1 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S11: The wearable inertial measurement unit is used to collect and process the caregiver's elbow joint motion data in real time, and the acquired three-dimensional angular velocity vector is analyzed and processed to calculate the synthetic linear velocity in the joint motion plane and generate the wrist joint linear velocity characteristic value; S12: Perform time-series tagging on the voice instruction data given to the caregiver, call the endpoint detection algorithm to segment the key instruction phrases of the nursing operation, and generate a word timestamp sequence; S13: The flexible sensor array is used to monitor the palm contact pressure in different areas, and dynamic delay compensation is performed on the pressure distribution vectors of the sacral area and thenar area of ​​the palm to generate a multi-device clock synchronization offset. S14: Perform time axis reconstruction on the wrist joint linear velocity feature values, word timestamp sequence, and multi-device clock synchronization offset, map the original data to a unified time base, and generate a time-aligned multimodal dataset.

[0007] In one embodiment, S2 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S21: Extract and process the turning action features from the joint motion data of the multimodal dataset, model the temporal dependency of spinal rotation during the turning operation through a gated recurrent unit network, and generate a dynamic encoding vector for the turning action; S22: parsing the voice command data in the multimodal dataset for nursing intent, calling a preset clinical term matching model to identify key action instructions for turning over, and generating a nursing operation type code; S23: Perform gradient analysis on the sacral region pressure data of the multimodal dataset and calculate the pressure distribution change rate for pressure ulcer prevention using a spatial difference algorithm; S24: Perform cross-modal fusion processing on the dynamic coding vector of turning action, nursing operation type coding and pressure ulcer prevention pressure distribution change rate, and integrate multi-source features through weighted attention mechanism to generate a spatiotemporal fusion feature sequence.

[0008] In one embodiment, S3 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S31: Perform venipuncture motion analysis on the spatiotemporal fusion feature sequence, identify the needle stability characteristics by analyzing the elbow flexion and extension angle change pattern, and generate venipuncture stability parameters; S32: Perform force control analysis on the spatiotemporal fusion feature sequence, identify pressure control defects by modeling the evolution of sacral pressure distribution, and generate pressure ulcer prevention force control parameters; S33: Perform fluency assessment on the intravenous puncture stability parameters and pressure ulcer prevention force control parameters, calculate the nursing operation continuity index through time window integration, and generate a mechanical parameter sequence. The mechanical parameter sequence is used to indicate the intravenous puncture stability level and the turning force control ability score.

[0009] In one embodiment, the calculation formulas for the venipuncture stability parameter and the pressure ulcer prevention force control parameter of a multimodal data-driven caregiver training method provided by the present invention are: ; in, is the venipuncture stability parameter, Control parameters for pressure ulcer prevention, is the total duration of the venipuncture procedure, is the elbow flexion and extension angle, is the sensitivity adjustment coefficient, is the absolute value of the elbow joint angle change rate, The number of sampling points for turning operation, For the The sacral pressure gradient change rate at the sampling point, is the pressure mutation threshold, is the normalization coefficient.

[0010] In one embodiment, S33 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S331: performing time window integration processing on the venipuncture stability parameters, calculating the average stability level of the elbow joint angle change rate through a sliding window, and generating a venipuncture stability grade; S332: Perform time window integration processing on the pressure ulcer prevention force control parameters, evaluate the smoothness of sacral pressure changes through a sliding window, and generate a turning force control ability score; S333: Perform temporal combination processing on the venipuncture stability grade and the turning force control ability score, construct a dual-parameter sequence according to the operation time sequence, and generate a mechanical parameter sequence.

[0011] In one embodiment, S4 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S41: Identify venipuncture defects based on the mechanical parameter sequence, analyze the deviation between the elbow joint angle deviation trend and the standard operation through a decision tree model, and generate a puncture stability defect report; S42: Identify turning force control defects based on the mechanical parameter sequence, locate abnormal change points of sacral pressure gradient through mutation detection algorithm, and generate a pressure ulcer prevention force control defect report; S43: Match the puncture stability deficiency report and the pressure ulcer prevention force control deficiency report to the course, search the standardized course library based on the pre-existing clinical nursing operation classification system, match the skill deficiency type with the training course mapping relationship, and generate a personalized training course combination; S44: Perform VR instruction conversion on the personalized training course combination, generate an executable training task instruction set through the virtual reality interface protocol, and send the training task instruction set to the virtual reality training terminal.

[0012] In a second aspect, the present invention provides a multimodal data-driven caregiver training device, which is configured with the following modules: The multimodal data time synchronization module is used to synchronize the joint motion data, voice command data, and hand pressure data collected during the caregiver's operation, align the timestamps of the joint motion sensor clock, the voice acquisition device clock, and the pressure sensor clock, and generate a time-aligned multimodal data set; The cross-modal feature fusion module is used to perform cross-modal feature fusion processing on the time-aligned multimodal dataset. It extracts the dynamic features of nursing actions through a temporal neural network and calculates the gradient features of palm pressure distribution to generate a spatiotemporal fusion feature sequence. The micro-motion analysis module is used to perform micro-motion analysis on the spatiotemporal fusion feature sequence. It decodes the elbow flexion and extension angle change rate and the pressure change parameters of the key areas of the palm through a preset probability generation model, calculates the continuity index of the nursing action fluency, and generates a mechanical parameter sequence. The adaptive feedback decision module is used to make adaptive feedback decisions on the mechanical parameter sequence, identify weak links in skills based on abnormal joint angle offset rate and pressure gradient mutation value, match the preset clinical nursing standardized course library, generate personalized training task instruction set, and send the training task instruction set to the virtual reality training terminal. The personalized training task instruction is used to indicate the intensive training content for the weak links in the caregiver's skills.

[0013] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements any of the above-mentioned multimodal data-driven caregiver training methods.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-mentioned multimodal data-driven caregiver training methods.

[0015] In summary, the multimodal data-driven caregiver training method provided by the present application achieves precise time synchronization of joint movement, voice commands and hand pressure data by dynamically compensating for the clock deviation of multiple source devices, which can eliminate the problem of motion-command coordination distortion caused by sensor clock asynchrony in traditional training, and achieve the technical effect of millisecond-level alignment of multimodal signals throughout the nursing operation process; based on the temporal neural network, the dynamic characteristics of joint movement and the pressure gradient distribution characteristics are integrated to achieve the coupled modeling of the rotational mechanical characteristics of the spine and the spatial evolution of sacral pressure during the turning operation, so as to solve the one-sided measurement of complex nursing actions by the single-modal evaluation system. Defects; decoding the elbow joint angle change rate and key area pressure parameters through a probabilistic generative model, combined with a time window integration mechanism to calculate the operation continuity index, can achieve a multi-dimensional biomechanical assessment of intravenous puncture stability and turning force control ability, overcoming the technical limitation of static assessment models that cannot fully analyze dynamic operation processes; finally, matching the clinical standard course library based on the joint angle deviation trend and pressure gradient mutation characteristics can achieve intelligent mapping of skill weaknesses with virtual reality training courses, generate personalized training task instruction sets to drive immersive reinforcement training, and address the core pain points of traditional training: delayed feedback and disconnection from actual operation. This method systematically constructs a technical closed loop of "data synchronization-feature fusion-defect identification-adaptive training", which can significantly improve the standardization and safety of nursing operations, and provide a quantifiable, traceable, and reusable intelligent solution for clinical nursing skills training.

[0016] For better understanding and implementation, the present invention is described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart of a multimodal data-driven caregiver training method provided in an embodiment of the present application; Figure 2 A schematic diagram of a process for generating a mechanical parameter sequence according to an embodiment of the present application; Figure 3 A flowchart of a personalized training task instruction set provided in an embodiment of the present application; Figure 4A schematic structural diagram of a multimodal data-driven caregiver training device provided in another embodiment of the present application. DETAILED DESCRIPTION

[0018] To facilitate understanding of the present invention, the present invention will be described more fully below with reference to the accompanying drawings. The drawings illustrate preferred embodiments of the present invention. However, the present invention may be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and comprehensive understanding of the present disclosure.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which this invention pertains. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0020] In one embodiment, Figure 1 As shown, a multimodal data-driven caregiver training method is provided. This embodiment uses the method applied to a terminal as an example for illustration. It is understandable that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps: S1: Time-synchronize the joint motion data, voice command data, and hand pressure data collected by the caregiver during the operation, align the timestamps of the joint motion sensor clock, voice acquisition device clock, and pressure sensor clock, and generate a time-aligned multimodal dataset.

[0021] Specifically, the system receives joint motion data, voice command data, and hand pressure data generated by the caregiver during the operation. The joint motion data is output by a wearable inertial measurement unit array distributed at the elbow, wrist, and metacarpophalangeal joints, and the output content includes timestamps, joint identifiers, three-dimensional acceleration, three-dimensional angular velocity, and three-dimensional angles; the voice command data is output by a microphone array, and the output content includes timestamps, voice frame sequences, and signal-to-noise ratios; and the hand pressure data is output by a flexible thin film pressure sensor array covering the palm and fingertips, and the output content includes timestamps, pressure matrices, and sensor temperature compensation values.

[0022] Exemplarily, the system connects all acquisition devices to a time synchronization gateway based on the Precision Time Protocol, keeps the clocks of each device consistent through a master-slave clock synchronization mechanism, runs a dynamic phase calibration program, and uses a sliding time window algorithm to monitor the data transmission delay of each modality in real time. The wireless transmission delay is calculated for joint motion data, the acoustic echo cancellation processing delay is calculated for voice data, and the analog-to-digital conversion delay is calculated for pressure data. Preferably, the system can predict and compensate for the monitored delays through a Kalman filter algorithm, generate a unified timestamp, and control the time deviation of joint motion data, voice command data, and hand pressure data within a preset range, thereby forming a time-aligned multimodal data set. The data structure of the data set is a corresponding set of time and each sub-data set, covering the entire time period from the start to the end of the operation.

[0023] S2: Perform cross-modal feature fusion processing on the time-aligned multimodal dataset, extract the dynamic features of nursing actions through the temporal neural network, calculate the gradient features of palm pressure distribution, and generate a spatiotemporal fusion feature sequence.

[0024] Specifically, the system can use a bidirectional long short-term memory network, which includes an input layer, a hidden layer, and an output layer. The input layer corresponds to three-dimensional acceleration and three-dimensional angular velocity, the hidden layer contains a certain number of neurons, and the output layer is a dynamic feature vector. The system divides the joint angle sequence through a sliding time window, extracts the joint motion trend characteristics, movement smoothness characteristics and motion trajectory entropy in each window, and outputs a dynamic feature matrix. The matrix dimension is related to the number of windows. The system preprocesses the pressure matrix in the hand pressure sub-dataset, performs mean filtering, filters the data with a convolution kernel, and then performs normalization processing to map the values ​​to the specified interval.

[0025] Preferably, the system can use the Sobel operator to calculate the pressure gradient field, obtain the horizontal gradient matrix and the vertical gradient matrix, calculate the gradient amplitude matrix through the formula, the matrix coordinates correspond to the position of the pressure sensor array, and extract the gradient maximum position, gradient direction entropy, and pressure gradient change rate, and output the pressure gradient feature matrix. The matrix dimension is related to the number of windows. The system constructs an attention mechanism fusion network, which includes a feature mapping layer, an attention weight calculation layer, and a fusion layer. Among them, the feature mapping layer maps the dynamic features and pressure features to the same dimensional space through a fully connected layer. The attention weight calculation layer calculates the correlation weights of the two types of features through a function, and the fusion layer obtains the fusion features through weighted summation. The system generates a spatiotemporal fusion feature sequence. The data structure is a corresponding set of window timestamps and fusion feature vectors. The fusion feature vector is the result of attention fusion of dynamic features and pressure gradient features.

[0026] S3: Perform micro-motion analysis on the spatiotemporal fusion feature sequence, decode the elbow flexion and extension angle change rate and the pressure change parameters of the key areas of the palm through the preset probability generation model, calculate the continuity index of the nursing action fluency, and generate a mechanical parameter sequence.

[0027] Specifically, the pre-set probabilistic generation model can be a deep belief network, consisting of a visible layer, a hidden layer, and an output layer. The visible layer corresponds to the spatiotemporal fusion features, the hidden layer contains a certain number of neurons, and the output layer is a parameter vector. This model is obtained through training, which includes unsupervised pre-training and supervised fine-tuning. Unsupervised pre-training can use the contrastive divergence algorithm, while supervised fine-tuning uses the backpropagation algorithm. The training dataset contains the spatiotemporal fusion feature sequences manipulated by experts and the corresponding annotated parameters. The annotated parameters are provided by biomechanics experts and include parameters such as the elbow angle change rate and palm pressure.

[0028] For example, the system inputs the spatiotemporal fusion feature sequence into a deep belief network and decodes it to obtain the elbow flexion and extension angle change rate and the pressure change parameters of key palm regions. The elbow flexion and extension angle change rate is calculated by comparing the elbow angle difference and time interval between consecutive frames. The angle difference is the difference in elbow angle at different moments. The key palm regions include the palm, thumb pad, and index finger pad. The pressure change parameters of each region are obtained by calculating the change rate of the regional pressure mean, which is the ratio of the difference in regional pressure mean at different moments to the time interval.

[0029] Based on the decoded parameters, the system calculates movement smoothness and continuity indicators, including angle change smoothness and pressure change continuity. Angle change smoothness is the standard deviation of the elbow angle change rate, and pressure change continuity is the mean of the absolute values ​​of the first-order derivatives of the palm pressure change rate. The system integrates the elbow angle change rate, pressure change parameters in key palm areas, angle change smoothness, pressure change continuity, and other auxiliary biomechanical parameters to form a mechanical parameter sequence. The sequence dimension includes multiple parameter items.

[0030] S4: Make adaptive feedback decisions on the mechanical parameter sequence, identify weak links in skills based on abnormal joint angle offset rate and pressure gradient mutation value, match the preset clinical nursing standardized course library, generate personalized training task instruction set, and send the training task instruction set to the virtual reality training terminal. The personalized training task instruction is used to indicate the intensive training content for the weak links in the caregiver's skills.

[0031] Specifically, the system receives a sequence of mechanical parameters and calls a preset biomechanical parameter threshold, which is obtained based on the statistics of the operation data of senior caregivers, including a threshold for the abnormal offset rate of joint angles and a threshold for the pressure gradient mutation value. The abnormal offset rate of joint angles is the deviation between the actual angle and the standard angle divided by the standard angle, and the pressure gradient mutation value is the difference between the pressure gradients of adjacent windows divided by the pressure gradient of the previous window. Exemplarily, the system scans the sequence of mechanical parameters through a rule engine. When the elbow joint angle change rate exceeds the threshold for multiple consecutive windows, it is determined to be insufficient joint stability; when the palm pressure gradient mutation value exceeds the threshold, it is determined to be a force mutation; when the movement smoothness continuity index does not meet the preset threshold, it is determined to be insufficient movement coherence. The system summarizes the judgment results and the frequency of occurrence to generate a list of weak links in skills.

[0032] Preferably, the system accesses a standardized clinical nursing curriculum library, which has a hierarchical structure consisting of a primary module, secondary submodules, and tertiary training units. The primary module includes venipuncture, turning, and pressure ulcer prevention, while the secondary submodules are subdivisions of the primary module, such as venipuncture, which includes needle angle control and needle holding stability training. Each training unit is associated with a standardized biomechanical parameter template, such as the range of elbow angle change rate and palm pressure gradient change during venipuncture, as well as corresponding virtual reality scene parameters, such as the virtual patient's body shape, vascular characteristics, and ambient lighting.

[0033] Preferably, based on a list of weak skill areas, the system can match training units in the course library using a cosine similarity algorithm, with the matching degree reaching a preset value. The system dynamically adjusts training parameters. For cases of insufficient joint stability, it adjusts the movement speed of the virtual blood vessels in the virtual reality scene and adjusts the needle insertion resistance fluctuation through a tactile feedback device. For cases of sudden force changes, a real-time visualization interface for sacral pressure is set in the turning training scene, triggering an early warning when the pressure suddenly changes. The system generates a training task instruction set in a specified format, including an instruction ID, a training module identifier, virtual reality scene parameters, training duration, assessment indicators, and feedback methods. The assessment indicators are biomechanical parameter thresholds, and the feedback methods include vision, touch, and voice. The system sends the training task instruction set to the virtual reality training terminal through a communication module that meets the transmission rate and latency requirements. The virtual reality training terminal includes a head-mounted display device, force feedback gloves, and a motion capture base station. The system control terminal parses the instructions, calls the engine to load the corresponding virtual scene, and starts the real-time motion capture and feedback module.

[0034] In summary, the multimodal data-driven caregiver training method provided by the present application achieves precise time synchronization of joint movement, voice commands and hand pressure data by dynamically compensating for the clock deviation of multiple source devices, which can eliminate the problem of motion-command coordination distortion caused by sensor clock asynchrony in traditional training, and achieve the technical effect of millisecond-level alignment of multimodal signals throughout the nursing operation process; based on the temporal neural network, the dynamic characteristics of joint movement and the pressure gradient distribution characteristics are integrated to achieve the coupled modeling of the rotational mechanical characteristics of the spine and the spatial evolution of sacral pressure during the turning operation, so as to solve the one-sided measurement of complex nursing actions by the single-modal evaluation system. Defects; decoding the elbow joint angle change rate and key area pressure parameters through a probabilistic generative model, combined with a time window integration mechanism to calculate the operation continuity index, can achieve a multi-dimensional biomechanical assessment of intravenous puncture stability and turning force control ability, overcoming the technical limitation of static assessment models that cannot fully analyze dynamic operation processes; finally, matching the clinical standard course library based on the joint angle deviation trend and pressure gradient mutation characteristics can achieve intelligent mapping of skill weaknesses with virtual reality training courses, generate personalized training task instruction sets to drive immersive reinforcement training, and address the core pain points of traditional training: delayed feedback and disconnection from actual operation. This method systematically constructs a technical closed loop of "data synchronization-feature fusion-defect identification-adaptive training", which can significantly improve the standardization and safety of nursing operations, and provide a quantifiable, traceable, and reusable intelligent solution for clinical nursing skills training.

[0035] In one embodiment, S1 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S11: The wearable inertial measurement unit is used to collect and process the caregiver's elbow joint motion data in real time, and the acquired three-dimensional angular velocity vector is analyzed and processed to calculate the synthetic linear velocity in the joint motion plane and generate the wrist joint linear velocity characteristic value.

[0036] Specifically, the system collects real-time data on the caregiver's elbow joint motion through a wearable inertial measurement unit (IMU). This wearable IMU is fixed to the caregiver's elbow joint and has a built-in 3D acceleration sensor and 3D angular velocity sensor. This unit continuously acquires the elbow joint's motion parameters in 3D space, and the output data includes a timestamp and the corresponding 3D angular velocity vector. After receiving the 3D angular velocity vector, the system determines the joint motion plane based on the anatomical structural parameters of the elbow joint. This plane is defined by the line containing the long axis of the humerus and the long axis of the ulna. The system then calculates the composite linear velocity of the wrist joint within this plane using a kinematic transformation formula based on the components of the 3D angular velocity vector within the joint motion plane and the anatomical distance parameters from the elbow to the wrist joint.

[0037] For example, the system decomposes the three-dimensional angular velocity vector into two orthogonal directions within the joint motion plane, obtaining the in-plane angular velocity components. These components are then multiplied by the corresponding lever arm lengths and vectorized to produce the composite linear velocity of the wrist joint. The system extracts features from the composite linear velocity, calculating the peak value, mean value, and rate of change of the linear velocity per unit time. These features are integrated to form the wrist joint linear velocity eigenvalue, which is stored as a data frame, with each frame containing a timestamp and corresponding characteristic parameters.

[0038] S12: Perform time-series tagging on the voice instruction data given to the caregiver, call the endpoint detection algorithm to segment the key instruction phrases of the nursing operation, and generate a word timestamp sequence.

[0039] Specifically, the system receives voice command data issued to caregivers. This data is acquired in real time by a voice acquisition device in the form of a continuous audio signal stream containing timestamp information. The system performs time-series tagging on the voice command data, adding a corresponding timing label to each sampling point in the audio signal stream and establishing a correspondence between the sampling point and time. The system then calls an endpoint detection algorithm to process the audio signal stream. This algorithm distinguishes between speech and non-speech signals by analyzing the energy value, zero-crossing rate, and spectral characteristics of the audio signal, thereby segmenting the key instruction phrases for the nursing operation.

[0040] During endpoint detection, the system sets energy and zero-crossing rate thresholds. When both the energy value and the zero-crossing rate of the audio signal exceed the preset thresholds, it is determined to be the speech start point. When both the energy value and the zero-crossing rate are below the preset thresholds and persist for a certain period of time, it is determined to be the speech end point. This determines the time interval of each key command phrase. The system performs speech recognition on the segmented key command phrases, converting them into corresponding text words. It also records the start and end time of each word in the audio signal stream, generating a word timestamp sequence. This sequence is stored in list form, with each list item containing the word text, a start timestamp, and an end timestamp.

[0041] S13: The palm contact pressure is monitored and processed in different areas through a flexible sensor array, and the pressure distribution vectors of the sacral area and thenar area of ​​the palm are dynamically delayed and compensated to generate a multi-device clock synchronization offset.

[0042] Specifically, the system monitors palm contact pressure in zones using a flexible sensor array. This array covers the caregiver's palm and the areas of contact with the patient, collecting pressure data from different areas of the palm and the patient's sacrum. The system divides the sensor array into multiple monitoring zones, including the palmar thenar eminence and the corresponding contact zone of the sacrum. The sensors in each zone output a pressure distribution vector for that zone. Each element in the vector corresponds to the pressure value of a single sensor in the zone, along with the acquisition timestamp corresponding to that pressure value.

[0043] For example, after receiving the pressure distribution vectors of the sacral region and thenar region of the palm, the system performs dynamic delay compensation processing on them, extracts the timestamp information in the pressure distribution vectors of the two regions, and calculates the acquisition time difference of the pressure data of the two regions at the same time. This time difference reflects the time delay of different sensor devices in the data acquisition process. Preferably, based on the time difference data at multiple times, the system can use a linear regression algorithm to fit the delay change trend, determine the clock deviation law between different devices, and then calculate the multi-device clock synchronization offset. This offset is expressed in the form of a time difference value and is used to quantify the degree of deviation between the clocks of different sensor devices.

[0044] S14: Perform time axis reconstruction on the wrist joint linear velocity feature values, word timestamp sequence, and multi-device clock synchronization offset, map the original data to a unified time base, and generate a time-aligned multimodal dataset.

[0045] Specifically, the system obtains wrist joint linear velocity feature values, word timestamp sequences, and multi-device clock synchronization offsets, using the time signal provided by a high-precision clock source as a unified time reference. This reference time signal has stable frequency and phase characteristics. For wrist joint linear velocity feature values, the system corrects the timestamp of each feature value based on its original timestamp and the corresponding device offset value in the multi-device clock synchronization offset, maps it to a unified time reference, and obtains a corrected wrist joint linear velocity feature value time series. For word timestamp sequences, the system adjusts the start and end timestamps of each word based on the clock synchronization offset corresponding to the voice acquisition device, so that they are consistent with the unified time reference, forming a corrected word time interval sequence.

[0046] For the pressure distribution-related data after dynamic delay compensation, the system also performs timestamp correction based on its corresponding clock synchronization offset, maps it to a unified time base, and integrates the corrected wrist joint linear velocity eigenvalue time series, word time interval series, and pressure distribution-related time series. The data items are arranged in chronological order according to the unified time base to ensure that the correspondence between different modal data in the time dimension is accurate, and finally generates a time-aligned multimodal data set, which contains a complete record of each modal data under a unified time base.

[0047] In one embodiment, S2 of a multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S21: Extract and process the turning action features of the joint motion data of the multimodal dataset, model the temporal dependency of spinal rotation during the turning operation through a gated recurrent unit network, and generate a dynamic encoding vector for the turning action.

[0048] Specifically, the system extracts joint motion data from a time-aligned multimodal dataset. This data contains the motion parameters of the caregiver's spine and related joints during the turning operation, including the rotation angle of each vertebra, the joint displacement, and the corresponding timestamp. The system calls a gated recurrent unit network to process this data. The network consists of an input layer, multiple gated recurrent unit layers, and an output layer. The input layer receives the feature vector of the joint motion data, which contains the rotational angular velocity, angular acceleration, and relative displacement between adjacent vertebrae of each spinal segment; the gated recurrent unit layer controls the flow of information through reset gates and update gates. The reset gate determines whether to ignore historical state information, and the update gate controls the degree of influence of historical state information on the current state. In this way, the temporal dependency of spinal rotation during the turning operation is modeled, capturing the association between spinal motion states at different moments, such as the influence of the thoracic vertebrae rotation angle at the previous moment on the lumbar vertebrae rotation angle at the current moment.

[0049] Preferably, the gated recurrent unit network can be parameter optimized through the back-propagation algorithm so that the output feature vector can accurately characterize the dynamic change law of the turning action. After network processing, the output layer generates a dynamic encoding vector of the turning action, which contains feature information of different stages such as the starting stage, rotation stage, and reset stage of the turning action, and is stored in the form of a fixed-dimensional numerical sequence.

[0050] S22: Perform nursing intention analysis on the voice command data of the multimodal dataset, call the preset clinical term matching model to identify the key action instructions of the turning operation, and generate the nursing operation type code.

[0051] Specifically, the system extracts voice command data from the multimodal dataset, which is a time-aligned word timestamp sequence and corresponding text information. The system calls a preset clinical term matching model, which is built based on clinical nursing operation specifications and contains a key term library related to turning operations. The term library stores command phrases related to turning operations such as "turn over to the left", "keep stable", and "rotate slowly" and corresponding operation type identifiers.

[0052] For example, the system performs word segmentation on the voice instruction text, divides the continuous text into independent words or phrases, removes meaningless function words, and obtains a valid word sequence. The clinical term matching model traverses the valid word sequence, compares each word with the key terms in the term library, calculates the semantic similarity, and when the similarity reaches a preset threshold, determines that the word is a key action instruction. The system summarizes the identified key action instructions and generates a nursing operation type code based on the corresponding operation type identifier in the term library. The code is a string of numbers or character combinations, and each code corresponds to a specific turning operation type, such as code "001" corresponds to "45-degree turning on the left side", code "002" corresponds to "90-degree turning on the right side", etc.

[0053] S23: Gradient analysis is performed on the sacral region pressure data of the multimodal dataset, and the pressure ulcer prevention pressure distribution change rate is calculated using a spatial difference algorithm.

[0054] Specifically, the system extracts sacral pressure data from a multimodal dataset. This data is a sequence of time-aligned pressure distribution matrices, each containing the pressure values ​​and corresponding timestamps for each monitoring point in the sacral region. The system processes the pressure distribution matrix using a spatial difference algorithm. First, the spatial coordinate system for the pressure distribution is determined. A three-dimensional rectangular coordinate system is established with the geometric center of the sacrum as the origin. The coordinates of each monitoring point are determined based on its actual position in the sacral region.

[0055] For the pressure distribution matrix at each time point, the system calculates the pressure difference between adjacent monitoring points and obtains the pressure gradient by the ratio of the pressure difference to the spatial distance. Preferably, the system performs neighborhood operations on the pressure distribution matrix using a spatial difference algorithm to obtain the pressure gradient of each monitoring point at that time point. x 、 y 、 z The pressure change rate in three directions is then vector-synthesized to obtain the pressure distribution change rate at that point. The system statistically analyzes the pressure distribution change rate at all monitoring points, calculating the average pressure change rate, maximum pressure change rate, and standard deviation of the pressure change rate distribution for the entire sacral region. These are used as characteristic parameters of the pressure distribution change rate for pressure ulcer prevention. These parameters are stored in numerical form and include corresponding time information.

[0056] S24: Perform cross-modal fusion processing on the dynamic coding vector of turning action, nursing operation type coding and pressure ulcer prevention pressure distribution change rate, and integrate multi-source features through weighted attention mechanism to generate a spatiotemporal fusion feature sequence.

[0057] Specifically, the system constructs an attention mechanism model, which includes a feature mapping module, a weight calculation module, and a fusion module. The feature mapping module maps the dynamic encoding vector of the turning action, the nursing operation type code, and the pressure distribution change rate for pressure ulcer prevention to the same high-dimensional feature space, making the features of different modalities comparable. The weight calculation module determines the attention weight by calculating the correlation between the features of different modalities. For example, the feature part related to the nursing operation type code in the dynamic encoding vector of the turning action is assigned a higher weight, and the pressure change feature corresponding to the turning action in the pressure distribution change rate for pressure ulcer prevention is assigned a higher weight. The fusion module performs a weighted summation of the mapped multi-source features according to the attention weight to obtain a fused feature vector. The system arranges the fused feature vectors at different times in chronological order to form a spatiotemporal fusion feature sequence. This sequence contains feature change information in the time dimension and multimodal feature association information in the spatial dimension. Each feature vector corresponds to a specific timestamp, which fully reflects the dynamic association of each modal data during the turning operation.

[0058] In one embodiment, Figure 2 As shown, S3 of the multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S31: Perform venipuncture motion analysis on the spatiotemporal fusion feature sequence, identify the needle stability characteristics by analyzing the elbow flexion and extension angle change pattern, and generate venipuncture stability parameters.

[0059] Specifically, the system extracts a feature subset related to venipuncture from the spatiotemporal fusion feature sequence. This subset includes characteristic information about the changes in elbow flexion and extension angles over time, as well as other auxiliary features related to the puncture action. Preferably, the system focuses on the change pattern of elbow flexion and extension angles, and uses a feature extraction algorithm to isolate the elbow angle data during the puncture process. This data includes elbow flexion and extension angle values ​​corresponding to different time points.

[0060] The system processes the extracted elbow flexion and extension angle data and calculates the elbow angle change rate, that is, the absolute value of the angle change rate is obtained by the ratio of the angle difference between adjacent time points to the time interval. Preferably, the system uses a preset formula to calculate the venipuncture stability parameter, which is: ; in, is the venipuncture stability parameter, The total duration of the venipuncture operation, that is, the time interval from the start to the end of the puncture; The elbow flexion and extension angles are obtained through the angle data in the feature subset; is the sensitivity adjustment coefficient, which is used to adjust the influence of the angle change rate on the stability parameters; is the absolute value of the elbow joint angle change rate. The system solves the formula through the numerical integration algorithm, and calculates the Perform integral calculation and divide by the total operation time , get the venipuncture stability parameter , a parameter used to quantify the stability characteristics of the needle during venipuncture.

[0061] S32: Perform force control analysis on the spatiotemporal fusion feature sequence, identify pressure control defects by modeling the evolution of sacral pressure distribution, and generate pressure ulcer prevention force control parameters.

[0062] Specifically, the system performs force control analysis on the spatiotemporal fusion feature sequence for turning over. A subset of sacral pressure distribution features related to the turning over operation is extracted from the spatiotemporal fusion feature sequence. This subset contains sacral pressure data at different sampling points and the corresponding time information. The system processes these pressure data, models the evolution of sacral pressure distribution, and identifies possible defects in the pressure control process by analyzing the changes in pressure distribution over time, such as pressure mutations and uneven pressure distribution. Based on the processed sacral pressure data, the system uses a preset formula to calculate the pressure ulcer prevention force control parameters. The formula is: ; in, Control parameters for pressure ulcer prevention, is the number of sampling points for the turning operation, that is, the total number of points where pressure data is collected during the turning process; For the The sacral pressure gradient change rate at the sampling point is obtained by calculating the spatial gradient of the sacral pressure distribution data; is the pressure mutation threshold, which is used to determine whether the pressure gradient change is a mutation; K is the normalization coefficient, which is used to map the calculation results to a specific numerical range. The system calculates , then sum these values ​​and divide them by the number of sampling points N and the normalization coefficient K, and subtract the result from 1 to obtain the pressure ulcer prevention force control parameter F, which is used to indicate the effect of pressure control during the turning operation.

[0063] S33: Perform fluency assessment on the intravenous puncture stability parameters and pressure ulcer prevention force control parameters, calculate the nursing operation continuity index through time window integration, and generate a mechanical parameter sequence. The mechanical parameter sequence is used to indicate the intravenous puncture stability level and the turning force control ability score.

[0064] Specifically, the system uses a time window integration method to set a fixed time window, integrates the stability parameter and force control parameter within the window, and generates a mechanical parameter sequence by calculating the nursing operation continuity index. This sequence not only contains the stability level information of the venipuncture, but also contains the force control ability score of the turning operation. Preferably, the mechanical parameter sequence is obtained by the following steps: S331: Perform time window integration processing on the venipuncture stability parameters, calculate the average stability level of the elbow joint angle change rate through a sliding window, and generate a venipuncture stability level.

[0065] Specifically, the system performs venipuncture stability parameters Perform time window integration processing, assuming that the sliding window length is 500ms, the sliding step is 100ms, and the window start time is Time from puncture To the end time , covering interval The system calculates the average stability level within the window: ; Where S(t) is the stability parameter at any time in the window, and the integration is done using the Simpson method: ; in, , the system calculates the fluctuation value of the angle change rate within the window: ; Where n is the number of sampling points in the window, is the angle change rate of the kth sampling point in the window, the average value .

[0066] The stability level is divided based on and Combination conditions: When In the highest range and When it is in the lowest range, the level is the highest level; when In the second highest range and When it is in the second lowest range, the level is the second highest, and so on. The level array contains timestamp, 、 And the corresponding stability level L.

[0067] S332: Perform time window integration processing on the pressure ulcer prevention force control parameters, evaluate the smoothness of sacral pressure changes through a sliding window, and generate a turning force control ability score.

[0068] Specifically, the system performs time window integration processing on the pressure sore prevention force control parameter F, assuming that the window length is 500ms and the step length is 100ms. Start by turning over To the end , interval Integration formula: ; The integration uses the Simpson method, which is consistent with step S331 and will not be repeated here. The system calculates the root mean square of the pressure gradient change rate within the window: ; Where m is the number of sampling points in the window, is the pressure gradient change rate of the kth sampling point in the window. The force control ability scoring rule is based on and Combination conditions: When In the highest range and When it is in the lowest range, the score is the highest level; when In the second highest range and When it is in the second lowest range, the score is the second highest, and so on. The score array contains timestamp, 、 And the corresponding score S, the operation preparation stage and the end stage are filled with zero values.

[0069] S333: Perform temporal combination processing on the venipuncture stability grade and the turning force control ability score, construct a dual-parameter sequence according to the operation time sequence, and generate a mechanical parameter sequence.

[0070] Specifically, the system performs time-series combination processing on the intravenous puncture stability level and the turning over force control ability score, extracts the start and end timestamps of the time window corresponding to each level in the intravenous puncture stability level sequence, and simultaneously extracts the start and end timestamps of the time window corresponding to each score in the turning over force control ability score sequence. The system compares the timestamps of the two sequences and associates the stability level and force control ability score to the same time reference through a time alignment algorithm to ensure that the two parameters at the same time point or within the time window can be accurately matched. For windows with time overlap, the system takes the time interval of the overlapping part as the common time marker of the corresponding parameters; for independent operation stages without overlap, each time marker and corresponding parameter value are retained.

[0071] The system arranges the matched venipuncture stability levels and turn-over force control scores in chronological order, constructing a two-parameter sequence. Each element contains a time stamp, a corresponding venipuncture stability level, and a turn-over force control score. This two-parameter sequence is the generated mechanical parameter sequence, fully recording the dynamic changes in venipuncture stability and turn-over force control throughout the nursing procedure.

[0072] In summary, the multimodal data-driven nursing training method provided by this application can realize accurate quantitative evaluation of micro-jitter phenomenon during intravenous puncture by analyzing the change pattern of elbow flexion and extension angle to identify the stability characteristics of the needle, so as to solve the technical limitation that traditional manual observation cannot capture instantaneous stability changes; based on the pressure control defect modeling of the sacral pressure distribution evolution process, it can achieve dynamic tracking of abnormal evolution of pressure gradient during turning operation, and overcome the defect of inaccurate judgment of force control continuity of single-point pressure monitoring; through the time window integration mechanism to calculate the nursing operation continuity index, it can realize the full process coupling evaluation of intravenous puncture stability and turning force control ability, so as to eliminate the disadvantage of static scoring model losing operation stage characteristics; the final generated mechanical parameter sequence can achieve fine-grained analysis of nursing operation quality through two-dimensional quantitative indicators, and provide traceable and verifiable biomechanical basis for adaptive training. This method systematically improves the objectivity and timeliness of clinical nursing operation evaluation, and converts abstract operation specifications into quantifiable dynamic parameters, which not only solves the problem of single evaluation dimension in traditional training, but also overcomes the defect of low skill transfer efficiency caused by feedback lag.

[0073] In one embodiment, Figure 3 As shown, S4 of the multimodal data-driven caregiver training method provided by the present invention specifically includes the following steps: S41: Identify venipuncture defects based on the mechanical parameter sequence, analyze the deviation between the elbow joint angle deviation trend and the standard operation through the decision tree model, and generate a puncture stability defect report.

[0074] Specifically, the system processes the mechanical parameter sequence for venipuncture defect identification and extracts a subsequence related to venipuncture from the mechanical parameter sequence. This subsequence contains the venipuncture stability level and the corresponding time stamp. The system then uses a pre-defined decision tree model trained based on standard clinical operation data. The model contains multiple decision nodes, each corresponding to a characteristic threshold for elbow joint angle deviation, such as angle deviation, deviation duration, and deviation frequency.

[0075] Exemplarily, the system inputs the elbow joint angle deviation trend features in the venipuncture subsequence into a decision tree model. Features include angle deviation values ​​at different time points, the rate of change of deviation within a continuous window, and so on. The decision tree model analyzes the degree of deviation between the elbow joint angle deviation trend and standard operation by comparing the feature values ​​and node thresholds layer by layer. For example, when the angle deviation exceeds the first threshold and lasts for more than a preset time, it is judged as a "moderate deviation defect"; when the deviation exceeds the second threshold and is accompanied by high-frequency fluctuations, it is judged as a "severe jitter defect."

[0076] The system summarizes the defect type, occurrence time, specific deviation values, and corresponding operation phases output by the model to generate a puncture stability defect report. The report includes the defect number, defect description, defect severity, occurrence time interval, and corresponding mechanical parameter abnormalities, providing a specific basis for subsequent course matching.

[0077] S42: Identify turning force control defects in the mechanical parameter sequence, locate abnormal change points of sacral pressure gradient through mutation detection algorithm, and generate a pressure ulcer prevention force control defect report.

[0078] Specifically, the system processes the mechanical parameter sequence to identify defects in turning force control. A subsequence corresponding to the turning operation is extracted from the mechanical parameter sequence. This subsequence contains information on the temporal changes in the turning force control ability score and sacral pressure-related parameters. The system then uses a mutation detection algorithm, which constructs a baseline model of the pressure gradient change rate based on a sliding window. Anomalies are identified by calculating the residual between the real-time pressure gradient change rate and the baseline model.

[0079] For example, the system sets a residual threshold. When the residual at a certain time point exceeds the threshold, the point is determined to be an abnormal change point of the sacral pressure gradient. The algorithm also analyzes the change amplitude, duration, and correlation with adjacent points of the abnormal point to distinguish between transient interference and substantial pressure mutations. For example, when the fluctuation of the pressure gradient change rate within three consecutive windows exceeds the preset range, and the single change exceeds the pressure mutation threshold, it is marked as a "pressure surge defect"; when the change rate is continuously lower than the normal range, it is marked as an "insufficient pressure defect."

[0080] The system integrates the identified anomaly information to generate a pressure ulcer prevention force control deficiency report. The report includes the anomaly timestamp, the difference between the measured pressure gradient change rate and the standard value, the defect type classification, and the corresponding turning operation stage, clarifying the specific manifestations of force control defects during the turning process.

[0081] S43: Match the puncture stability defect report and the pressure ulcer prevention force control defect report to the courses, search the standardized course library based on the pre-existing clinical nursing operation classification system, match the skill defect type and the training course mapping relationship, and generate a personalized training course combination.

[0082] Specifically, the system performs course matching processing on puncture stability defect reports and pressure ulcer prevention force control defect reports, and calls the pre-stored clinical nursing operation classification system, which divides subcategories according to operation type, skill module, and defect level. For example, the intravenous puncture module contains subcategories such as "needle insertion stability" and "angle control", and the turning module contains subcategories such as "pressure distribution" and "force application rhythm".

[0083] Exemplarily, the system retrieves a standardized course library based on a classification system, and the course library stores attribute labels for various training courses, including corresponding skill modules, targeted defect types, training objectives, and difficulty levels. The system establishes a mapping relationship table between skill defect types and training courses. Each record in the table contains a defect type identifier, a recommended course number, and a matching weight. The system compares the defect type in the defect report with the mapping relationship table, calculates the matching degree, and filters out the courses with the highest matching weight. For complex defects, the system sorts them by defect severity, prioritizes matching courses with high weights, and forms a combination plan containing multiple courses. The generated personalized training course combination includes the course name, course number, training duration, targeted defect points, and expected improvement goals to ensure that the training content accurately corresponds to the skill shortcomings.

[0084] S44: Perform VR instruction conversion on the personalized training course combination, generate an executable training task instruction set through the virtual reality interface protocol, and send the training task instruction set to the virtual reality training terminal.

[0085] Specifically, the system converts VR commands for personalized training course combinations. The system loads the virtual reality interface protocol, which defines the instruction format for training scenario parameters, interaction logic, and feedback mechanisms, supporting standardized communication with virtual reality training terminals. The system parses the content of the personalized training course combinations and converts the training objectives within the courses into specific parameters for the VR scenarios, such as the diameter, position, and movement speed of the virtual blood vessels in the venipuncture course, and the body shape parameters and pressure feedback thresholds of the virtual patients in the turning course. Simultaneously, operational requirements are converted into interactive commands, such as the angle range of the puncture action, the numerical range of the force applied, and the timing control of the turning action. For example, the system generates an executable training task instruction set in a protocol format. The instruction set includes scenario initialization parameters, dynamic adjustment rules, assessment indicators, and feedback signal definitions. This instruction set is then transmitted via wired or wireless communication to a virtual reality training terminal, which includes a head-mounted display (HMD), motion capture sensors, and force feedback devices. Once the system confirms that the terminal has received the instruction, it triggers the terminal to initiate the corresponding training scenario, enabling immersive presentation of personalized training content.

[0086] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0087] Based on the same inventive concept, the embodiments of the present application also provide a multimodal data-driven caregiver training device for implementing the multimodal data-driven caregiver training method involved above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more embodiments of the multimodal data-driven caregiver training device provided below can be found in the above-mentioned limitations of the multimodal data-driven caregiver training method, and will not be repeated here.

[0088] Preferably, if Figure 4 As shown, the present invention provides a multimodal data-driven caregiver training device 500, which is configured with the following modules: The multimodal data time synchronization module 510 is used to synchronize the joint motion data, voice command data, and hand pressure data collected during the caregiver's operation, align the timestamps of the joint motion sensor clock, the voice acquisition device clock, and the pressure sensor clock, and generate a time-aligned multimodal data set; Cross-modal feature fusion module 520 is used to perform cross-modal feature fusion processing on the time-aligned multi-modal dataset, extract the dynamic features of the nursing action through a temporal neural network and calculate the gradient features of the palm pressure distribution to generate a spatiotemporal fusion feature sequence; The micro-motion analysis module 530 is used to perform micro-motion analysis on the spatiotemporal fusion feature sequence, decode the elbow flexion and extension angle change rate and the palm key area pressure change parameters through a preset probability generation model, calculate the nursing action fluency continuity index, and generate a mechanical parameter sequence; The adaptive feedback decision module 540 is used to make adaptive feedback decisions on the mechanical parameter sequence, identify weak links in skills based on the abnormal offset rate of joint angles and the sudden change value of pressure gradient, match the preset clinical nursing standardized course library, generate a personalized training task instruction set, and send the training task instruction set to the virtual reality training terminal. The personalized training task instructions are used to indicate the intensive training content for the weak links in the caregiver's skills.

[0089] Preferably, the multimodal data time synchronization module 510 provided in this application is configured with the following units: The elbow joint motion data acquisition and analysis unit is used to collect and process the caregiver's elbow joint motion data in real time through a wearable inertial measurement unit, perform motion velocity analysis on the acquired three-dimensional angular velocity vector, calculate the synthetic linear velocity within the joint motion plane, and generate the wrist joint linear velocity characteristic value; The voice command timing tagging unit is used to perform timing tagging on the voice command data given to the caregiver, calling the endpoint detection algorithm to segment the key instruction phrases of the nursing operation and generate a word timestamp sequence; The palm pressure monitoring and compensation unit is used to monitor and process the palm contact pressure in different areas through a flexible sensor array, and to perform dynamic delay compensation processing on the pressure distribution vectors of the sacral area and thenar area of ​​the palm to generate a multi-device clock synchronization offset; The multimodal data timeline reconstruction unit is used to perform timeline reconstruction processing on the wrist joint linear velocity feature values, word timestamp sequences, and multi-device clock synchronization offsets, mapping the original data to a unified time base to generate a time-aligned multimodal dataset.

[0090] Preferably, the cross-modal feature fusion module 520 provided in this application is configured with the following units: The turning action feature extraction unit is used to extract and process the turning action features from the joint motion data of the multimodal dataset, model the temporal dependency of spinal rotation during the turning operation through a gated recurrent unit network, and generate a dynamic encoding vector for the turning action; The nursing intention parsing unit is used to parse the voice instruction data of the multimodal data set for nursing intention, call the preset clinical term matching model to identify the key action instructions of the turning operation, and generate the nursing operation type code; The pressure gradient analysis unit is used to perform gradient analysis on the sacral region pressure data of the multimodal data set and calculate the pressure distribution change rate for pressure ulcer prevention using a spatial difference algorithm; The cross-modal feature fusion processing unit is used to perform cross-modal fusion processing on the dynamic coding vector of the turning action, the nursing operation type coding and the pressure distribution change rate of pressure ulcer prevention, and to generate a spatiotemporal fusion feature sequence by weighted integration of multi-source features through the attention mechanism.

[0091] Preferably, the micro-action analysis module 530 provided in this application is configured with the following units: The venipuncture motion analysis unit is used to perform venipuncture motion analysis on the spatiotemporal fusion feature sequence, identify the needle stability characteristics by analyzing the elbow flexion and extension angle change pattern, and generate venipuncture stability parameters; The turning force control analysis unit is used to perform turning force control analysis on the spatiotemporal fusion feature sequence, identify pressure control defects by modeling the evolution of sacral pressure distribution, and generate pressure ulcer prevention force control parameters; The operation fluency assessment unit is used to perform fluency assessment on the intravenous puncture stability parameters and pressure ulcer prevention force control parameters, calculate the nursing operation continuity index through time window integration, and generate a mechanical parameter sequence. The mechanical parameter sequence is used to indicate the intravenous puncture stability level and the turning force control ability score.

[0092] Preferably, the operation fluency assessment unit includes a venipuncture stability level calculation subunit, a turning force control ability score calculation subunit, and a mechanical parameter sequence construction subunit. Among them, the venipuncture stability level calculation subunit is used to perform time window integration processing on the venipuncture stability parameters, calculate the average stability level of the elbow joint angle change rate through a sliding window, and generate the venipuncture stability level; the turning force control ability score calculation subunit is used to perform time window integration processing on the pressure sore prevention force control parameters, evaluate the smoothness of sacral pressure changes through a sliding window, and generate the turning force control ability score; the mechanical parameter sequence construction subunit is used to perform time sequence combination processing on the venipuncture stability level and the turning force control ability score, construct a dual parameter sequence according to the operation time sequence, and generate the mechanical parameter sequence.

[0093] Preferably, the adaptive feedback decision module 540 provided in this application is configured with the following units: Venipuncture defect recognition unit, used to identify venipuncture defects based on mechanical parameter sequences, analyze the deviation between elbow joint angle deviation trends and standard operations through a decision tree model, and generate a puncture stability defect report; The turning force control defect identification unit is used to identify turning force control defects in the mechanical parameter sequence, locate abnormal change points of sacral pressure gradient through mutation detection algorithm, and generate a pressure ulcer prevention force control defect report; The training course matching unit is used to match courses for puncture stability deficiency reports and pressure ulcer prevention force control deficiency reports. Based on the pre-existing clinical nursing operation classification system, the standardized course library is searched and the mapping relationship between skill deficiency types and training courses is matched to generate personalized training course combinations. The VR training instruction conversion unit is used to convert the personalized training course combination into VR instructions, generate an executable training task instruction set through the virtual reality interface protocol, and send the training task instruction set to the virtual reality training terminal.

[0094] In one embodiment, the present application further provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned multimodal data-driven caregiver training method when executing the computer program.

[0095] In one embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned multimodal data-driven caregiver training method when executed by a processor.

[0096] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and integrate different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0097] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0098] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A multimodal data-driven caregiver training method, characterized in that: The following steps are involved: S1: Time-synchronize the joint motion data, voice command data, and hand pressure data collected by the caregiver during the operation, align the timestamps of the joint motion sensor clock, voice acquisition device clock, and pressure sensor clock, and generate a time-aligned multimodal dataset; S2: performing cross-modal feature fusion processing on the time-aligned multimodal dataset, extracting dynamic features of nursing actions through a temporal neural network and calculating gradient features of palm pressure distribution to generate a spatiotemporal fusion feature sequence; S3: Performing micro-motion analysis on the spatiotemporal fusion feature sequence, decoding the elbow flexion and extension angle change rate and the palm key area pressure change parameters through a preset probability generation model, and calculating the nursing action fluency continuity index to generate a mechanical parameter sequence; S4: Perform adaptive feedback decision on the mechanical parameter sequence, identify weak links in skills based on abnormal joint angle offset rate and pressure gradient mutation value, match the preset clinical nursing standardized course library, generate a personalized training task instruction set, and send the training task instruction set to the virtual reality training terminal. The personalized training task instruction is used to indicate the intensive training content for the weak links in the caregiver's skills.

2. The method according to claim 1, characterized in that Said S1 comprises: S11: The wearable inertial measurement unit is used to collect and process the caregiver's elbow joint motion data in real time, and the acquired three-dimensional angular velocity vector is analyzed and processed to calculate the synthetic linear velocity in the joint motion plane and generate the wrist joint linear velocity characteristic value; S12: Perform time-series tagging on the voice instruction data given to the caregiver, call the endpoint detection algorithm to segment the key instruction phrases of the nursing operation, and generate a word timestamp sequence; S13: The flexible sensor array is used to monitor the palm contact pressure in different areas, and dynamic delay compensation is performed on the pressure distribution vectors of the sacral area and thenar area of ​​the palm to generate a multi-device clock synchronization offset. S14: Performing time axis reconstruction processing on the wrist joint linear velocity feature value, the word timestamp sequence and the multi-device clock synchronization offset, mapping the original data to a unified time reference, and generating a time-aligned multimodal data set.

3. The method according to claim 1, characterized in that The S2 includes: S21: performing turning action feature extraction processing on the joint motion data of the multimodal dataset, modeling the temporal dependency of spinal rotation in the turning operation through a gated recurrent unit network, and generating a turning action dynamic encoding vector; S22: performing nursing intention analysis on the voice instruction data in the multimodal dataset, calling a preset clinical term matching model to identify key action instructions for turning over, and generating a nursing operation type code; S23: performing gradient analysis on the sacral region pressure data of the multimodal dataset, and calculating a pressure ulcer prevention pressure distribution change rate using a spatial difference algorithm; S24: Perform cross-modal fusion processing on the dynamic coding vector of the turning action, the nursing operation type code and the pressure ulcer prevention pressure distribution change rate, and weightedly integrate multi-source features through the attention mechanism to generate a spatiotemporal fusion feature sequence.

4. The method according to claim 1, wherein The S3 includes: S31: performing venipuncture motion analysis on the spatiotemporal fusion feature sequence, identifying needle stability features by analyzing the elbow flexion and extension angle change pattern, and generating venipuncture stability parameters; S32: performing turning force control analysis processing on the spatiotemporal fusion feature sequence, identifying pressure control defects by modeling the evolution of sacral pressure distribution, and generating pressure ulcer prevention force control parameters; S33: Perform a fluency assessment on the intravenous puncture stability parameter and the pressure ulcer prevention force control parameter, calculate the nursing operation continuity index through time window integration, and generate a mechanical parameter sequence. The mechanical parameter sequence is used to indicate the intravenous puncture stability level and the turning force control ability score.

5. The method according to claim 4, characterized in that The calculation formulas for the venipuncture stability parameter and the pressure sore prevention force control parameter are: ; in, is the venipuncture stability parameter, Control parameters for pressure ulcer prevention, is the total duration of the venipuncture procedure, is the elbow flexion and extension angle, is the sensitivity adjustment coefficient, is the absolute value of the elbow joint angle change rate, The number of sampling points for turning operation, For the The sacral pressure gradient change rate at the sampling point, is the pressure mutation threshold, is the normalization coefficient.

6. The method according to claim 4, characterized in that The S33 includes: S331: performing time window integration processing on the venipuncture stability parameter, calculating the average stability level of the elbow joint angle change rate through a sliding window, and generating a venipuncture stability level; S332: performing time window integration processing on the pressure sore prevention force control parameter, evaluating the smoothness of sacral pressure changes through a sliding window, and generating a turning force control ability score; S333: Performing time-series combination processing on the venipuncture stability grade and the turning-over force control ability score, constructing a dual-parameter sequence according to the operation time sequence, and generating a mechanical parameter sequence.

7. The method according to any one of claims 1 to 6, characterized in that The S4 includes: S41: performing venipuncture defect identification on the mechanical parameter sequence, analyzing the degree of deviation between the elbow joint angle deviation trend and the standard operation through a decision tree model, and generating a puncture stability defect report; S42: identifying turning force control defects on the mechanical parameter sequence, locating abnormal change points of sacral pressure gradient using a mutation detection algorithm, and generating a pressure ulcer prevention force control defect report; S43: Matching the puncture stability deficiency report and the pressure sore prevention force control deficiency report to a course, searching a standardized course library based on a pre-stored clinical nursing operation classification system, matching skill deficiency types with training course mappings, and generating a personalized training course combination; S44: Perform VR instruction conversion on the personalized training course combination, generate an executable training task instruction set through a virtual reality interface protocol, and send the training task instruction set to the virtual reality training terminal.

8. A multimodal data-driven caregiver training device, characterized in that: The device comprises: The multimodal data time synchronization module is used to synchronize the joint motion data, voice command data, and hand pressure data collected by the caregiver during the operation, align the timestamps of the joint motion sensor clock, the voice acquisition device clock, and the pressure sensor clock, and generate a time-aligned multimodal data set; A cross-modal feature fusion module is used to perform cross-modal feature fusion processing on the time-aligned multimodal dataset, extract the dynamic features of the nursing action through a temporal neural network, calculate the gradient features of the palm pressure distribution, and generate a spatiotemporal fusion feature sequence; A micro-motion analysis module is used to perform micro-motion analysis on the spatiotemporal fusion feature sequence, decode the elbow flexion and extension angle change rate and the palm key area pressure change parameters through a preset probability generation model, calculate the nursing action fluency continuity index, and generate a mechanical parameter sequence; An adaptive feedback decision module is used to make adaptive feedback decisions on the mechanical parameter sequence, identify weak links in skills based on abnormal joint angle offset rates and pressure gradient mutation values, match a preset clinical nursing standardized course library, generate a personalized training task instruction set, and send the training task instruction set to a virtual reality training terminal. The personalized training task instruction is used to indicate the intensive training content for the caregiver's weak links in skills.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Hospital nursing assessment management system

    CN111863218A

  • Unmanned public nursing and first-aid training system and training method

    CN111915943A

  • Nursing practice teaching demonstration system based on big data

    CN114743428A

  • Intelligent caregiver massage examination system based on computer vision

    CN118485337A

  • Ball action evaluation method and device based on large model, equipment and medium

    CN120123702A

Cited By

  • Nursing operation evaluation system and method based on multi-modal large model

    CN122222459A